Why can a Linux program call clock_gettime() millions of times per second without paying the cost of entering the kernel on each call?
answer
- some calls need no privilege at all
- the kernel maps a page into you
- shared read-only timekeeping data
- linux-vdso.so.1, a library with no file
basics
~20 sThe kernel maps a small shared object, the vDSO, into every process. For a few read-only calls such as clock_gettime it exports real user-space code that reads kernel-maintained timekeeping data from a shared page, so the call returns without any trap into kernel mode.
solid answer
~50 sSome system calls only read information the kernel is happy to publish, and trapping for them is pure overhead. Linux therefore maps a small kernel-provided shared object — the vDSO, which shows up as `linux-vdso.so.1` in a process's mappings — into every address space at exec time. It contains real user-space implementations of a handful of calls, notably `clock_gettime()`, `gettimeofday()` and `getcpu()`, which compute their answer by reading a kernel-maintained page of timekeeping data plus the CPU's timestamp counter. glibc resolves those symbols through the vDSO when it can, so the call is an ordinary function call with no mode switch. It is not unconditional: if the system's clocksource does not support the fast path, the vDSO routine falls back to the real system call, and the same code can suddenly become an order of magnitude more expensive on that host.
code
c · 18 lines#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <time.h>
int main(void) {
struct timespec start, end;
clock_gettime(CLOCK_MONOTONIC, &start);
for (long i = 0; i < 10000000L; i++) {
struct timespec t;
clock_gettime(CLOCK_MONOTONIC, &t); /* served from the vDSO when it can be */
}
clock_gettime(CLOCK_MONOTONIC, &end);
double ns = (end.tv_sec - start.tv_sec) * 1e9 + (end.tv_nsec - start.tv_nsec);
printf("%.1f ns per call\n", ns / 10000000.0);
return 0;
}go deeper
It is enough to know that reading the clock is unusually cheap on Linux because the kernel shares the data instead of requiring a trap, and that the mechanism is called the vDSO.
Explain the mechanism: a kernel-provided shared object mapped into every process, a read-only data page of timekeeping parameters, and user-space code that scales a counter reading to produce the answer.
Bring the operational angle: the fast path is conditional on the active clocksource, virtualised hosts often lose it, and the resulting cost difference is a real and frequently misdiagnosed hotspot.
Frame it as a general design pattern — moving read-only, high-frequency kernel state into a shared mapping — and be ready to say where that pattern is safe to apply and where a privilege check makes the trap unavoidable.
## The problem the vDSO solves Time is the most-called thing in a server. Loggers stamp every line, metrics libraries bracket every operation, timeout logic re-reads the clock in loops. If every one of those cost a full trap into the kernel and back, a busy service would spend a measurable share of its CPU doing nothing but asking what time it is. But asking the time does not actually need kernel privilege. The kernel maintains the data required to answer — a base wall-clock time, a base counter value, and multiply/shift factors that convert counter ticks to nanoseconds — and there is no security reason a process cannot read those numbers. What it needs is a trustworthy, consistent copy and code that knows how to use it. ## What the vDSO is The virtual dynamic shared object is a small ELF shared library that lives inside the kernel image and is mapped into the address space of every process at `execve()` time. It is not a file on disk: you will never find it under `/lib`, and the dynamic linker does not open it. It appears in a process's mapping list under the name `linux-vdso.so.1`, which is why that entry shows up in every process yet corresponds to nothing in the filesystem. The kernel tells the process where it is through the auxiliary vector — the `AT_SYSINFO_EHDR` entry that the kernel places on the stack at exec alongside `argv` and the environment. The C library's startup code reads that entry, parses the small ELF object it points at, and records the addresses of the symbols it exports. Statically linked programs do this too; the vDSO is not a dynamic-linking feature, it is a kernel-to-process handoff. ## How the fast path works Alongside the code page the kernel maps a read-only data page containing the current timekeeping parameters, which it updates whenever the clock is adjusted or the timer tick advances. The vDSO's `clock_gettime()` implementation reads that data under a sequence-lock protocol — read the sequence counter, read the values, read the counter again, retry if it changed — so it can get a consistent snapshot without any lock a user process could hold or corrupt. It then reads the CPU's timestamp counter with `rdtsc` on x86-64, scales the delta by the published factors, adds the base, and returns. The whole thing is a handful of instructions and a couple of cache-resident loads. There is no privilege change, no register save, no kernel stack switch — nothing that distinguishes it from calling any other library function. ## Which calls get this treatment The exported set is deliberately tiny and architecture-specific. On x86-64 it covers `clock_gettime()`, `gettimeofday()`, `time()` and `getcpu()`, with a corresponding `clock_getres()`. That is the whole idea: only calls that are read-only, extremely frequent, and require no permission check are candidates. There is no vDSO `read()` or `open()`, because those genuinely require the kernel to act on shared state and enforce policy. ## When it silently falls back This is the part worth knowing for production. The fast path depends on the active clocksource being one the vDSO can read directly from user space — in practice the CPU timestamp counter, which the kernel selects only when it has judged it constant and reliable across cores. If the kernel has instead selected a clocksource that must be read through kernel-only means, the vDSO routine cannot compute the answer itself and issues the real system call. That produces one of the more baffling performance reports: identical binaries, identical load, and one host where a timestamp costs tens of nanoseconds and another where it costs hundreds or more. Virtualised environments are the usual culprit, because the hypervisor may not expose a counter the guest kernel is willing to trust. The current selection is published under `/sys/devices/system/clocksource/`, and it is one of the first things to check when time-stamping suddenly becomes a hotspot. Other consequences follow from the same mechanism. Because the vDSO is user-space code reading a shared page, a syscall-level sandbox or tracer sees nothing at all for these calls — there is no trap to intercept. And because the returned time derives from a counter reading, `CLOCK_MONOTONIC` and `CLOCK_REALTIME` differ exactly as they do through the syscall path: monotonic never jumps backwards, realtime can, so timing intervals with realtime remains a bug regardless of how the call is serviced. ## The deprecated ancestor An older mechanism, `vsyscall`, mapped a fixed page at a fixed address for the same purpose. Its fixed address made it a useful gadget source for exploits, so it was replaced by the vDSO, which is placed at a randomised address like any other mapping. Modern kernels keep only a compatibility emulation of it for very old binaries, and it is not something new code should reference.
- Why are only a few calls exposed through the vDSO rather than most of them?Because the technique only works for calls that read published state and enforce no policy. Timekeeping data can be safely shared read-only with every process, so user-space code can compute the answer itself. Anything that mutates shared state or must apply a permission check — opening files, sending packets, creating processes — fundamentally requires the kernel to act, so no amount of clever mapping removes the trap.
- What does the vDSO imply for a tool or sandbox that works at the system call boundary?It means those calls are invisible to it. A vDSO-served `clock_gettime()` never traps, so a syscall tracer records nothing and a seccomp filter has no decision point to act on. That is usually harmless for time, but it is worth remembering when reasoning about what a syscall-level observation or restriction actually covers — the boundary is complete only for calls that genuinely cross it.
- Two identical hosts show very different timestamp costs. Where would you look?At the active clocksource. The vDSO fast path needs a counter it can read from user space, and if the kernel has selected a different clocksource — common under virtualisation where the guest cannot trust the CPU counter — the vDSO routine falls back to a real system call and the cost jumps by an order of magnitude. The kernel publishes the current selection under /sys/devices/system/clocksource/.
saying these in an interview costs you the question
- Calls the vDSO a shared library loaded from disk
- Thinks the vDSO can accelerate any system call
- Assumes clock_gettime never traps into the kernel
- Says the vDSO lets user code write kernel memory