skip to content

A Linux network service burns most of its CPU in kernel mode because it issues a very large number of small system calls. Why is each crossing of the user/kernel boundary expensive, and what changes make the same work cost fewer crossings?

level: seniorimportance: should knowfreq 45%

answer

  1. the crossing itself is the cost
  2. fixed overhead, whatever the payload
  3. mode switch is not a context switch
  4. same bytes, fewer calls

basics

~20 s

Each crossing costs a privilege transition plus register saving, argument validation and, on mitigated CPUs, page-table and speculation work — hundreds of nanoseconds of pure overhead per call. The fix is to move more bytes or more events per call: bigger buffers, vectored and zero-copy I/O, and batching event APIs.

solid answer

~50 s

A system call is not a function call. The CPU changes privilege level, the kernel switches to a kernel stack, saves and later restores the user register state, validates every argument, and on CPUs carrying speculative-execution mitigations it also swaps page tables and runs branch-predictor barriers on entry and exit. That is fixed overhead per call, independent of how much work the call actually does, so a workload doing thousands of tiny reads and writes pays it thousands of times. The remedy is amortisation, not micro-optimisation of the call itself: buffer in user space so one `write()` carries many messages; use `writev()`/`readv()` to hand the kernel several buffers at once; use `sendfile()` or `splice()` so file-to-socket data never round-trips through user memory; and let a readiness API return many events per call rather than polling descriptors one at a time.

code

c · 18 lines
c
#include <string.h>
#include <sys/uio.h>
#include <unistd.h>

int main(void) {
    char *status  = "HTTP/1.1 200 OK\r\n";
    char *headers = "Content-Length: 2\r\n\r\n";
    char *body    = "ok";

    struct iovec iov[3] = {
        { status,  strlen(status)  },
        { headers, strlen(headers) },
        { body,    strlen(body)    },
    };

    writev(STDOUT_FILENO, iov, 3);   /* one crossing instead of three */
    return 0;
}

go deeper

for a junior

Know that a system call is much more expensive than a normal function call and that buffering output instead of writing a line at a time is the standard reason programs go faster.

for a middle

Explain what the transition actually does — privilege change, register save and restore, argument validation, page-table work on mitigated CPUs — and why that overhead is fixed regardless of payload size.

for a senior

Diagnose it: separate transition overhead from work done inside the kernel, distinguish a mode switch from a context switch, and pick the right amortisation — vectored I/O, zero-copy transfer, batched readiness, or submission queues.

for a principal

Own the tradeoff. Batching buys throughput at the cost of latency and complexity, and newer submission-queue interfaces raise the kernel-version floor and the security surface; be ready to say when the syscall rate is genuinely the constraint worth redesigning around.

## What the crossing actually costs The expensive part of a system call is rarely the work it names; it is the transition. On entry the CPU switches privilege level, the kernel switches to the thread's kernel stack, and the user-mode register state is saved so it can be restored exactly. The syscall number is bounds-checked and dispatched. Every pointer argument is treated as untrusted and copied through validating helpers, because the kernel must never dereference a user address directly. On exit the whole thing unwinds, plus a check for pending work such as signals or a reschedule request. Speculative-execution mitigations made this dearer. With kernel page-table isolation, entering the kernel switches to a different set of page tables and leaving switches back, which invalidates translation-buffer entries and costs real time; branch-predictor barriers add more. The upshot is that on a mitigated x86-64 machine a trivial system call costs on the order of hundreds of nanoseconds where an in-process function call costs single-digit nanoseconds. That ratio is the whole reason syscall-bound workloads exist. None of this is wasted on a call that transfers a megabyte. It is entirely wasted on a call that transfers sixteen bytes. ## Mode switch is not a context switch The two get conflated constantly, and the distinction matters for diagnosis. A **mode switch** changes privilege level within the same thread: same address space (mitigations aside), same scheduling entity, same caches mostly warm. A **context switch** replaces the running thread with a different one: the scheduler runs, register and floating-point state is swapped, and if the new thread belongs to another process the address space changes, so translation buffers and caches lose their working set. A context switch is typically an order of magnitude more disruptive than a mode switch, and its real cost is the cache damage afterwards rather than the switch itself. A system call causes a context switch only if it blocks — waiting on I/O, on a lock, on an empty socket. A non-blocking call that is satisfied immediately, such as a `read()` served from the page cache, is a mode switch and nothing more. So "high kernel-mode CPU time" and "high context-switch rate" are different symptoms with different remedies: the first says you are calling too often or doing too much per call, the second says your threads keep going to sleep. ## Making the same work cost fewer crossings Every technique here is a form of amortisation — the same bytes, fewer transitions. **Buffer in user space.** The classic case is writing a log line or a protocol message at a time. Accumulating output and flushing when the buffer fills turns thousands of `write()` calls into a handful. This is exactly what stdio does, and it is why a program that writes with `printf()` performs so differently from one that writes each line with an unbuffered `write()`. **Vectored I/O.** When the pieces cannot be concatenated cheaply — a status line, headers and a body held in separate allocations — `writev()` takes an array of buffers and writes them in one call. `readv()` is the mirror image. One transition, several buffers, no memory copying to join them. **Zero-copy transfers.** Serving a file over a socket the naive way costs a `read()` into user memory and a `write()` out of it: two crossings and two copies of every byte. `sendfile()` moves the data inside the kernel with one call and no user-space copy; `splice()` generalises that to pipes. For a static-content server this is the difference between being CPU-bound and being network-bound. **Batch the event loop.** A readiness API that returns many ready descriptors from a single call amortises the crossing across all of them. `epoll_wait()` fills an array of events per call; the pathological alternative is asking about descriptors individually in a loop. Combine that with reading until the socket is drained rather than one small chunk per wakeup. **Submission-queue I/O.** `io_uring` takes the idea further: the application and the kernel share ring buffers in memory, so many operations can be submitted and completed with few or no transitions at all. It is the right tool when you have already exhausted batching and the syscall rate is still the wall you are hitting. **Do not call at all when you need not.** Reading the clock is the common offender, and it is already handled by the kernel mapping timekeeping data into every process. But application-level equivalents exist: caching a value your code re-derives from the kernel on every request often removes an entire class of calls. ## Where this reasoning stops Syscall overhead is a real ceiling, but it is not the explanation for every slow service. If the calls are blocking on disk or network, the transitions are noise next to the wait, and batching buys nothing; the fix is concurrency or fewer round trips. The question to answer first is whether kernel time is being spent *entering* the kernel or *inside* it doing genuine work — high kernel CPU with a huge call count and small payloads points at the boundary, while high kernel CPU with a modest call count points at the work itself, such as checksum, copy or lock contention in the network stack.

  • How do you tell a syscall-overhead problem from a blocking-I/O problem?
    Look at what the kernel time is buying. Overhead shows up as a very high call rate with tiny payloads and CPU that stays busy — the machine is working, just on transitions. Blocking shows up as threads spending their time off-CPU waiting, with a high context-switch or sleep rate rather than high kernel CPU. Batching fixes the first and does nothing for the second, where the answers are concurrency or fewer round trips.
  • Why does sendfile() help more than simply using a larger read buffer?
    A larger buffer removes crossings but still copies every byte into user memory and back out. `sendfile()` removes both: one system call, and the data moves inside the kernel from the page cache to the socket without ever being copied into the process's address space. On a static-content server that is the difference between CPU-bound and network-bound, because you stop paying twice for memory bandwidth you never needed.
  • Does io_uring eliminate system calls entirely?
    Not entirely, but it can come close. The application and kernel share submission and completion ring buffers in memory, so many operations are queued and reaped by writing and reading memory. A call is still needed to tell the kernel there is work, though even that can be avoided in polled modes. The real win is that the number of transitions stops scaling with the number of operations.

saying these in an interview costs you the question

  • Calls a system call just a slower function call
  • Treats mode switch and context switch as the same thing
  • Thinks more threads reduces syscall overhead
  • Assumes any high kernel CPU means the syscall boundary
  • Believes a bigger buffer removes the data copy

context