skip to content

Describe what happens between a C program calling read() and the kernel returning data on Linux x86-64, and explain how a failed system call is reported back to the caller.

level: middleimportance: must knowfreq 60%

answer

  1. the libc wrapper is not the call
  2. number and arguments go in registers
  3. kernel returns one small negative value
  4. wrapper turns that into -1 and errno

basics

~20 s

The libc wrapper loads the syscall number and arguments into registers, executes the syscall trap, and the kernel dispatches, validates and returns a value in the same register. A negative value means failure, so the wrapper returns -1 and stores the positive error code in the thread-local errno.

solid answer

~50 s

`read()` in your program is a glibc wrapper, not the system call itself. The wrapper puts the syscall number in `rax` and the three arguments — file descriptor, buffer pointer, count — in `rdi`, `rsi` and `rdx`, then executes the `syscall` instruction. The CPU switches to kernel mode at a fixed entry point; the kernel saves the user registers, dispatches through the system call table, and runs the handler, which resolves the descriptor, copies bytes into the user buffer with pointer-checking helpers, and returns. The return value comes back in `rax`. The kernel's own error convention is a small negative number such as `-EBADF`; the wrapper sees a value in that negative range, converts it to a positive code stored in the thread-local `errno`, and returns `-1` to you. That is why you check the return value first and only then read `errno`.

code

c · 17 lines
c
#include <errno.h>
#include <stdio.h>
#include <string.h>
#include <unistd.h>

int main(void) {
    char buf[16];
    ssize_t n = read(-1, buf, sizeof buf);   /* deliberately invalid fd */

    if (n == -1) {
        int saved = errno;                   /* save it before anything else runs */
        printf("read failed: %s\n", strerror(saved));
        return 1;
    }
    printf("read %zd bytes\n", n);
    return 0;
}

go deeper

for a junior

Know the calling convention you actually write: a system call wrapper returns -1 on failure and sets errno, and you check the return value first. Naming a couple of common errno values such as EBADF or EINTR helps.

for a middle

Walk the whole path out loud: wrapper, syscall number and arguments in registers, trap, table dispatch, argument validation in the kernel, and the single return register carrying either a result or a negative error code.

for a senior

Show the operational consequences: that short reads and EINTR must be handled, that errno is thread-local and easily clobbered, and that the count of library calls in your code is not the count of system calls on the machine.

for a principal

Be ready to discuss the syscall ABI as a stability contract — numbers and semantics the kernel cannot break — and what that implies for shipping software that bypasses libc, such as static binaries and language runtimes with their own syscall layer.

## The wrapper is not the system call When your C code calls `read()`, the symbol you bind to lives in the C library — glibc on most distributions, musl on Alpine. It is ordinary user-space code. Its whole job is to marshal arguments into the registers the kernel expects, execute the trap, and translate the result into the convention C programmers expect. The system call is only the trap instruction in the middle of that wrapper. Keeping the two ideas apart explains a lot of otherwise confusing behaviour: why the same C function can be a real trap on one call and pure user-space work on another, why the C library can retry or transform calls, and why the number of `read()` calls in your source is not necessarily the number of system calls the machine performs. ## Registers and the calling convention The system call ABI is deliberately different from the ordinary C function ABI. On x86-64 Linux: - `rax` carries the system call number and, on return, the result; - arguments go in `rdi`, `rsi`, `rdx`, `r10`, `r8`, `r9` — note `r10`, not `rcx`, because the `syscall` instruction clobbers `rcx` and `r11` to save the return address and flags; - at most six arguments fit, which is why calls needing more take a pointer to a struct. The numbers themselves are architecture-specific and stable within an architecture, because they are the kernel's binary contract with userspace: number 0 is `read` and 1 is `write` on x86-64, but the numbering differs on 32-bit x86 and on ARM. That is why a syscall number is meaningless without the architecture that goes with it. ## Inside the kernel The trap lands at an entry point the kernel installed at boot. The kernel switches to the calling thread's kernel stack, saves the user-mode registers, and range-checks the number in `rax` against the system call table. An out-of-range number is not a crash; it returns `-ENOSYS`. The handler then does its work under the assumption that every argument is untrusted. The descriptor number is looked up in the process's open-file table, which is where `EBADF` comes from. The user buffer pointer is never dereferenced directly; the kernel copies through helpers that validate the address is user-space memory and recover from a fault, which is where `EFAULT` comes from. Only after those checks does the actual read happen, which may serve from the page cache immediately or block the thread while I/O completes. ## How errors come back The kernel has exactly one return channel: the value in `rax`. It signals errors by returning a small negative number — the negated `errno` constant, such as `-9` for `EBADF`. Successful results are non-negative: a byte count, a descriptor, an offset. The wrapper checks whether the return value falls in the reserved negative error range; if it does, it negates it, stores it in `errno`, and returns `-1`. This is why the two-step discipline exists. `errno` is meaningful only after a call has reported failure through its return value. The kernel and the library never reset `errno` to zero on success, so a stale value from an earlier failure sits there indefinitely. Reading `errno` after a call that succeeded tells you about some unrelated earlier call. ```c ssize_t n = read(fd, buf, len); if (n == -1) { /* only now is errno meaningful */ if (errno == EINTR) { /* interrupted by a signal, retry */ } } ``` A related trap: anything you do between the failing call and reading `errno` can overwrite it, including `printf()`, a logging call, or a destructor. Save it into a local `int` immediately if you are not going to use it on the very next line. ## errno is per-thread In a threaded program, `errno` cannot be a single global — two threads failing concurrently would clobber each other. In glibc `errno` is a macro that expands to a dereference of a thread-local location, so each thread has its own. Practical consequences: you may not declare `extern int errno` yourself, you must include `<errno.h>`, and you cannot pass `&errno` around expecting it to mean the same thing in another thread. ## Raw system calls, and why you rarely want them glibc exposes a generic `syscall()` function, declared in `<unistd.h>`, taking a number from `<sys/syscall.h>` such as `SYS_read`. It is useful for calls glibc has no wrapper for. But it is still a library wrapper: it also returns `-1` and sets `errno`. The raw negative-return convention is visible only if you write the trap in assembly yourself. Bypassing the wrapper also bypasses whatever the library does around the call — cancellation handling, the transparent substitution of a newer, better system call for an older one, and the vDSO fast paths for a few time-related calls. Unless you are implementing a runtime or a sandbox, call the wrapper. ## Return semantics worth knowing `read()` returning a positive number smaller than the requested count is not an error; short reads are normal on pipes, sockets and terminals, and code that assumes a full read is a latent bug. Zero means end of file. And a blocking call interrupted by a signal returns `-1` with `EINTR`, which application code must be prepared to retry.

  • Why must you check the return value before reading errno rather than the other way round?
    Because nothing ever clears `errno` on success. The kernel and the C library only ever write to it when a call fails, so a stale value from a much earlier failure persists indefinitely. Testing `errno != 0` after a successful call will therefore report a phantom error. The rule is: the return value tells you whether it failed, and `errno` tells you why — in that order, and with `errno` read before any other call can overwrite it.
  • How does errno behave in a multi-threaded program?
    It is per-thread. In glibc `errno` is a macro that resolves to thread-local storage, so two threads failing at the same time do not overwrite each other's error. That is also why you must include `<errno.h>` instead of declaring `extern int errno`, and why you cannot hand `&errno` to another thread and expect it to refer to that thread's error state.
  • What does a read() call returning fewer bytes than requested mean?
    It is a normal, successful short read, not an error. On pipes, sockets and terminals the kernel returns whatever is available rather than waiting for the full count, so code that treats a short return as failure — or that assumes the buffer is full — is buggy. A return of 0 means end of file, and only -1 signals an error. Correct code loops until it has what it needs.

saying these in an interview costs you the question

  • Says read() in C is itself the system call
  • Reads errno without first checking the return value
  • Assumes errno is reset to zero on success
  • Thinks errno is a global shared by all threads
  • Treats a short read as an error condition

context