Describe what happens between a C program calling read() and the kernel returning data on Linux x86-64, and explain how a failed system call is reported back to the caller.
answer
- the libc wrapper is not the call
- number and arguments go in registers
- kernel returns one small negative value
- wrapper turns that into -1 and errno
basics
~20 sThe libc wrapper loads the syscall number and arguments into registers, executes the syscall trap, and the kernel dispatches, validates and returns a value in the same register. A negative value means failure, so the wrapper returns -1 and stores the positive error code in the thread-local errno.
solid answer
~50 s`read()` in your program is a glibc wrapper, not the system call itself. The wrapper puts the syscall number in `rax` and the three arguments — file descriptor, buffer pointer, count — in `rdi`, `rsi` and `rdx`, then executes the `syscall` instruction. The CPU switches to kernel mode at a fixed entry point; the kernel saves the user registers, dispatches through the system call table, and runs the handler, which resolves the descriptor, copies bytes into the user buffer with pointer-checking helpers, and returns. The return value comes back in `rax`. The kernel's own error convention is a small negative number such as `-EBADF`; the wrapper sees a value in that negative range, converts it to a positive code stored in the thread-local `errno`, and returns `-1` to you. That is why you check the return value first and only then read `errno`.
code
c · 17 lines#include <errno.h>
#include <stdio.h>
#include <string.h>
#include <unistd.h>
int main(void) {
char buf[16];
ssize_t n = read(-1, buf, sizeof buf); /* deliberately invalid fd */
if (n == -1) {
int saved = errno; /* save it before anything else runs */
printf("read failed: %s\n", strerror(saved));
return 1;
}
printf("read %zd bytes\n", n);
return 0;
}go deeper
Know the calling convention you actually write: a system call wrapper returns -1 on failure and sets errno, and you check the return value first. Naming a couple of common errno values such as EBADF or EINTR helps.
Walk the whole path out loud: wrapper, syscall number and arguments in registers, trap, table dispatch, argument validation in the kernel, and the single return register carrying either a result or a negative error code.
Show the operational consequences: that short reads and EINTR must be handled, that errno is thread-local and easily clobbered, and that the count of library calls in your code is not the count of system calls on the machine.
Be ready to discuss the syscall ABI as a stability contract — numbers and semantics the kernel cannot break — and what that implies for shipping software that bypasses libc, such as static binaries and language runtimes with their own syscall layer.
## The wrapper is not the system call When your C code calls `read()`, the symbol you bind to lives in the C library — glibc on most distributions, musl on Alpine. It is ordinary user-space code. Its whole job is to marshal arguments into the registers the kernel expects, execute the trap, and translate the result into the convention C programmers expect. The system call is only the trap instruction in the middle of that wrapper. Keeping the two ideas apart explains a lot of otherwise confusing behaviour: why the same C function can be a real trap on one call and pure user-space work on another, why the C library can retry or transform calls, and why the number of `read()` calls in your source is not necessarily the number of system calls the machine performs. ## Registers and the calling convention The system call ABI is deliberately different from the ordinary C function ABI. On x86-64 Linux: - `rax` carries the system call number and, on return, the result; - arguments go in `rdi`, `rsi`, `rdx`, `r10`, `r8`, `r9` — note `r10`, not `rcx`, because the `syscall` instruction clobbers `rcx` and `r11` to save the return address and flags; - at most six arguments fit, which is why calls needing more take a pointer to a struct. The numbers themselves are architecture-specific and stable within an architecture, because they are the kernel's binary contract with userspace: number 0 is `read` and 1 is `write` on x86-64, but the numbering differs on 32-bit x86 and on ARM. That is why a syscall number is meaningless without the architecture that goes with it. ## Inside the kernel The trap lands at an entry point the kernel installed at boot. The kernel switches to the calling thread's kernel stack, saves the user-mode registers, and range-checks the number in `rax` against the system call table. An out-of-range number is not a crash; it returns `-ENOSYS`. The handler then does its work under the assumption that every argument is untrusted. The descriptor number is looked up in the process's open-file table, which is where `EBADF` comes from. The user buffer pointer is never dereferenced directly; the kernel copies through helpers that validate the address is user-space memory and recover from a fault, which is where `EFAULT` comes from. Only after those checks does the actual read happen, which may serve from the page cache immediately or block the thread while I/O completes. ## How errors come back The kernel has exactly one return channel: the value in `rax`. It signals errors by returning a small negative number — the negated `errno` constant, such as `-9` for `EBADF`. Successful results are non-negative: a byte count, a descriptor, an offset. The wrapper checks whether the return value falls in the reserved negative error range; if it does, it negates it, stores it in `errno`, and returns `-1`. This is why the two-step discipline exists. `errno` is meaningful only after a call has reported failure through its return value. The kernel and the library never reset `errno` to zero on success, so a stale value from an earlier failure sits there indefinitely. Reading `errno` after a call that succeeded tells you about some unrelated earlier call. ```c ssize_t n = read(fd, buf, len); if (n == -1) { /* only now is errno meaningful */ if (errno == EINTR) { /* interrupted by a signal, retry */ } } ``` A related trap: anything you do between the failing call and reading `errno` can overwrite it, including `printf()`, a logging call, or a destructor. Save it into a local `int` immediately if you are not going to use it on the very next line. ## errno is per-thread In a threaded program, `errno` cannot be a single global — two threads failing concurrently would clobber each other. In glibc `errno` is a macro that expands to a dereference of a thread-local location, so each thread has its own. Practical consequences: you may not declare `extern int errno` yourself, you must include `<errno.h>`, and you cannot pass `&errno` around expecting it to mean the same thing in another thread. ## Raw system calls, and why you rarely want them glibc exposes a generic `syscall()` function, declared in `<unistd.h>`, taking a number from `<sys/syscall.h>` such as `SYS_read`. It is useful for calls glibc has no wrapper for. But it is still a library wrapper: it also returns `-1` and sets `errno`. The raw negative-return convention is visible only if you write the trap in assembly yourself. Bypassing the wrapper also bypasses whatever the library does around the call — cancellation handling, the transparent substitution of a newer, better system call for an older one, and the vDSO fast paths for a few time-related calls. Unless you are implementing a runtime or a sandbox, call the wrapper. ## Return semantics worth knowing `read()` returning a positive number smaller than the requested count is not an error; short reads are normal on pipes, sockets and terminals, and code that assumes a full read is a latent bug. Zero means end of file. And a blocking call interrupted by a signal returns `-1` with `EINTR`, which application code must be prepared to retry.
- Why must you check the return value before reading errno rather than the other way round?Because nothing ever clears `errno` on success. The kernel and the C library only ever write to it when a call fails, so a stale value from a much earlier failure persists indefinitely. Testing `errno != 0` after a successful call will therefore report a phantom error. The rule is: the return value tells you whether it failed, and `errno` tells you why — in that order, and with `errno` read before any other call can overwrite it.
- How does errno behave in a multi-threaded program?It is per-thread. In glibc `errno` is a macro that resolves to thread-local storage, so two threads failing at the same time do not overwrite each other's error. That is also why you must include `<errno.h>` instead of declaring `extern int errno`, and why you cannot hand `&errno` to another thread and expect it to refer to that thread's error state.
- What does a read() call returning fewer bytes than requested mean?It is a normal, successful short read, not an error. On pipes, sockets and terminals the kernel returns whatever is available rather than waiting for the full count, so code that treats a short return as failure — or that assumes the buffer is full — is buggy. A return of 0 means end of file, and only -1 signals an error. Correct code loops until it has what it needs.
saying these in an interview costs you the question
- Says read() in C is itself the system call
- Reads errno without first checking the return value
- Assumes errno is reset to zero on success
- Thinks errno is a global shared by all threads
- Treats a short read as an error condition