On Linux, what is the difference between user space and kernel space, and by what mechanism does a user-space process get the kernel to do work on its behalf?
answer
- two CPU privilege levels
- hardware enforces the split
- one controlled entry point
- a trap, not a function call
basics
~20 sUser space is the unprivileged CPU mode, where a process sees only its own virtual memory and cannot touch hardware; kernel space is the privileged mode. A process crosses over only by issuing a system call, a hardware trap that enters the kernel at one fixed, kernel-chosen entry point.
solid answer
~50 sModern CPUs run code at two privilege levels, and Linux uses both. User space is the unprivileged one: a process there can address only its own virtual memory, cannot execute privileged instructions, and cannot talk to devices. Kernel space is privileged — the kernel can map all physical memory and drive hardware. The hardware enforces the split, so a process cannot simply jump into kernel code; it has to ask. That ask is a **system call**: the process places a syscall number and arguments in agreed registers and executes a trap instruction (`syscall` on x86-64), which switches the CPU into kernel mode and transfers control to a single kernel entry point. The kernel validates the number and the arguments, does the work, and returns to user mode with a result. Everything that leaves a program's own address space — opening a file, sending a packet, creating a process — goes through that one door.
go deeper
Be ready to state plainly that user mode is unprivileged, kernel mode is privileged, and a system call is the only way across. Naming one or two real system calls, such as read or write, is enough at this level.
Explain the mechanics: the trap instruction, the syscall number and arguments in registers, the fixed kernel entry point, and the fact that the C library wrapper is user-space code that merely ends in that trap.
Show you reason about the boundary in production: which library calls actually enter the kernel, why kernel-mode CPU time is a distinct signal from user time, and why the kernel must validate every argument it receives.
Own the isolation argument. Be prepared to compare this hardware boundary with the weaker or stronger boundaries around it, and to say what a kernel bug costs you relative to an application bug when you choose a workload isolation strategy.
## Two privilege levels, one CPU The user/kernel split is not a software convention that programs politely observe — it is a property of the processor. On x86-64 the CPU has privilege rings, and Linux uses exactly two of them: ring 3 for user code and ring 0 for the kernel. On 64-bit ARM the equivalent is exception levels, with applications at EL0 and the kernel at EL1. In either case the current privilege level is CPU state, and instructions that configure memory mapping, mask interrupts, or touch I/O ports simply fault when attempted at the unprivileged level. The second half of the enforcement is the page tables. Every virtual address a process uses is translated through page tables the kernel controls, and kernel memory is mapped with a flag that makes it inaccessible from user mode. So a user process cannot read kernel data structures even by guessing an address: the memory management unit refuses the translation and raises a fault, which Linux normally turns into a `SIGSEGV`. ## What the split buys you Because the boundary is enforced by hardware, the kernel can be the sole arbiter of everything shared: physical memory, the CPU schedule, the filesystem, the network stack, devices. A buggy or hostile process cannot corrupt another process's memory, cannot bypass file permissions by writing to disk directly, and cannot monopolise a device. It also means a crashing application is just a process that dies, while a bug in kernel code can take the machine down — which is why so little runs in ring 0 and why drivers are the classic source of kernel panics. A point candidates often miss: *root is not kernel mode*. A process running as UID 0 is still ring 3 user code. It gets far more permissive answers when it asks the kernel for things, but it is still asking. The only code that runs in kernel mode is the kernel itself and the modules loaded into it — which is exactly why loading a kernel module is such a privileged, security-sensitive operation. ## Crossing the boundary The transition is a deliberate, controlled trap. On x86-64 the process puts the system call number in `rax` and up to six arguments in `rdi`, `rsi`, `rdx`, `r10`, `r8` and `r9`, then executes the `syscall` instruction. The CPU switches to kernel mode and jumps to an address the kernel installed at boot — the process does not choose where it lands, which is the whole point. The kernel switches to that thread's kernel stack, saves the user register state, looks the number up in the system call table, and dispatches. Inside, the kernel treats every argument as hostile. Pointers coming from user space are copied in and out with dedicated helpers that check the address is really user-space memory and handle a fault gracefully; that is why a bad pointer produces an `EFAULT` return rather than a kernel crash. When the work is done, the kernel restores the user registers, puts the result in `rax`, and returns to ring 3 at the instruction after the trap. ```c /* what glibc's read() wrapper does on x86-64, in essence */ mov $0, %eax /* syscall number 0 = read */ syscall /* trap: CPU enters kernel mode here */ /* execution resumes here, in user mode, with the result in %rax */ ``` ## Where the C library sits Almost no application writes that trap by hand. Instead it calls a C library function — `read()`, `open()`, `connect()` — and glibc (or musl on Alpine) supplies a thin wrapper that loads the registers, traps, and translates the kernel's return value into the familiar "return -1 and set `errno`" convention. Because the wrapper and the system call usually share a name, people conflate them, but they are different things: the wrapper is ordinary user-space code that happens to end in a trap instruction. That distinction matters because many library functions are *not* system calls at all. `printf()` formats a string in user space and may buffer it for a long time before it ever calls `write()`. `malloc()` normally hands back memory the allocator already owns and only occasionally asks the kernel for more with `brk()` or `mmap()`. `strlen()` never enters the kernel. Arithmetic, function calls within your program, and reads of your own heap are all pure user-space work — the boundary is crossed only when a program needs something it cannot do itself. ## The practical consequence Since every interaction with the outside world is a system call, the list of system calls a process makes is a complete description of its effect on the machine. That is why the boundary is the natural place to observe a program, to sandbox it, and to reason about its cost: it is the one interface a process cannot route around.
- Is a process running as root running in kernel mode?No. A root process is still ordinary user-mode code in ring 3; it makes the same system calls as anyone else, it just passes more of the kernel's permission checks. The only code executing in kernel mode is the kernel and its loaded modules. That is why loading a kernel module is treated as a far more dangerous privilege than merely being UID 0 — it puts your code inside the boundary rather than in front of it.
- Does every C library function call cross into the kernel?No, and assuming so leads to badly wrong performance models. `strlen()` and `memcpy()` never leave user space. `printf()` formats in user space and buffers output, so many `printf()` calls can collapse into one `write()`. `malloc()` serves most requests from memory the allocator already holds, asking the kernel via `brk()` or `mmap()` only when it needs more. The kernel is entered only for things a process cannot do itself.
- What happens if a program passes an invalid pointer to a system call?The kernel never dereferences user pointers directly; it uses copy helpers that validate the address range and recover from a fault. So the call returns `-1` with `errno` set to `EFAULT` rather than crashing the machine. This defensive copying is a standing rule for kernel code, because a system call argument is by definition attacker-controlled input.
Think of the kernel as a bank vault and your process as a customer in the lobby. You cannot walk into the vault; you slide a numbered request slip through one window, and a teller who works for the bank decides whether to honour it and hands the result back.
saying these in an interview costs you the question
- Says kernel space is just memory a process may read
- Thinks running as root means running in kernel mode
- Believes every libc function call enters the kernel
- Claims the process picks which kernel address it jumps to
- Confuses a system call with a normal function call