skip to content

questions

4

What is the difference between a process and a thread, and what does it actually mean to say that each process has its own address space?

level: juniorimportance: must knowfreq 85%

answer

  1. process = resource ownership; thread = scheduling
  2. page tables + MMU enforce isolation
  3. threads share heap/handles, not stack/registers
  4. pointer is meaningless outside its process
  5. crash blast radius: one process vs all threads

basics

~20 s

A process owns an isolated address space and OS resources; threads are execution contexts inside one process that share that memory. Each thread has its own stack, registers and program counter. Separate processes cannot touch each other's memory and must use explicit inter-process communication.

solid answer

~50 s

A **process** is the operating system's unit of *resource ownership*: it owns a virtual address space, open file handles, credentials and accounting state. A **thread** is the unit of *scheduling*: a program counter, register set and stack that a CPU can run. One process can hold many threads. Those threads share heap, globals and code, so handing data between them costs a pointer. Each thread still keeps its own stack and registers, because that is its private execution position. "Own address space" is a hardware fact, not a convention. The CPU's memory-management unit translates virtual addresses through per-process page tables, so address 0x4000 in process A and in process B refer to different physical memory, and touching an unmapped address faults instead of reading a neighbour's data. That isolation buys fault containment and a security boundary. It costs cheap sharing: cross-process communication must go through pipes, sockets, signals or explicitly mapped shared memory, and usually involves copying.

code

text · 5 lines
text
Process A:  addr 0x4000 -> page table A -> physical frame 17
Process B:  addr 0x4000 -> page table B -> physical frame 92

Thread 1 and Thread 2 inside Process A:
  both:     addr 0x4000 -> page table A -> physical frame 17   (same bytes)

go deeper

for a junior

Recall the core split: process owns memory and resources, thread is what gets scheduled, threads inside a process share the heap but not stacks.

for a middle

Explain the mechanism — page tables and the MMU — and derive the consequences: sharing needs IPC, crashes are contained, switching costs differ.

for a senior

Frame it as a design boundary: what blast radius and trust boundary you get for what communication cost, and where hybrid designs sit.

for a principal

Discuss isolation as an architectural primitive — fault domains, privilege reduction, per-tenant separation — and the throughput price of copying versus the operability gain.

## Two different units Operating systems separate two ideas that are easy to merge. - A **process** is the unit of **resource ownership**. It owns a virtual address space, a table of open file and socket handles, a working directory, user credentials, resource limits and accounting information. - A **thread** is the unit of **scheduling**. It is a program counter, a register set, a stack and a small amount of thread-private storage. The scheduler picks threads, not processes, when it decides what runs on a core. A process always contains at least one thread. A single-threaded process is simply a process whose one thread does all the work. ## What threads share and what they do not Threads inside a process share: the code, the heap, global and static data, open file handles, and the memory mapping itself. They do **not** share: the call stack, the register state (including the program counter), and any thread-local storage slots. That asymmetry explains the classic bug pattern. Two threads reading and writing the same heap object need synchronization, because they truly touch the same bytes. Two threads each using their own local variables never conflict, because those live on separate stacks. ## What address-space isolation really is Each process has its own **page tables**: a mapping from virtual addresses (the numbers the program uses) to physical frames of RAM. The memory-management unit consults those tables on every access. A page not mapped for the current process cannot be reached at all — the hardware raises a fault and the OS typically kills the offender. The consequences are worth stating plainly: 1. **A pointer is only meaningful inside its own process.** Sending the numeric value of a pointer to another process is useless; the same number means something else there. 2. **Isolation is enforced, not agreed.** A buggy or malicious process cannot scribble on another's data structures the way a buggy thread can scribble on its siblings'. 3. **Sharing must be requested.** Two processes only share memory when both explicitly map the same object (shared-memory segments or memory-mapped files). Otherwise every byte exchanged is copied through the kernel. ## What that trade buys **Fault containment.** A crash — a bad pointer, an out-of-memory kill, a stack overflow — takes down one process. Its siblings keep running, and a supervisor can restart it. In a multi-threaded process the same fault usually ends the entire process, taking every thread with it. **A security boundary.** Processes carry credentials and can be sandboxed, given reduced privileges, or confined so that even total compromise of one yields little. Threads inside a process all run with the same authority; there is no meaningful privilege boundary between them. **Freedom from shared-mutable-state hazards.** Two processes that communicate only by messages cannot race on a shared object, because there is no shared object. The hazards move elsewhere (protocol design, message ordering, partial failure), but the classic data race disappears by construction. ## What it costs **Creation and footprint.** A process needs its own page tables, handle table and kernel bookkeeping; a thread needs a stack and a small control block. Processes are heavier by an order of magnitude or more. **Communication.** Sending a megabyte between threads is a pointer assignment. Between processes it is a copy — often two, into and out of a kernel buffer — plus serialization if the data is not a flat byte blob. **Switching.** Switching between threads of the same process keeps the address space; switching between processes changes page tables and can invalidate address-translation caches, so it is more expensive. ## How to choose Use threads when the units of work must share large mutable state cheaply and trust each other. Use processes when you need a failure or trust boundary: untrusted code, plugins, per-tenant isolation, or a component whose crash must not take the whole system down. Many mature systems do both — a small number of isolated processes, each internally multi-threaded.

  • If threads share the heap, why does each thread still need its own stack?
    The stack records where a thread currently is: its call chain, return addresses, and local variables. Two threads execute different call chains at the same moment, so a shared stack would be meaningless and instantly corrupted. Locals living on private stacks is also why they need no synchronization.
  • Can two processes ever share memory directly, and what changes when they do?
    Yes — both can map the same shared-memory object or file into their address spaces, after which reads and writes hit the same physical pages with no copying. Once they do, they are back in shared-mutable-state territory: they need cross-process synchronization and the isolation guarantee no longer covers that region.
  • Which is cheaper to switch between, two threads of one process or two processes, and why?
    Two threads of the same process, because the address space stays the same: the kernel swaps registers and stacks but keeps the page tables and the address-translation caches valid. A process switch changes page tables, which can invalidate translation caches and force expensive re-population.

A process is an apartment with its own locked door; threads are flatmates inside it. Flatmates share the fridge and can spoil each other's food; neighbours have to knock and hand things through the door.

saying these in an interview costs you the question

  • Saying threads have their own memory, or that each thread gets its own heap.
  • Claiming isolation is enforced by the language or runtime rather than by hardware page tables.
  • Believing one process can read another's variables by address without shared memory being set up.
  • Saying a thread crash only kills that thread — an unhandled fault normally kills the whole process.
  • Treating 'process' and 'program' as the same thing; one program can run as many processes.

context

open as a page

Explain the fork-then-exec model of creating a new process, including what copy-on-write does, and why creating a process is normally more expensive than creating a thread.

level: middleimportance: should knowfreq 50%

basics

~20 s

Fork clones the calling process, giving the child a duplicate address space and handle table; exec then replaces the child's program image with a new one. Copy-on-write makes the duplicate lazy — pages are shared read-only until written. Processes still cost more than threads: page tables, handle tables and kernel bookkeeping versus one stack.

open as a page

Compare the main inter-process communication mechanisms — pipes, shared memory, sockets and signals — and explain how you would choose among them for moving a high volume of data between two processes on the same machine.

level: seniorimportance: should knowfreq 45%

basics

~20 s

Pipes are simple one-way byte streams via kernel buffers. Sockets are the same idea but routable across machines. Shared memory is the only zero-copy option — mapped by both sides, needing its own synchronization. Signals carry notification, not data. For high volume on one host, use shared memory with a ring buffer, or pipes if simplicity wins.

open as a page

You are designing a server that runs untrusted extension code and must stay available when one unit of work crashes. Make the case for a multi-process architecture versus a multi-threaded one, and state honestly what each choice costs.

level: principalimportance: should knowfreq 40%

basics

~20 s

Multi-process gives hardware-enforced fault and trust boundaries: a crash, leak or compromise is contained and the supervisor restarts one worker. It costs memory per worker, copying on every message, and slower startup. Multi-threaded is cheaper and shares state for free, but one bad pointer or hostile extension takes down everything.

open as a page