You are designing a server that runs untrusted extension code and must stay available when one unit of work crashes. Make the case for a multi-process architecture versus a multi-threaded one, and state honestly what each choice costs.
answer
- choose a fault domain, not a primitive
- untrusted code -> process, no debate
- isolation bill: memory, copies, supervisor
- hybrid: few processes, threads inside
- recycle workers, backoff restarts
basics
~20 sMulti-process gives hardware-enforced fault and trust boundaries: a crash, leak or compromise is contained and the supervisor restarts one worker. It costs memory per worker, copying on every message, and slower startup. Multi-threaded is cheaper and shares state for free, but one bad pointer or hostile extension takes down everything.
solid answer
~50 sWith untrusted code the decision is nearly made for you. Threads share one address space and one set of privileges, so an extension that misbehaves can read every secret in the process and a single fault kills all in-flight work. Processes give a **hardware-enforced** boundary: each worker can be sandboxed, privilege-reduced and resource-capped, and its crash costs one request while a supervisor restarts it. What that costs: per-worker memory and startup, message passing instead of pointer passing (serialize, copy, deserialize), harder sharing of caches, and a distributed-systems flavour inside one machine — partial failure, restart storms, versioning between workers and supervisor. The multi-threaded design wins on raw efficiency: cheap creation, a shared cache, no serialization. It is the right answer when all code is trusted and correctness is under your control. Most mature systems land in the middle: a small number of isolated processes, each internally multi-threaded, with an explicit message contract between them and a per-worker crash budget.
code
text · 15 linespool = spawn_workers(N) # pre-warmed, sandboxed, privilege-reduced
loop:
unit = accept()
w = pool.pick_idle()
send(w, unit) # message, not pointer
on worker_exit(w, reason):
fail_only(w.current_unit) # one unit fails, not the server
record(reason)
if crash_rate > budget: alert()
respawn(w, backoff) # avoid restart storms
on worker_units_handled > K or worker_rss > limit:
drain(w); respawn(w) # recycle to bound leaksgo deeper
Say the core fact: separate processes contain crashes and untrusted code, threads share everything so one failure takes the process down.
Add the concrete costs — per-worker memory, serialization instead of pointer passing, slower startup — and mention pre-warmed worker pools.
Reason about blast radius, worker recycling to bound leaks, crash budgets, and how a worker crash surfaces as a clean per-unit error.
Present it as fault-domain and trust-boundary design: where the seams go, what isolation converts a class of bugs into, the supervision and rollout machinery it obliges, and when a shared-memory fast path is worth reopening the boundary.
## Frame the decision as choosing a fault domain The real question is not "threads or processes" but **how much should fail together**. A fault domain is a set of work that a single defect can take down. Threads in one process are one fault domain; separate processes are separate domains, enforced by the memory-management hardware rather than by discipline. Three kinds of failure decide the answer: 1. **Crashes.** An invalid access, a stack overflow, or a fatal runtime error normally ends a whole process. In a threaded server that means every in-flight request dies, including the ones that were perfectly healthy. In a process-per-unit design, one request dies and the supervisor replaces the worker. 2. **Resource exhaustion.** Memory leaks, runaway allocation and file-handle leaks accumulate per process. Isolated workers can be capped and recycled after N units of work; in a shared process a leak in one code path degrades everything. 3. **Compromise.** Threads share credentials and address space, so "untrusted extension in a thread" is not a security boundary at all — anything it can address, it can read. A process can drop privileges, enter a sandbox and hold only the handles it needs, so a compromised worker yields little. Untrusted extension code makes (3) decisive and (1) very likely, which is why browsers, hardened web servers and plugin hosts converge on process isolation. ## What isolation costs Be honest about the bill, or the answer sounds like ideology. **Memory and startup.** Each worker carries its own runtime, its own caches and its own page tables. Shared read-only pages and copy-on-write reduce but do not remove this. Startup cost per worker pushes you toward pre-forked or pre-warmed pools rather than spawn-per-request. **Communication.** Pointer passing becomes message passing: serialize, copy through the kernel, deserialize. For small control messages this is irrelevant; for large payloads it can dominate, which pushes toward shared-memory or memory-mapped fast paths for bulk data — and those partially reopen the isolation you were buying. **Shared state that is genuinely shared.** Caches, connection pools and rate-limit counters are trivial in one address space and become a design problem across many: replicate per worker (memory cost, weaker hit rate), centralize in one owner process (a hop and a new failure mode), or externalize. **Operational complexity.** You now own a supervisor: health checks, restart policy, backoff so a crash loop does not become a restart storm, orphan cleanup, and a compatibility story between supervisor and workers during rolling upgrades. You are running a small distributed system inside a machine, with partial failure as a normal state. ## What the threaded design genuinely offers Cheap creation, cheap switching, zero-copy sharing, one warm cache, one set of connections, simple deployment. If all code is yours and memory-safe, and a crash is a genuine bug rather than an expected outcome, threads deliver more work per unit of hardware. Whole classes of systems are correct and fast this way. The hazards move rather than disappear: shared mutable state means synchronization, and synchronization means races, deadlocks and hard-to-reproduce bugs. Message-passing designs sidestep that specific class by construction — there is no shared object to race on — and pay for it in copies and in a new class of protocol and partial-failure problems. ## The hybrid, and how to choose the seams The common production answer is layered: a modest number of processes forming trust and fault boundaries, each internally multi-threaded for efficiency inside its boundary. The design work is choosing the seams: - **Put a process boundary where trust changes.** Untrusted extension, third-party codec, tenant-supplied code, anything parsing hostile input. - **Put one where failure semantics change.** A component whose crash must not cost unrelated work. - **Do not put one where two components exchange large data at high frequency**, unless you are prepared to build a shared-memory fast path. Then size it: how many workers, how much concurrency inside each, whether a worker is recycled after N units or a memory threshold, what the crash budget is before you consider the whole service unhealthy, and how a crash surfaces to the caller (a clean error for that unit, not a dropped connection for all). ## What a strong answer sounds like State the decision rule first — trust boundary and blast radius decide it, efficiency only breaks ties — then name the costs you are accepting, then describe the hybrid seam and the operational machinery (supervision, restart backoff, recycling policy, per-worker limits) that makes isolation actually deliver availability rather than just fragmenting the system. Mentioning that isolation converts memory-safety bugs into an availability problem you can budget for, while threads leave them as a correctness problem you must eliminate, shows you understand what you are actually buying.
- How do you keep process isolation from turning into a restart storm under a systemic bug?Track crash rate as a first-class signal with a budget, and apply exponential backoff with jitter to respawns so a bug that kills every worker instantly does not become a spawn loop. Above the budget, stop treating it as a local fault: shed load, degrade the feature, and alert. Isolation converts a crash into an availability question, so it needs an availability policy.
- Two isolated workers each need a large read-mostly cache. How do you avoid paying for it N times?Options in increasing complexity: back it with a memory-mapped read-only file so all workers share the same physical pages; move it into a single owner process queried over a local channel, accepting a hop and a dependency; or externalize it to a shared store. The read-only mapping is usually the best fit because it keeps zero copying without reintroducing shared mutable state.
- Where does a hybrid design put its process boundaries?Where trust changes and where failure semantics change — untrusted extensions, hostile-input parsers, per-tenant work, and any component whose crash must not cost unrelated requests. It avoids boundaries between components that exchange large payloads at high frequency, since every such boundary becomes a copy or forces a shared-memory fast path.
Threads are one open-plan kitchen: fast to pass a pan, but one fire closes the restaurant. Processes are separate kitchens behind fire doors: slower to move ingredients, but service continues while one is rebuilt.
saying these in an interview costs you the question
- Claiming a thread-level sandbox gives a real security boundary for untrusted code.
- Choosing processes purely 'for safety' without pricing memory, serialization and supervision.
- Assuming a crashed thread only kills itself, leaving the rest of the process healthy.
- Forgetting the supervisor's own concerns: restart backoff, crash budgets, worker recycling.
- Treating the choice as global when the right answer is usually a seam in one place — processes at the trust boundary, threads inside.