skip to content

What distinguishes multiprocessing's fork, spawn and forkserver start methods, and which is default where?

level: middleimportance: must knowfreq 60%

answer

  1. Three ways to get a child's initial state
  2. Copy, re-import, or fork from a helper
  3. Speed on one axis, inherited hazards on the other
  4. Defaults differ per platform and moved in 3.14
  5. Unix other than macOS now starts a server

basics

~20 s

fork clones the parent's memory, so it is fast but inherits threads and held locks. spawn starts a clean interpreter and re-imports, so it is slow but safe. forkserver forks each worker from a small single-threaded server, combining most of both.

solid answer

~50 s

`fork` copies the parent process, so children start in microseconds with every import and module-level object already in place — but they also inherit locks held by threads that do not exist in the child, which is why it is unsafe in a threaded program. `spawn` launches a fresh interpreter that re-imports the main module and receives only pickled arguments; it costs tens to hundreds of milliseconds per worker and inherits nothing, so it is the safe default. `forkserver` starts one small, single-threaded server process on first use and forks every worker from *that*, so children are cheap to create and free of the parent's threads; `multiprocessing.set_forkserver_preload()` pre-imports modules into the server to make them cheaper still. On CPython 3.14 the default is `forkserver` on Unix other than macOS, and `spawn` on macOS and Windows; `fork` must now be asked for explicitly.

code

python · 14 lines
python
import multiprocessing as mp

LOADED_AT_IMPORT = ["parent-only"]

def show():
    print(mp.current_process().name, mp.get_start_method(), LOADED_AT_IMPORT)

if __name__ == "__main__":
    LOADED_AT_IMPORT.append("added-after-import")
    for method in mp.get_all_start_methods():
        ctx = mp.get_context(method)
        p = ctx.Process(target=show)
        p.start()
        p.join()

go deeper

for a junior

Recall the three names and the one-line difference: fork copies the parent, spawn starts a clean interpreter, forkserver forks from a helper process. Know that the platform default is not the same everywhere.

for a middle

Explain the mechanics and the trade-off in both directions — start-up cost versus inherited threads, locks and module state — and state Python 3.14's defaults per platform without hedging.

for a senior

Demonstrate that you would pick per workload and measure it: worker churn, import weight, whether the parent is threaded, and what state your workers were quietly inheriting before the choice changed.

for a principal

Own the upgrade story: a fleet moving to 3.14 changes start method underneath itself, so the questions are which services relied on inherited state, what per-worker start-up cost the new default adds, and how the change is rolled out and verified.

The three start methods differ in one thing — where the child's initial state comes from — and everything else follows from that. ### fork The child is a copy of the parent's address space at the moment of the call. Every module already imported is there, every module-level object is there, open file objects are there, and start-up cost is a fraction of a millisecond because nothing is imported or pickled. Copy-on-write means large read-only data structures appear free at first, though CPython's reference counts live inside the object headers, so merely *touching* inherited objects dirties their pages and the sharing erodes. The cost is correctness. `fork` duplicates only the calling thread: any other thread simply does not exist in the child, and any lock those threads were holding stays locked forever. It also duplicates things that are not safe to duplicate — pooled connections, buffered file objects with unflushed data, cryptographic random state, and runtimes below CPython that assume a single process. `fork` is only available on Unix. ### spawn The child is a brand-new interpreter, started from `sys.executable`, that imports the parent's main module and unpickles a small payload describing the target callable and its arguments. It inherits nothing, so it is immune to the fork hazards, and it is the only method Windows supports. The price is start-up time — a fresh interpreter plus a full import of your application's module graph, which for an import-heavy program can be hundreds of milliseconds per worker — and a stricter contract: the main module needs a `if __name__ == "__main__":` guard, and everything the worker receives must be picklable. ### forkserver On first use the parent starts a dedicated server process, then every worker is forked from that server rather than from the parent. Because the server is created early and is single-threaded, the forked children never inherit application threads or their held locks. Because forking is still just a fork, worker creation stays cheap. The server starts nearly empty, so a worker only has whatever the server imported; `multiprocessing.set_forkserver_preload([...])` names modules to import into the server once, so the fork gives every worker those imports for free. Like `spawn`, `forkserver` imports the main module in the worker and requires the guard, and arguments must be picklable. It is Unix-only. ### Defaults, and how they got that way On CPython 3.14, `multiprocessing.get_start_method()` returns `forkserver` on Unix platforms other than macOS, and `spawn` on macOS and Windows. macOS moved to `spawn` in 3.8, because forking a process that has touched higher-level system frameworks tends to crash the child. Python 3.12 began emitting a `DeprecationWarning` when the implicit default `fork` was used from a multi-threaded parent, and 3.14 completed the retreat by changing the default outright. `fork` remains available on Unix; you now have to request it. `multiprocessing.get_all_start_methods()` reports what the running platform supports. ### Choosing Use `spawn` when you need cross-platform behaviour, when the parent is threaded or embeds another runtime, or when you want children whose state is exactly what you passed them. Use `forkserver` on Linux when worker churn matters — many short-lived workers, or a pool that recycles workers — and preload the heavy modules. Reach for `fork` only in a deliberately single-threaded parent where inheriting a large in-memory structure without re-loading it is the point, and even then know that you are opting out of the default for a reason you can state. ### The state-inheritance trap The most common surprise when moving off `fork` is not a crash but a wrong answer. Under `fork`, a module-level list, cache or configuration object that the parent mutated after import arrives in the child already populated; under `spawn` or `forkserver`, the child re-imports the module and sees only the value that the import itself created. A route-optimisation job that builds a distance table at start-up and mutates a module-level default in `main()` will happily produce different routes per worker after an upgrade changes the start method — with no exception anywhere. Pass such state explicitly as an argument, or rebuild it in a pool initializer, instead of relying on inheritance.

  • Why is copy-on-write less of a win in CPython than the theory suggests?
    Reference counts are stored in each object's header, so reading an inherited object writes to its page. Traversing a large inherited structure gradually copies most of it, and the garbage collector touching objects has the same effect. Sharing survives for `bytes`-like blobs and memory-mapped data far better than for graphs of Python objects.
  • What does set_forkserver_preload actually change?
    It names modules for the fork server to import once, before it starts forking. Every worker then inherits those imports for free instead of paying for them individually, which is the main way to close the start-up gap between `forkserver` and `fork` for an import-heavy application. It has no effect under `spawn` or `fork`.
  • When is fork still the right choice on Linux?
    When the parent is genuinely single-threaded and the point is to hand children a large in-memory structure without re-loading or pickling it — and when you can state that no library in the process holds a lock, a connection or a native runtime across the fork. It is now an explicit opt-in with a justification, not a default.

fork clones the whole workshop mid-job, spawn builds an empty workshop and hands over a parts list, and forkserver keeps one pristine template workshop and clones that.

saying these in an interview costs you the question

  • Says fork is still the Linux default on 3.14
  • Claims spawn is Windows-only behaviour
  • Thinks forkserver forks each worker from the parent
  • Assumes copy-on-write keeps inherited data free forever
  • Believes the start method changes what can be pickled

context