skip to content

Start Methods: spawn, fork, forkserver

How a child is born changes everything: fork copies the parent's memory along with its locks and threads, while spawn re-imports your module. The __main__ guard and fork's retreat are standard fodder.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why does multiprocessing code need an `if __name__ == "__main__":` guard under spawn?

level: juniorimportance: must knowfreq 70%

answer

  1. The child does not inherit memory
  2. Something has to be imported first
  3. Top-level code runs again in the worker
  4. Guard it or recurse forever
  5. spawn and forkserver, never fork

basics

~20 s

Under the spawn start method the child launches a fresh interpreter and re-imports your main module to reach the target function. Without the guard, the module-level code that launched the process runs again in the child, so CPython raises RuntimeError.

solid answer

~50 s

The `spawn` start method does not copy the parent's memory. It starts a brand-new interpreter, which **imports the parent's main module** (as `__mp_main__`) so that the pickled target function can be resolved by module and qualified name. Any statement at module level therefore executes a second time inside every child. If one of those statements is `Process(...).start()` or `Pool(...)`, each child would try to create more children forever, so CPython detects the situation during bootstrap and raises `RuntimeError` telling you to add the `if __name__ == "__main__":` guard and, for frozen executables, `multiprocessing.freeze_support()`. The same requirement applies to `forkserver`, which also imports the main module. Only `fork` escapes it, because the child is a memory copy and never re-imports anything. Since Python 3.14 the Unix default is no longer `fork`, so the guard is now needed almost everywhere.

code

python · 9 lines
python
import multiprocessing as mp

def work(n):
    return n * n

if __name__ == "__main__":
    ctx = mp.get_context("spawn")
    with ctx.Pool(2) as pool:
        print(pool.map(work, [1, 2, 3]))

go deeper

for a junior

Be ready to say that a spawned worker starts a fresh interpreter and imports your script to find the function, so unguarded top-level code runs again in every child. Know the fix by heart.

for a middle

Explain the mechanism: functions pickle by module-and-name, the child imports the main module as __mp_main__, and CPython raises RuntimeError when a process is started during bootstrap. Say which start methods need the guard.

for a senior

Show you know what the re-import costs in production: per-worker import time, side effects repeated once per worker, and module-level state that the parent filled in after import silently arriving empty in the child.

for a principal

Own the migration angle — Python 3.14 changed the Unix default away from fork, so a codebase full of guardless scripts and fork-inherited globals needs an audit before the upgrade, not a hotfix after it.

A **start method** is the recipe `multiprocessing` uses to bring a worker process into existence. CPython 3.14 ships three: `fork`, `spawn` and `forkserver`. The `__main__` guard question is really a question about what `spawn` has to do that `fork` does not. ### What spawn actually does When you call `Process.start()` under `spawn`, the parent launches a *new* Python interpreter (roughly `sys.executable` with a bootstrap argument), hands it a pipe, and sends it a small pickled payload: the target callable, the arguments, and some process state. Nothing of the parent's memory is inherited — no imported modules, no open objects, no module-level variables. That creates a bootstrapping problem. A function is not pickled by value; it is pickled **by reference**, as "the name `work` in module `__main__`". For the child to unpickle that reference it must first *import* the module the function came from. So the child imports your script — under the name `__mp_main__`, precisely so `__name__ == "__main__"` is false there — and only then resolves the target and runs it. ### Why the guard is mandatory Importing your script executes every top-level statement in it a second time, in every child. If a top-level statement starts processes, each child would start more children, which would import the module again, and so on. Rather than let you fork-bomb your own machine, CPython checks during child bootstrap whether a process is being started before the current process has finished bootstrapping, and raises `RuntimeError` with a message that spells out the fix: put the process-starting code inside `if __name__ == "__main__":`, and call `multiprocessing.freeze_support()` if the program is being frozen into an executable. The rule of thumb is therefore: **definitions above the guard, actions below it.** Imports, function and class definitions, and constants must stay at module level, because the child needs them to resolve the target. Anything with a side effect — starting processes or pools, parsing `sys.argv`, opening files or network connections, writing logs, running a long computation — belongs under the guard, or it will run once per worker. ### The same rule for forkserver, not for fork `forkserver` also imports the main module in the worker, so it needs the guard too. `fork` does not: the child is a copy of the parent's address space and simply continues after the fork point, so nothing is re-imported and no guard is required. That is exactly why so much older code omits the guard and still worked on Linux — and why it started failing on Python 3.14, where the default on Unix other than macOS became `forkserver`. macOS has defaulted to `spawn` since 3.8, and Windows has only ever supported `spawn`. ### Related consequences of the re-import Because the child re-imports rather than inherits, module-level mutable state that the parent filled in *after* import does not travel. A cache populated inside `main()` is empty in a `spawn` or `forkserver` child, while under `fork` it would arrive pre-filled. Code that quietly depended on that inheritance breaks when the start method changes, and the failure looks like a logic bug rather than a concurrency bug. The re-import also means the main module must be **importable at all**. Interactive sessions have no importable `__main__`, which is why `spawn` workers defined in a REPL cannot resolve their target; the target function must live in a real, importable module. ### What the error looks like A missing guard under `spawn` surfaces as a `RuntimeError` raised in the child during bootstrap, usually repeated for each worker and often accompanied by a partially started pool. It is not a hang and not a pickling error — recognising the bootstrap message is the fastest route from symptom to fix. If instead you see the *parent's* top-level work being redone with no error, you are usually looking at side-effecting module-level code that is legal but wasteful: it runs once per worker, multiplying startup cost by the pool size. ### The cost that follows from the same fact Because every worker imports your module graph, a spawned worker's start-up cost scales with how heavy your imports are rather than with how much work the task does. A pool created once at start-up pays that cost once per worker; a pool created inside a request handler pays it on every request, which is how a change of start method turns a fast endpoint into a slow one without a single line of the task changing. The same reasoning explains why an expensive table or model loaded at module level is a poor fit for `spawn`: it is rebuilt in each worker instead of being inherited. Pass it as an argument, load it once in a pool initializer, or keep the pool alive across requests.

  • Which start methods require the guard, and which does not?
    `spawn` and `forkserver` both import the parent's main module in the child, so both require it. `fork` does not: the child is a copy of the parent's address space and continues from the fork point without re-importing. Since 3.14 the Unix default (other than macOS) is `forkserver`, so guardless code that ran on older Linux now fails.
  • What belongs above the guard and what belongs below it?
    Above: imports, constants, and the function and class definitions that a worker must import to resolve its pickled target. Below: anything with a side effect — starting processes or pools, argument parsing, opening files or connections, logging setup, the actual computation. Side-effecting top-level code is legal but runs once per worker.
  • Why does the child module get the name `__mp_main__`?
    So that re-importing it does not make `__name__ == "__main__"` true a second time. The child imports the parent's script under that alias, which keeps the guarded block from executing while still making the module's top-level definitions available for unpickling the target callable.

fork photocopies a running kitchen mid-service; spawn hands the new cook the recipe book and says "start from the top" — so anything you wrote outside the recipe gets cooked again.

saying these in an interview costs you the question

  • Claims the guard is only a style convention
  • Says the child inherits the parent's imports under spawn
  • Thinks fork also re-imports the main module
  • Puts pool creation at module level and blames pickling
  • Believes the guard is a Windows-only requirement

context

open as a page

What distinguishes multiprocessing's fork, spawn and forkserver start methods, and which is default where?

level: middleimportance: must knowfreq 60%

basics

~20 s

fork clones the parent's memory, so it is fast but inherits threads and held locks. spawn starts a clean interpreter and re-imports, so it is slow but safe. forkserver forks each worker from a small single-threaded server, combining most of both.

open as a page

Why can multiprocessing's fork start method deadlock a child of a threaded service?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Forking duplicates only the calling thread but copies every lock exactly as it was. A lock another thread held at fork time arrives in the child locked forever, with no owner left to release it, so the child hangs.

open as a page

When should a library use multiprocessing.get_context() instead of set_start_method()?

level: middleimportance: nice to knowfreq 25%

basics

~10 s

multiprocessing.set_start_method() mutates process-global state and is meant to be called once by the application. A library should call multiprocessing.get_context("spawn") and use that context's Process, Pool and Queue, leaving everyone else's choice untouched.

open as a page