skip to content

Why does a metrics scraper's __main__ module body re-run inside every spawned worker process?

level: seniorimportance: should knowfreq 34%

answer

  1. The child starts with none of your objects
  2. Something has to rebuild the entry module
  3. Rebuilt by executing the same file again
  4. Bound under a different module name
  5. Body re-runs, guarded block does not

basics

~20 s

A child built from a fresh interpreter inherits none of the parent's memory, so multiprocessing rebuilds the entry module by executing its file again — under a substitute module name, so the module body runs but the name guard block does not.

solid answer

~50 s

A child created by `spawn` or `forkserver` starts a brand-new interpreter with none of the parent's objects, yet it still has to unpickle a target function that refers to the main module. `multiprocessing` reconstructs that module by re-executing the parent's file in the child and binding it to a substitute module name it reserves — not to `"__main__"`. So **module-level code runs again in every worker, and code behind `if __name__ == "__main__":` does not.** For a scraper that builds a collector list at module level and passes it as a default argument, each worker gets its own fresh empty list, so accumulated state silently stops accumulating; likewise every module-level connection, timer or banner is paid per worker. The fix is discipline, not configuration: cheap side-effect-free module bodies, all startup behind the guard, and worker state passed explicitly rather than read from globals.

code

python · 14 lines
python
import os

SEEN = []
print("module body ran in pid", os.getpid())


def collect(sample, seen=SEEN):
    seen.append(sample)
    return len(seen)


if __name__ == "__main__":
    print("guarded startup, workers:", 11)
    print("collected:", collect("scrape"))

go deeper

for a junior

Recall the rule of thumb: with a fresh-interpreter start method, anything written at module level happens again in each worker, so put startup behind the name guard and never create processes at module level.

for a middle

Explain the mechanism: the child has no copy of the parent's memory, callables pickle by reference, so the entry module is rebuilt by executing its file again under a substitute name, which is why the guarded block is skipped.

for a senior

Diagnose from symptoms — a banner printed once per worker, reset counters, per-worker copies of what looked like shared state — and fix it by making module bodies pure, guarding startup, and passing worker state explicitly instead of reading globals.

for a principal

Own the platform question: the 3.14 default change makes fresh-interpreter children the norm, so set the convention that module bodies are import-safe, decide where expensive initialisation happens once, and treat 'works because fork copied it' as a defect rather than a design.

## What the child actually has to rebuild When `multiprocessing` starts a worker with `fork`, the child is a copy of the parent's address space: every module already imported, every global already built. Nothing needs re-running. Under `spawn` and `forkserver` the child is a *fresh interpreter* — it shares no objects with the parent at all. To run your target callable it must unpickle it, and a function pickles **by reference**: module name plus qualified name. If the callable was defined in the parent's `__main__` module, the child has to produce a module that contains it. It does that by locating the parent's main file and executing it again in the child. Crucially, it executes it under a **substitute module name that `multiprocessing` reserves for exactly this purpose**, not under `"__main__"`. That single detail explains the whole behaviour: - Everything at module level — imports, constants, class and function definitions, and any side effects — **runs again, in every child**. - Everything inside `if __name__ == "__main__":` — **does not**, because the comparison is false there. ## Why this now bites almost everyone In Python 3.14 the default start method on Unix platforms other than macOS became `forkserver`; macOS and Windows already defaulted to `spawn`, and `fork` must now be requested explicitly. `forkserver` forks from a clean server process that does not carry your application state, so it re-imports the main module just as `spawn` does. Code that quietly relied on `fork`'s copied memory on Linux — module-level singletons, an already-open pool, a populated cache — changes behaviour on 3.14 without a line of it being edited. For a team sharing one CLI, that shows up as "it works on the release box and not on mine", because the two platforms had different defaults for years. ## The failure shape to recognise A metrics scraper whose module body builds a collector and then hands it out as a default argument is the compact version of the bug: ```python import os SEEN = [] # rebuilt in every re-import of this module print("module body, pid", os.getpid()) def collect(sample, seen=SEEN): seen.append(sample) return len(seen) ``` In a single process this is the familiar shared-mutable-default idiom: one list, bound once at `def` time, growing across calls. Fan the scraper out across eleven workers and every child re-executes the module body, so every child binds a *different* fresh list as the default. Deduplication that "worked" silently stops working, each worker re-scrapes what another already had, counts come back low, and nothing raises. The banner printed at module level appears once per worker, which is usually the first visible clue. The other classic symptom is recursion: if a worker is started at module level with no guard, the child re-executes the body, which starts another worker, and so on. `multiprocessing` detects this and raises a `RuntimeError` whose message tells you to add the `__name__` guard — and mentions `multiprocessing.freeze_support()`, which is the extra call a frozen executable needs. ## What to do about it 1. **Guard all startup.** Argument reading, logging configuration, connections, timers, process creation: below the line. `multiprocessing.Process` must never be constructed at module level. 2. **Keep module bodies cheap and pure.** Definitions and small constants only. Anything that costs time is paid once per worker, so an expensive module body turns into a slow, mysteriously CPU-hungry startup for a big pool. 3. **Pass state explicitly.** Do not treat a module-level object as shared: with a fresh-interpreter start method there is no sharing, and each worker gets its own copy of whatever the module body built. Send the data as arguments, or use an explicit shared mechanism. 4. **Beware what module-level code touches.** An open file descriptor, a logging handler pointed at the same file, or a random seed fixed at module level is now created once per worker, with all the interleaving that implies. 5. **Do not "fix" it by requesting `fork`.** It removes the symptom by reintroducing the hazards that motivated the default change — forking a process with threads or held locks is the reason `forkserver` became the default in the first place. ## The interview answer Name the cause (a fresh interpreter that must rebuild the main module by re-executing its file), name the precise consequence (module body yes, guarded block no, because the child binds a substitute module name), and name the fix (thin, pure module bodies with everything else behind the guard). Mention that 3.14 changed the Unix default to `forkserver`, so this is now the normal case rather than a Windows quirk.

  • Under spawn, does the code inside the __name__ guard also run in the child?
    No. The child re-executes the parent's main file, but binds it to a substitute module name that `multiprocessing` reserves, so `__name__ == "__main__"` is false there. Only the module body runs. That asymmetry is what makes the guard the correct place for startup: it is executed once, in the parent, no matter how many workers are created.
  • What error appears if a worker process is created at module level with no guard?
    The child re-executes the module body, which creates another worker, which re-executes it again. `multiprocessing` detects the recursion and raises a `RuntimeError` whose message says to add the `if __name__ == "__main__":` guard, and mentions `multiprocessing.freeze_support()` for frozen executables. The fix is to move the process creation below the guard.
  • Which multiprocessing default changed in Python 3.14, and why does it matter for this?
    On Unix platforms other than macOS the default start method became `forkserver`; macOS and Windows stay on `spawn`, and `fork` must be asked for explicitly. `forkserver` forks from a clean helper process, so like `spawn` it re-imports the main module. Linux code that depended on `fork` copying an already-initialised module suddenly sees module bodies run again and globals reset.
  • How do you keep expensive setup out of every worker?
    Build it once behind the guard and hand the result to workers as an argument, or build it lazily inside the worker function so it is created deliberately rather than as an import side effect. Module bodies should hold definitions and cheap constants only, because whatever they cost is multiplied by the worker count under a fresh-interpreter start method.

The child is not a photocopy of the parent's desk, it is a new hire handed the same instruction sheet: everything written on the sheet gets done again, except the paragraph addressed only to the original.

saying these in an interview costs you the question

  • Says children inherit the parent's globals under spawn
  • Thinks the guarded block also executes in the child
  • Treats a module-level object as shared across processes
  • Creates worker processes at module level with no guard
  • Believes fork is still the default on Linux in 3.14
  • Blames pickling alone without mentioning main-module re-execution

context