skip to content

Sharing Data Across Processes

Worker processes do not share objects, so every argument is pickled and every child copies memory. The fixes are a shared buffer, a memory-mapped file, and freezing the heap before fork.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why don't changes a multiprocessing child process makes to a global variable reach the parent?

level: juniorimportance: must knowfreq 55%

answer

  1. Processes are not threads
  2. Two address spaces, two copies
  3. The child's copy diverges immediately
  4. Results travel over a pipe, pickled
  5. Sharing needs shared_memory or a Manager

basics

~20 s

Each multiprocessing worker is a separate OS process with its own address space and its own copy of every module-level name. What the child rebinds or mutates changes only the child's memory; results must be sent back explicitly.

solid answer

~50 s

`multiprocessing` starts real OS processes, not threads, so parent and child have separate address spaces and separate copies of every module-level name. Whatever the child assigns, appends to or deletes touches only its own copy, and the parent's copy is untouched. What the child starts with depends on the start method: `fork` hands it a copy-on-write duplicate of the parent's memory, while `spawn` and `forkserver` start a fresh interpreter that re-imports your modules, so import-time globals are rebuilt and anything the parent set up after import is simply absent. On 3.14 the default is `forkserver` on Unix other than macOS, and `spawn` on macOS and Windows. To move data you need an explicit channel — a `multiprocessing.Queue`, a `multiprocessing.Pipe`, an executor's return value — or deliberately shared state such as `multiprocessing.Value`, a `multiprocessing.Manager` proxy or a `multiprocessing.shared_memory.SharedMemory` block.

go deeper

for a junior

Recall the one-liner: separate processes, separate memory, so a global changed in the child is invisible to the parent. Be ready to name at least one legitimate way back, such as returning the value or putting it on a queue.

for a middle

Explain the mechanics: which start method the child used, whether it inherited the parent's heap copy-on-write or re-imported the module, and why both rebinding and mutating stay local. Name the explicit channels and say that they pickle.

for a senior

Show the design judgment: decide what crosses the boundary and how often, avoid Manager proxies in hot loops because each access is an IPC round trip, and know that a codebase relying on inherited globals is fragile against a start-method change.

for a principal

Own the platform tradeoff. Inherited-global designs bind you to fork on Linux, which is now non-default and unsafe alongside threads; a design that passes an explicit handle or a shared buffer keeps the same code working on every start method and every OS.

## Two processes means two address spaces The single most important fact about `multiprocessing` is that it is not `threading`. A thread shares the interpreter's heap with every other thread in the process, so a global really is one object seen by everybody. A `multiprocessing.Process` is a separate operating-system process: the kernel gives it its own virtual address space, its own heap, its own `sys.modules` and therefore its own copy of every module-level name your code informally calls "a global". Parent and child agree on values only at the instant the child starts. From that instant the two copies evolve independently, and nothing propagates in either direction unless you write code to move it. ## What the child actually starts with depends on the start method Three start methods exist, and they give the child very different starting states. * **fork** — the child is a copy-on-write duplicate of the parent process. Every object the parent had is visible in the child at the moment of the fork, including globals built at runtime. Writes in either process are private. * **spawn** — a brand-new interpreter is launched. It imports your entry module afresh, so globals are re-created by re-running module-level code. Anything the parent computed *after* import does not exist in the child unless it was passed as a (pickled) argument. * **forkserver** — a small clean server process is forked once, imports the modules you ask it to, and every worker is forked from that server rather than from your possibly large, possibly multi-threaded main process. On Python 3.14 the default changed: Unix platforms other than macOS now default to `forkserver`, while macOS and Windows default to `spawn`. `fork` must be requested explicitly through `multiprocessing.get_context("fork")` or `multiprocessing.set_start_method("fork")`. Code written before 3.14 that quietly relied on the child inheriting a big preloaded global on Linux is exactly the code that breaks on upgrade: under `forkserver` the worker re-imports the module and the global is either rebuilt from scratch or missing. Because `spawn` and `forkserver` re-import the entry module, the `if __name__ == "__main__":` guard is mandatory around code that starts processes — without it the re-import restarts the pool recursively. ## Rebinding and mutating both stay local Two variants of the same misconception show up in interviews. The first is rebinding: the worker does `counter = counter + 1` on a module-level `counter` and the parent still prints the original value. The second is mutating: the worker appends to a module-level list inherited through `fork` and expects the parent's list to grow. Neither works. Copy-on-write means the child could *read* the parent's list pages, but the moment it writes, the kernel gives the child a private copy of the affected page, and from then on the two lists are unrelated objects at the same address in two different address spaces. ## How data actually crosses the boundary Everything that crosses does so through an explicit mechanism, and each has a cost: * **Message passing** — `multiprocessing.Queue`, `multiprocessing.Pipe`, or the return value of a worker function submitted to a process pool. The payload is pickled in the sender, copied through a pipe or socket, and unpickled in the receiver. Simple, but it copies. * **Shared primitives** — `multiprocessing.Value` and `multiprocessing.Array` place a small C-typed value or array in a shared mapping that all children see. Access still needs a lock for read-modify-write sequences. * **Manager proxies** — `multiprocessing.Manager()` starts a *server process* that owns the real objects; children hold proxies, and every attribute read or method call is an IPC round trip. Convenient and slow, and a proxied list mutated element-by-element pays a round trip per element. * **Raw shared buffers** — `multiprocessing.shared_memory.SharedMemory` maps one named block of bytes into several processes, and a memory-mapped file via `mmap` does the same for on-disk data. Both avoid copying but give you no structure and no synchronization. ## Why the question is asked It is the fastest way to find out whether a candidate has actually shipped multiprocessing code or has only pattern-matched it onto threading. The follow-through matters too: once you accept that nothing is shared by default, the design question becomes *what* to move and *how often* — which is why the expensive part of most process-parallel Python is not the computation but the data crossing the boundary.

  • If the child mutates a list it inherited through fork instead of rebinding the name, does the parent see it?
    No. `fork` gives the child copy-on-write pages, not shared pages. The first write to a page makes the kernel hand the child a private copy, so the append lands in the child's copy only. Copy-on-write is an optimization to avoid copying memory eagerly, never a sharing mechanism.
  • How do you get a value computed in a worker process back to the parent?
    Send it over an explicit channel: put it on a `multiprocessing.Queue`, write it to a `multiprocessing.Pipe`, or return it from the function you submitted to a process pool so the pool ships it back. All three pickle the object in the child and unpickle it in the parent, so return small results and keep bulk data in a shared buffer if it is large.
  • Why do spawn and forkserver require the entry module to be guarded by a main check?
    Both start a fresh interpreter that imports the entry module to rebuild the worker's globals. Without `if __name__ == "__main__":` around the code that creates processes, that import re-runs process creation, and each new child does it again — an unbounded recursion of processes. `fork` does not re-import, which is why the bug appears only after switching start methods.

Forking a process is like photocopying a filing cabinet: the copy starts identical, but writing on a page in one cabinet never changes the other.

saying these in an interview costs you the question

  • Says child processes share memory the way threads do
  • Expects a global assigned in a worker to update the parent
  • Thinks copy-on-write means the parent sees the child's writes
  • Assumes a mutated list crosses the process boundary in place
  • Confuses multiprocessing.Queue with a plain shared list
  • Believes a Manager proxy is free because it looks like a normal object

context

open as a page

Why does passing a large list as a multiprocessing task argument cost so much?

level: middleimportance: should knowfreq 58%

basics

~20 s

Worker processes share no objects, so every argument is pickled in the parent, pushed through a pipe, and unpickled in the child, and the result makes the same trip back. A large list pays that cost on every task.

open as a page

How does multiprocessing.shared_memory.SharedMemory avoid pickling a large buffer between processes?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It allocates one named block of raw memory that the operating system maps into every process that attaches to it. Only the short name travels between processes; reads and writes through the block's memoryview hit the same physical bytes, with no serialization.

open as a page

Why does a forked Python worker's memory climb even when it only reads inherited data?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Copy-on-write shares pages until something writes to them, and CPython writes constantly: touching any object updates its reference count, and the cyclic collector writes to object headers as it traverses. Those writes dirty pages, so read-only access privately copies the heap anyway.

open as a page