skip to content

Sharing State Between Processes

Processes share no memory, so every value crosses through a queue, a pipe, a shared-memory block or a Manager proxy. Interviewers probe the copy cost that makes multiprocessing lose on small tasks.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why doesn't a global list mutated inside a multiprocessing.Process appear changed in the parent?

level: juniorimportance: must knowfreq 70%

answer

  1. Two processes, two address spaces
  2. The child mutates a different object
  3. Copy-on-write diverges on first write
  4. Sharing must be named, not assumed
  5. Queue, Pipe, Value/Array, shared_memory, Manager

basics

~20 s

Each multiprocessing.Process is a separate OS process with its own memory, so the child mutates its own copy and nothing flows back. Sharing needs an explicit channel: a Queue, a Pipe, Value/Array, a shared memory block or a Manager proxy.

solid answer

~40 s

Threads share one address space; processes do not. When `multiprocessing.Process` starts a child, that child either re-imports your module (under the `spawn` and `forkserver` start methods) or gets a copy-on-write snapshot of the parent's pages (under `fork`). Either way the child's `items` is a *different* list object living in different memory, so `items.append(...)` is invisible to the parent, and even under `fork` the first write breaks the copy-on-write page and diverges. To actually move data you must name a mechanism: `multiprocessing.Queue` or `multiprocessing.Pipe` to send objects by copy, `multiprocessing.Value`/`Array` for a small block of shared C-typed memory, `multiprocessing.shared_memory.SharedMemory` for a raw shared byte buffer, or `multiprocessing.Manager()` for proxied `list`/`dict` objects served by a helper process. Each of those has a different copy cost.

code

python · 13 lines
python
import multiprocessing as mp

items = []

def add():
    items.append("child")
    print("in child:", items)

if __name__ == "__main__":
    p = mp.Process(target=add)
    p.start()
    p.join()
    print("in parent:", items)  # []

go deeper

for a junior

Be ready to state the core fact in one sentence: separate processes have separate memory, so a global changed in the child is invisible to the parent. Know that a Queue is the usual way to get a result back.

for a middle

Explain the mechanics: what the child inherits under fork versus what it re-imports under spawn and forkserver, why copy-on-write diverges on first write, and the four transports available with their rough copy costs.

for a senior

Show that you design for it. Name which transport carries each piece of state in a real pipeline, and diagnose the silent-data-loss shape where a worker mutates a global that nobody ever reads back.

for a principal

Own the tradeoff: whether the workload justifies process isolation at all given the serialization tax, and whether shared mutable state across processes should exist rather than being replaced by returning results and aggregating in one place.

### The one fact everything else follows from A thread shares its parent's address space; a *process* does not. `multiprocessing.Process` creates a real operating-system process, and the operating system gives it a private virtual address space. A module-level name like `items` is not one object seen by two processes — after the child starts there are two `items` lists at two addresses, and mutating either one tells the other nothing. This surprises people because the code *looks* like the threading code that works. With `threading.Thread`, `items.append("x")` in the worker really is visible in the main thread, because both run inside one interpreter in one address space. Swap in `multiprocessing.Process` and the identical line silently does nothing useful. ### What the child actually starts with How the child gets its initial state depends on the start method, and the two families behave differently enough to be worth knowing: * **`spawn` and `forkserver`** — the child is a fresh interpreter that **imports your module again**. Module-level code runs a second time, so `items` is rebuilt as a brand-new empty list. Anything the parent computed after import time simply is not there unless it was pickled and passed as an argument. This is why a target function and its arguments have to be picklable, and why module-level code needs an `if __name__ == "__main__":` guard. * **`fork`** — the child begins as a duplicate of the parent, so `items` *starts out* holding whatever the parent had. That is a trap, not a feature: the pages are **copy-on-write**, so the first mutation gives the child its own private copy of the page and the two diverge from that instant. Reads look shared; writes never are. On Python 3.14 the default start method on Unix other than macOS is `forkserver`; macOS and Windows use `spawn`. `fork` must be asked for explicitly. So on a modern default install the child usually re-imports and starts from a clean module state. A further wrinkle: even *reading* under `fork` is not free. CPython touches an object's reference count on almost every access, and the refcount lives in the object header, so merely iterating a large parent list in the child writes to those pages and copies them. "Copy-on-write means my big read-only dict is free" is wrong in CPython. ### The mechanisms that do share Once you accept that nothing is shared implicitly, sharing is a choice among four shapes: 1. **Message passing by copy** — `multiprocessing.Queue`, `multiprocessing.SimpleQueue`, or the pair of connection objects from `multiprocessing.Pipe`. Objects are serialized in the sender, written to an OS pipe or socket, and rebuilt in the receiver. Nothing is shared; a *copy* arrives. This is the default answer and the right one for results, work items and events. 2. **Shared C-typed memory** — `multiprocessing.Value` and `multiprocessing.Array` allocate a small block of memory mapped into every child, holding one C value or a fixed-length array of them. Real sharing, no copy, but only for machine types, and you must synchronize writes. 3. **A raw shared byte block** — `multiprocessing.shared_memory.SharedMemory` maps a named region that any process can attach to by name and read or write through a `memoryview`. Zero-copy for large buffers; you own its lifetime and its layout. 4. **A proxied server** — `multiprocessing.Manager()` starts a helper process that owns real `list`, `dict`, `Namespace`, `Lock` and similar objects, and hands out proxies. The proxy makes the code look like ordinary Python, but every method call is a round trip to that process. ### The mental model to carry into the interview Say it as a rule: **in multiprocessing, state does not leak — it is transported.** Then, for any variable in a proposed design, ask *which* of the four transports carries it and what that costs. A worker computing a summary should return it (queue or pool result); a large read-mostly buffer belongs in shared memory; a small counter belongs in `Value` under its lock; a convenient shared dict for coordination is what a `Manager` proxy is for, at proxy-round-trip prices. The corresponding anti-pattern is a global that a child mutates and a parent later reads. It does not raise, it does not warn, and under `fork` it can even look plausible in a small test because the pre-fork contents are there. It just quietly loses every update.

  • Under the fork start method the child can already see the parent's data. Why is that not real sharing?
    Because the pages are copy-on-write. The child starts with the same physical pages mapped read-only-ish, and the first write to a page gives the child a private copy; from then on the two processes diverge. Reads can look shared, writes never propagate in either direction. In CPython even reading is not free, since touching an object updates its reference count and dirties the page.
  • If you need a shared counter across four worker processes, which mechanism do you reach for and why?
    `multiprocessing.Value` with its default lock, because a counter is a single machine integer and belongs in shared C-typed memory rather than in a serialized message. A `Manager` proxy would also work and is more convenient, but every increment becomes a round trip to the manager process. Better still, have each worker count locally and send one total back through a queue, so the shared write happens once per worker.
  • Why must the target function and its arguments be picklable at all?
    Because under `spawn` and `forkserver` there is no memory to inherit: the parent serializes the callable reference and the argument tuple and writes them down a pipe to a fresh interpreter, which reconstructs them. That is why a lambda, a local closure, an open socket or a `threading.Lock` cannot be passed, and why the target usually has to be a module-level function reachable by import.

Threads are two people editing one shared document; processes are two people who were each handed a photocopy. Marking up your photocopy never changes anyone else's.

saying these in an interview costs you the question

  • Thinking processes share globals the way threads do
  • Claiming copy-on-write makes a large read-only structure free in CPython
  • Assuming a mutation in the child propagates back after join
  • Confusing multiprocessing.Queue with an in-process queue.Queue
  • Believing a passed object is shared rather than copied

context

open as a page

Why does multiprocessing.Queue deadlock when you join the child before draining it?

level: middleimportance: must knowfreq 55%

basics

~20 s

A multiprocessing.Queue is a bounded OS pipe fed by a background thread in the writing process. If the pipe fills, that thread blocks until the reader drains it, and the child cannot exit — so the parent's join() waits forever. Drain first, then join.

open as a page

Why does counter.value += 1 on a multiprocessing.Value race, and how do you fix it?

level: middleimportance: should knowfreq 45%

basics

~10 s

counter.value += 1 is a read, an add and a write, and another process can land between them, so increments are lost. Hold the value's own lock around the whole read-modify-write with with counter.get_lock():.

open as a page

When does multiprocessing.shared_memory beat a Manager proxy for a 200 MB invoice-bitmap payload?

level: seniorimportance: should knowfreq 35%

basics

~20 s

When the payload is large and read-mostly. A Manager proxy pickles a round trip to a server process on every access, so a 200 MB bitmap is copied repeatedly; multiprocessing.shared_memory.SharedMemory maps one block that workers read in place, at the price of owning its lifecycle and layout yourself.

open as a page