skip to content

Why did multiprocessing.Pool slow down a translation-memory updater that ships each record's full context?

level: seniorimportance: should knowfreq 50%

answer

  1. Parallelism is not free at the edges
  2. Something has to move between the processes
  3. Compare transfer cost against per-task work
  4. The parent can become the bottleneck
  5. Fewer, larger tasks amortize the fixed cost

basics

~20 s

Every argument and every result is pickled, piped and unpickled, so a task that ships a large context but does milliseconds of work pays more at the boundary than it saves. Send small handles and batch the records.

solid answer

~40 s

Parallel speedup only appears when per-task compute exceeds the cost of moving the task. In a pool that cost is a full round trip: the parent pickles the arguments, writes them to a pipe, a worker unpickles them, then the result is pickled back and unpickled in the parent. If each task carries a whole translation-memory context but only updates one record, the parent becomes a serialization bottleneck feeding four idle workers. The fixes are all about payload shape: batch many records into one task so the fixed cost is amortized, send identifiers or ranges instead of whole objects, and let each worker load the large read-only structure once for its lifetime rather than receiving it per task. Measure before guessing — time `pickle.dumps` on a representative payload against the per-record compute.

code

python · 9 lines
python
import pickle
import time

records = [{"id": i, "text": "x" * 200} for i in range(20_000)]

start = time.perf_counter()
blob = pickle.dumps(records)
elapsed = time.perf_counter() - start
print(f"{len(blob) / 1e6:.1f} MB pickled in {elapsed * 1000:.0f} ms")

go deeper

for a junior

Remember the core fact: data passed to a worker process is copied, not shared, and copying costs time proportional to its size. Small arguments and small return values are the safe default.

for a middle

Explain the full round trip — pickle in the parent, pipe, unpickle in the worker, and the same again for the result — and why batching records into fewer, larger tasks amortizes that fixed cost.

for a senior

Show the diagnosis: measure payload serialization against per-task compute, spot the parent pinned at full CPU while workers idle, look at the payload size distribution for outliers, and reshape the payload rather than adding workers.

for a principal

Own the decision framework: characterise the workload, price the boundary, and choose threads, processes or a single optimised loop accordingly. Also own the migration risk that 3.14's forkserver default turns previously free inherited data into explicit payload.

## The cost model A pool buys you parallel CPU and charges you serialization. For each task: 1. the parent calls `pickle.dumps` on the callable reference and the arguments, 2. writes the bytes into a pipe, 3. and a worker reads and unpickles them; 4. when the task finishes, the return value makes the same journey back. That is **two serializations and two deserializations per task**, plus the memory to hold the byte strings and the rebuilt objects on both sides. Parallelism pays only when per-task compute is comfortably larger than that fixed cost. ## How the pathology looks A translation-memory updater that fans out per record, passing each record together with the surrounding context it might need, hits the bad end of the curve: a few hundred microseconds of real work wrapped in a payload of tens or hundreds of kilobytes. Because a pool serializes tasks in the parent, the parent process pins one core at full CPU pickling while the workers sit idle waiting for their next task — the classic signature is **a parent at 100% CPU and workers well under it**, with wall-clock time no better than, or worse than, a plain loop. Result payloads cause the mirror image: if each task returns an enriched record rather than a small summary, the parent spends its time unpickling instead. ## Measure, do not guess Two numbers settle the argument. 1. First, `pickle.dumps` a representative argument and time it, and note the byte length; do the same for a representative return value. 2. Second, time the worker function called directly in-process with `time.perf_counter`. If serialization is a meaningful fraction of compute, the boundary is the bottleneck and adding workers cannot help. Payload *distribution* matters as much as the mean: with a shared updater maintained by a four-person team, one record type that carries an enormous segment list will produce an outlier task whose transfer dwarfs everything else, and because tasks flow through a single pipe, that one giant payload stalls the workers queued behind it. An intermittent timeout that never reproduces in a single-process run is very often that outlier, not a deadlock. ## Amortize with bigger tasks The most effective fix is to make each task do **more work per byte sent**. Group records into batches and send a batch per task; the per-task overhead is then divided across the batch, and the number of pipe round trips drops by the same factor. The tuning parameter is the ratio, not the worker count: aim for tasks that run for a noticeable slice of time — tens of milliseconds upward — rather than microseconds. Over-batching costs you load balance at the tail, so the sweet spot is many more batches than workers, but far fewer batches than records. ## Shrink what travels - Send identifiers, offsets or key ranges and let the worker fetch or compute the rest locally. - Return a small verdict rather than a rebuilt object. - Strip fields the worker does not read; a dict with ten keys where the function uses two is 80% waste on every task. - Where the large structure is read-only and shared by all tasks, load it once per worker process at start-up so it never appears in a task payload at all. ## What Python 3.14 changed The default start method on Unix platforms other than macOS is now `forkserver`, not `fork` (macOS and Windows already used `spawn`). - **Under `fork`**, a large structure built in the parent before the pool started was inherited by every child and could be read without ever being sent — an implicit, copy-on-write freebie. - **Under `forkserver`** the children come from a clean server process that only imported modules, so that structure is no longer there and must either be sent explicitly or rebuilt per worker. Code that was quietly fast on Linux can therefore get slower on 3.14 with no source change, and the diagnosis is exactly this payload accounting. ## When the boundary is the wrong tool If the per-record work is I/O-bound, or is dominated by a C extension that releases the GIL around a long computation, threads do the same job with zero serialization because they share the objects outright. Processes earn their overhead only for **CPU-bound pure-Python work with a small, cheap payload**; when the data is big and the work is small, the honest engineering answer is often to keep it in one process and optimise the loop instead.

  • How would you separate serialization cost from worker compute time?
    Measure both in isolation. Time `pickle.dumps` and `pickle.loads` on a representative argument and on a representative return value, and record the byte sizes; then time the worker function called directly in-process. If the round trip is a meaningful fraction of the compute, the boundary is the bottleneck. Confirm it at runtime by watching CPU: a parent pinned near 100% while workers idle is the serialization signature.
  • One record intermittently times out under the pool but never in a single process. Where do you look?
    At the payload size distribution, not the average. One record carrying an unusually large context produces a task whose transfer dwarfs the rest, and since tasks flow through one pipe from the parent, that single giant payload stalls every worker queued behind it. Log the pickled byte length per task, look at the tail rather than the mean, and cap or stream the outliers instead of raising the timeout.
  • When is a process pool simply the wrong tool for this workload?
    When the work is I/O-bound, or when it happens inside a C extension that releases the GIL around a long computation — threads then share the data outright and pay nothing at a boundary. Processes earn their overhead only for CPU-bound pure-Python work with a small payload. If the data is large and the per-item work is small, optimising the single-process loop usually beats parallelising it.

It is like couriering a filing cabinet across town so a colleague can correct one typo. The correction takes a second; the van does not. Send a box of a thousand corrections instead, or ask them to keep their own cabinet.

saying these in an interview costs you the question

  • Assumes more worker processes always means more throughput
  • Forgets that return values are pickled back too
  • Sends the whole dataset with every single task
  • Blames the GIL for a multiprocessing slowdown
  • Times the worker only, never the parent's serialization
  • Believes children still see parent globals under spawn

context