A search-index rebuilder moved from threads to ProcessPoolExecutor and got slower while memory climbed — which costs of the process model explain that?
answer
- Parallel gain minus three overheads
- Startup, transfer and memory per worker
- Pickles cross a pipe both ways
- Send identifiers, not whole payloads
- Bound submissions or the parent grows
basics
~20 sProcesses pay for interpreter startup, per-worker memory and pickling every argument and result through a pipe. Large payloads and per-worker copies of module-level caches can cost more than the parallelism gains, and unconsumed results pile up in memory.
solid answer
~50 sThree costs bite. **Startup**: each worker is a fresh interpreter that re-imports your modules and rebuilds module-level state; with `forkserver` or `spawn` that is seconds, not microseconds. **Transfer**: every argument and every return value is pickled, written to a pipe and unpickled, so shipping whole documents to workers that do a little work per document can dominate the run. **Memory**: N workers means N copies of the imported modules and any cached structure they build, which is exactly how a rebuild that fitted before starts climbing. On top of that, submitting the whole corpus up front materializes every pending argument and every completed result in the parent. The fixes are to send identifiers rather than payloads, chunk the work so each task earns its transfer cost, bound submission instead of queueing everything, and share large read-only buffers through `multiprocessing.shared_memory`.
code
python · 7 linesimport pickle
docs = [{"doc_id": i, "body": "term " * 200} for i in range(50_000)]
per_item = len(pickle.dumps(docs[0]))
print(f"{per_item} bytes per document")
print(f"{per_item * len(docs) / 1e6:.1f} MB if each document is shipped once")
print(f"{len(pickle.dumps([d['doc_id'] for d in docs])) / 1e6:.2f} MB if only ids are shipped")go deeper
Remember that processes do not share memory: everything sent to a worker is copied, and each worker is a full interpreter with its own imports and its own RAM.
Explain the three cost lines — startup, pickling through a pipe, and per-worker memory — and the standard mitigations: chunking, sending identifiers instead of payloads, and choosing a sensible worker count.
Diagnose it with numbers: payload size per task, first-completion latency, per-worker resident memory over time. Then say which of bigger chunks, shared memory, fewer workers, or a different model you would apply, and why.
Frame it as a capacity decision: the memory ceiling a process model imposes per host, what it does to deployment density and cost, and when the honest answer is to move the heavy stage out of Python rather than tune the pool.
## Why the process model is not free Swapping a thread pool for a process pool changes three cost lines at once, and a workload that is only mildly CPU-bound can lose on all three. **Interpreter startup.** A worker process is a whole new interpreter. It imports your package tree again, re-executes module-level code, rebuilds any caches your modules construct at import, and only then takes its first task. Under `spawn` and `forkserver` that is a fresh interpreter every time; only `fork` inherits the parent image cheaply, and it does so at the price of copying a process mid-flight with locks and threads in unknown states. Python 3.14 changed the default on Unix other than macOS to `forkserver`; macOS and Windows already used `spawn`, and `fork` must now be requested explicitly. If your job previously relied on forked workers inheriting a warm in-memory structure for free, 3.14 is where that stopped happening silently. **Data transfer.** Arguments and results cross the boundary as pickles through a pipe. The parent serializes, the OS copies, the child deserializes — and the reverse for results. The rule of thumb is to compare bytes moved against work done: shipping a full document to a worker that spends a millisecond on it is a losing trade; shipping a document identifier and letting the worker fetch or memory-map the content is a winning one. Chunking helps for the same reason — one task per item pays the round-trip per item, while a task carrying a few hundred items amortizes it. **Memory multiplication.** Every worker has its own heap. Module-level caches, compiled patterns, warm dictionaries, an in-memory vocabulary: whatever the parent built once, eight workers now build eight times. This is the usual cause of a rebuild whose resident memory climbs steadily after a switch to processes, and it gets worse if each worker accumulates state across tasks because the pool reuses workers for the whole run. ## The unbounded-growth pattern specifically Memory that grows without bound during a process-pool run usually comes from the parent, not the workers. Submitting every unit of work up front creates one pending future per unit, each holding its pickled arguments alive; results then land and are held until something consumes them. A 27-minute rebuild that submits a million documents in a loop and iterates results afterwards will hold both sides in RAM at the peak. Two disciplines fix it: bound the number of in-flight submissions (submit a window, consume completions, top the window up), and iterate results as a stream rather than materializing a list of them. A worker-side leak is possible too — a module-level cache that only grows — and `maxtasksperchild`-style worker recycling exists precisely to cap that. ## Deciding whether processes were the right model at all Before optimizing the process pool, re-ask the classification. If the hot frames were inside a C extension that releases the interpreter lock, the original thread pool was already parallel and the move to processes bought nothing while adding all three costs. If the units of work are small and numerous, the transfer cost may exceed the compute per item no matter how you tune it, and batching is not optional. If the workers all need the same large read-only structure, the choices are: build it once and expose it through `multiprocessing.shared_memory`, memory-map it from disk, or reconsider the free-threaded build where threads share it natively. ## What to measure Quantify rather than argue. Measure the pickled size of one task payload with `pickle.dumps` and multiply by the task count; time one worker startup by timing the pool's first completion versus its tenth; watch per-worker resident memory over the run rather than only the total. Those three numbers tell you immediately whether the answer is bigger chunks, fewer bytes, fewer workers, or a different model entirely.
- How would you decide between shipping the data and sharing it?By comparing bytes moved to work done per task. If each task does milliseconds of work on megabytes of input, transfer dominates and you should ship an identifier and let the worker read the content, memory-map the file, or attach to a `multiprocessing.shared_memory` block the parent created once. If tasks do seconds of work on kilobytes, pickling is noise and simplicity wins.
- What in 3.14 could make a previously fast process pool slower without any code change?The start-method default. On Unix other than macOS it is now `forkserver` rather than `fork`, so workers no longer inherit the parent's warm heap; they start from a clean server process and import your modules. Any design that relied on forked inheritance of a large preloaded structure now pays for building it per worker, or must move it to shared memory.
- Where does the unbounded memory usually live — parent or workers?Most often the parent. Submitting every unit up front keeps a future plus its pickled arguments alive for each one, and completed results accumulate until consumed. Bound the in-flight window and stream results. A worker-side cause is a module-level cache that only grows, which worker recycling after a fixed number of tasks caps.
Hiring eight contractors instead of eight helpers: each brings their own toolbox and needs the blueprints couriered over and back. Worth it for big jobs, ruinous for many tiny ones.
saying these in an interview costs you the question
- Treats processes as free parallelism with no transfer cost
- Submits every unit of work up front and calls it streaming
- Ignores that each worker re-imports modules and duplicates caches
- Assumes workers inherit the parent heap on every platform
- Adds more workers when the bottleneck is serialization
- Never checks whether the hot code already released the interpreter lock