skip to content

When is maxtasksperchild on a multiprocessing.Pool worth setting?

level: seniorimportance: nice to knowfreq 22%

answer

  1. A voluntary exit, not a crash
  2. Bounds whatever a process accumulates
  3. Memory is one accumulation; staleness is another
  4. Every recycle pays a process start
  5. Size it against task duration

basics

~20 s

Set it when workers accumulate what they should not keep across tasks — leaked memory or stale per-process cached state. The worker retires after that many tasks and a fresh one replaces it; each recycle costs a process start.

solid answer

~50 s

`multiprocessing.Pool(maxtasksperchild=N)` retires each worker after it has completed N tasks and starts a replacement, so nothing a worker accumulates lives longer than N tasks. It buys two things: a bound on memory growth you cannot fix at the source — typically a leak inside a native extension — and a guarantee that per-process cached state cannot go stale, which matters when a long-lived worker memoizes something that changes underneath it. The cost is a fresh interpreter start per recycle, which is significant under the `spawn` and `forkserver` start methods and dominates entirely when tasks are short. So it is a containment tool, not a fix: use it when the underlying leak is not yours to repair or is not worth repairing now, size N so a recycle is amortized over real work, and record that it is masking something. `concurrent.futures.ProcessPoolExecutor` has the same knob spelled `max_tasks_per_child`, added in Python 3.11.

code

python · 11 lines
python
import os
import multiprocessing as mp


def worker_pid(_):
    return os.getpid()


if __name__ == "__main__":
    with mp.Pool(2, maxtasksperchild=1) as pool:
        print(len(set(pool.map(worker_pid, range(6)))))

go deeper

for a junior

Know that pool workers normally live for the whole run and that this setting retires each one after a fixed number of tasks so a fresh process replaces it. That is all you need at this level.

for a middle

Explain both accumulations it bounds — leaked memory and stale per-process cached state — and the cost, which is a full interpreter start per recycle and therefore start-method dependent.

for a senior

Show that you size it from measurements: worker memory against tasks completed, worker start cost against task duration, and headroom for workers that do not recycle in lockstep. Say plainly that it is containment and that the underlying defect stays on the list.

for a principal

Frame it as a decision about where the fix belongs: your own leak gets repaired, a dependency's leak gets contained, and stale cached state ideally gets a real invalidation channel rather than a rotation policy. Make sure the knob is documented with the reason it exists.

### What the parameter does By default a pool worker lives for the pool's whole lifetime, handling task after task in the same interpreter. `maxtasksperchild=N` changes that contract: after finishing N tasks the worker exits cleanly, and the pool starts a fresh process to take its place. The pool size stays constant; only the identity of the processes rotates. `concurrent.futures.ProcessPoolExecutor` gained the same knob as `max_tasks_per_child` in Python 3.11. A worker exiting under this policy is an *orderly* exit, not a crash: it finishes the task it is holding, and nothing outstanding is lost. That is what distinguishes it from every other worker-death story. ### The two problems it actually solves **Unbounded memory growth you cannot fix at the source.** Long-lived worker processes are where slow leaks become visible, because the same interpreter runs thousands of tasks. Pure-Python leaks are usually fixable — a module-level cache with no eviction, a growing registry, a reference cycle holding large objects. Leaks inside a compiled extension often are not: the allocation happens outside the Python heap, so it does not show up in Python-level allocation tracing, and you have no way to release it short of ending the process. Recycling the process ends it for you. This is containment, and the right way to talk about it is as a stopgap with the real defect written down. **Per-process cached state going stale.** Consider an ad-auction bidder whose scoring workers each memoize a price table on first use to avoid re-reading it per task. That is a sensible optimization while the table is fresh. But when the table is refreshed on disk and the workers are hours old, every worker is serving a stale cached value, and the batch produces confidently wrong bids with no error anywhere. Nothing crashes; the numbers are simply out of date. `maxtasksperchild` bounds the age of that state: with N tasks per worker, no cached value survives longer than N tasks' worth of work. It is a blunt instrument compared with a real invalidation signal, and it should be named as such — but on a batch pipeline with no back-channel to the workers, it is often the pragmatic answer. ### The cost, and how to size N Every recycle pays for a process start. Under `spawn` that means a brand-new interpreter that re-imports your modules; under `forkserver` a fresh fork of the pre-imported server process, which is cheaper but not free; under `fork` it is cheapest but `fork` is no longer the default on Unix other than macOS as of Python 3.14, and is unsafe in a parent with threads. If your task takes 50 milliseconds and a worker start takes 300, then `maxtasksperchild=1` means you spend most of the run starting interpreters. If a task takes 30 seconds, a recycle every task is nearly free. So size N from the ratio: pick the largest N that still keeps the accumulated growth or the staleness inside the tolerance you care about, and check that the amortized start cost is a small fraction of task time. Measure the resident memory of a worker as a function of tasks completed and pick N below the point where the machine gets uncomfortable — leaving headroom, because all pool workers do not recycle in lockstep and the peak matters more than the average. A 27-minute batch that recycles a handful of times is invisible; the same batch recycling every task might not finish at all. ### What it does not do It does not rescue a worker that has already been killed: the exit it schedules is voluntary, so a process taken out by an out-of-memory kill or a segfault is a different failure entirely, with its own recovery story. It does not help with a single task whose peak footprint alone exhausts memory — that needs the task fixed or routed elsewhere, since even a fresh worker will die on it. It does not make a leak smaller, only shorter-lived. And it does not bound *parent* memory: if the parent is the one accumulating results, no amount of worker rotation helps. ### The judgement being tested An interviewer asking this wants to hear that you know it is a mitigation. The strong answer names the two legitimate uses, prices the recycle against task duration, sizes N from measurements rather than a round number, and says out loud that a leak inside your own code should be fixed rather than recycled around — while a leak inside a third-party native library, or a cached value with no invalidation channel, is exactly what this knob exists for.

  • How would you choose the value of N rather than guessing?
    Measure a worker's resident memory as a function of tasks completed, and measure a cold worker start under the start method you actually use. Pick the largest N whose accumulated growth stays inside your headroom, then check that the start cost amortized over N tasks is a small fraction of task time. Leave slack because workers do not recycle in lockstep, so the peak is what the machine has to survive.
  • Does maxtasksperchild help when a worker is killed by the OOM killer?
    No. The recycle it schedules is a voluntary, orderly exit after a completed task; a worker killed from outside never gets there. It can *prevent* some of those kills by capping slow growth before it reaches the limit, but once a task's own peak footprint is what exhausts memory, a fresh worker dies just as fast. That case needs the task fixed, the input routed away, or the pool sized by memory instead of cores.
  • Why might recycling be the wrong answer to a memory leak?
    Because it hides the defect while keeping the cost. If the leak is in your own Python code — an unbounded cache, a growing registry, a cycle holding large objects — it is usually cheap to find and fix, and recycling instead leaves a knob nobody understands two years later. Recycling earns its place when the leak is in a compiled dependency you do not control, and even then the real defect belongs in the tracker.

It is the rental-car policy of retiring a vehicle at a fixed mileage: you are not repairing the wear, you are capping how much of it any one customer inherits.

saying these in an interview costs you the question

  • Setting it to 1 by default without pricing the process start
  • Calling it a fix for a memory leak rather than containment
  • Expecting it to save a worker already killed from outside
  • Ignoring that a fresh worker must re-import and re-warm
  • Assuming it bounds the parent process's memory too

context