What does the chunksize argument to multiprocessing.Pool.map control?
answer
- It is a batching knob, not a worker count
- Trades fixed overhead against load balance
- The map default is derived, not fixed
- Roughly four batches per worker
- Watch the chunk count against the worker count
basics
~20 schunksize is how many items of the iterable are batched into a single task handed to one worker. Larger chunks mean fewer dispatches and less per-item overhead; smaller chunks spread work more evenly and shorten the idle tail.
solid answer
~50 sThe pool does not send one item per task. It slices the iterable into chunks of `chunksize` consecutive items, and each chunk is one unit of scheduling: one trip through the internal task queue, one hand-off to a worker, one result batch back. Bigger chunks amortise that fixed per-task cost over more items, which matters when each item is tiny. But a chunk is also the granularity of load balancing — a worker that grabs a chunk owns it to the end — so oversized chunks leave workers idle at the tail while one finishes. `Pool.map` with no chunksize computes `divmod(len(iterable), workers * 4)`, rounding up, so it aims for about four chunks per worker: enough batching to be cheap, enough chunks to rebalance. `imap` and `imap_unordered` default to 1 because they optimise for time-to-first-result instead.
code
python · 17 linesimport multiprocessing as mp
import time
def parse_batch(batch_id):
time.sleep(0.2)
return batch_id
if __name__ == "__main__":
for chunksize in (1, 8):
chunks = -(-96 // chunksize)
start = time.perf_counter()
with mp.Pool(11) as pool:
pool.map(parse_batch, range(96), chunksize=chunksize)
print(f"chunksize={chunksize} chunks={chunks} "
f"elapsed={time.perf_counter() - start:.1f}s")go deeper
Know that it batches several items into one unit of work for a worker, and that leaving it out is fine to start with because a sensible value is computed for you.
Explain both directions of the tradeoff and the derived default of roughly four chunks per worker, and be able to say why the streaming forms default to one item per chunk instead.
Show that you check the chunk count against the worker count, not just the chunk size, and that you diagnose a tail-latency regression that no per-call profile would reveal.
Own the tuning policy. Decide whether the team hard-codes a value, derives one from measured per-item cost, or leaves the default alone — and insist that whatever is chosen is measured on production-shaped input sizes.
`chunksize` is the pool's batching knob, and understanding it is understanding that a `Pool` has two costs: - the work itself, - and the fixed overhead of getting each task to a worker and its result back. ## What it actually does Given `pool.map(func, items, chunksize=k)`, the pool slices `items` into consecutive runs of `k` and puts each run on the internal task queue as one task. A worker pops a whole run, applies `func` to every element in it, and sends back a list of `k` results. `func` is still called once per item — chunking changes the *packaging*, never the semantics, and never the order of the list `map` returns. ## Why raising it helps Every task carries a **fixed cost** independent of the work: - queueing, - crossing the process boundary in both directions, - and the bookkeeping the pool does per result. If each item takes 20 microseconds of real work, that overhead dominates completely and the pool can be slower than a plain loop. Batching a thousand items into one task pays the fixed cost once instead of a thousand times. ## Why raising it hurts A chunk is the **unit of scheduling**. A worker that takes a chunk finishes it before taking another, so the pool cannot rebalance mid-chunk. Total wall time is set by the *last* chunk to finish, and that has a sharp boundary effect. Concretely: a clinical-lab result loader parses 96 batches of results, each batch costing about 0.2 seconds, across a pool of 11 worker processes. - With `chunksize=8` that is exactly 12 chunks — one more than the worker count. Eleven workers each take a chunk and run for 1.6 seconds; the twelfth chunk cannot start until one of them frees up, so it runs *after*, and the whole map takes about 3.3 seconds while ten of the eleven workers sit idle for the entire second half. - Drop to `chunksize=1` and the same work finishes in roughly 1.9 seconds. One extra chunk past the worker boundary nearly doubled the runtime — and nothing in the code looks wrong. That is the classic **off-by-one of pool tuning**: it is not the chunk *size* that bit, it is the chunk *count* landing just past a multiple of the worker count. ## The default is a compromise, and a good one `Pool.map` and `map_async` with `chunksize=None` compute `divmod(len(iterable), len(pool) * 4)` and add one if there is a remainder — roughly four chunks per worker. Four, rather than one, precisely to survive the case above: if items have uneven costs, a worker that draws an expensive chunk still has three chunks' worth of slack for its peers to absorb. It also explains why `map` insists on knowing the length: an iterable without `__len__` is passed through `list()` first, because there is no chunk size without a total. ## The streaming forms default differently `imap` and `imap_unordered` take `chunksize=1` by default because their contract is incremental delivery. Batching would mean waiting for an entire chunk to complete before the first result is yielded, which defeats the point. Raising `chunksize` on `imap` is a deliberate trade: - fewer dispatches, - more latency to the first item, - and coarser interleaving. ## How to choose in practice 1. Estimate the per-item work. - If it is milliseconds or less and the item count is large, batch — start near the default and raise it until the throughput curve flattens. - If items take seconds and their costs vary, keep chunks small so the scheduler can rebalance; `chunksize=1` is correct for a few dozen long, uneven jobs. 2. Then sanity-check the chunk *count* against the worker count: `ceil(len(items) / chunksize)` wants to be either well above the worker count or an exact multiple of it, never worker-count-plus-one. 3. And measure — the tail effect above is invisible in a profile of `func` itself, because no individual call got slower. ## What it is not - `chunksize` has nothing to do with how many worker processes exist; that is the pool's own constructor argument. - It does not reorder results, - does not change which function runs, - and does not bound memory on its own — a large chunk means a worker holds that many results at once before sending them back.
- Why does Pool.map aim for about four chunks per worker rather than one?One chunk per worker would be optimal only if every item cost the same. Four gives the scheduler slack: a worker that draws an expensive chunk still has other chunks left for its idle peers to pick up, so uneven item costs get absorbed instead of setting the wall time. It is a cheap hedge — four times the dispatch overhead, far better tail behaviour.
- When is chunksize=1 the right choice?When there are relatively few items, each item is expensive, and their costs vary. Then per-task overhead is negligible next to the work, and maximum scheduling granularity is exactly what you want so no worker is stuck holding a batch of slow items. It is also the default for the streaming forms, where the point is delivering the first result quickly.
- Does raising chunksize change the order of results from Pool.map?No. Chunking only changes how items are packaged into tasks; `map` reassembles the per-chunk result lists into one list indexed exactly like the input, whatever the chunk size. Ordering is a property of the method you called, not of the batching. `imap_unordered` is the only form that gives that guarantee up.
Handing out lab work by the tray instead of by the tube: fewer trips to the bench, but whoever draws the last tray works alone while everyone else stands idle.
saying these in an interview costs you the question
- Thinking chunksize sets the number of worker processes
- Claiming a larger chunksize is always faster
- Assuming Pool.map with no chunksize sends one item per task
- Believing chunksize can reorder the list map returns
- Tuning chunk size without checking the resulting chunk count
- Ignoring per-item cost when picking a value