How do multiprocessing.Pool's map, imap and imap_unordered differ?
answer
- One returns a container, two return iterators
- Think about when the first result arrives
- One of the three makes no ordering promise
- A slow first item punishes exactly one of them
- Memory profile differs more than wall time does
basics
~20 smap() blocks and returns a list of every result in input order. imap() returns a lazy iterator that yields results in input order as they arrive. imap_unordered() yields each result the moment it is ready, in completion order.
solid answer
~50 sAll three fan the same work across the same worker processes; they differ in *when* results reach you and how much memory that costs. `map()` materialises the input (a generator is passed through `list()` first) and the whole output list, then blocks until the last item is done — O(n) memory and no result until everything finishes. `imap()` streams: you can start consuming while workers are still busy, but because it must preserve input order, one slow early item stalls delivery of everything behind it, and those finished results buffer in the parent meanwhile. `imap_unordered()` drops the ordering promise, so results arrive as they complete and nothing buffers waiting for a straggler. `map_async()` returns an `AsyncResult` immediately instead of blocking. Ordering is the tradeoff, not throughput: total wall time is roughly the same in all three.
code
python · 14 linesimport multiprocessing as mp
import time
def slow_first(n):
time.sleep(1.0 if n == 0 else 0.05)
return n
if __name__ == "__main__":
with mp.Pool(4) as pool:
print("imap :", list(pool.imap(slow_first, range(4))))
with mp.Pool(4) as pool:
print("imap_unordered:", list(pool.imap_unordered(slow_first, range(4))))go deeper
Recall the shapes: one call gives you a finished list in input order, the other two give you an iterator you consume as results arrive. Know that only one of the three drops ordering.
Explain the mechanics behind the contract: the up-front list() conversion, the reorder buffer that ordered streaming needs, and why chunk-size defaults differ between the blocking and streaming forms.
Demonstrate the judgement in production: pick the form by memory profile and time-to-first-result, carry identity in results so unordered output stays usable, and know how a worker exception surfaces mid-iteration.
Frame the choice as a pipeline-design decision. Streaming with unordered delivery lets downstream stages overlap and bounds memory; ordered batch delivery is simpler but couples the whole batch's latency to its slowest item.
The four fan-out methods on `multiprocessing.Pool` do the same underlying thing — push work items onto an internal task queue that idle worker processes pull from — and differ only in the *result-delivery contract* they offer the caller. Interviewers ask this because choosing the wrong one is a memory bug or a latency bug that never shows up in a small test. ## The four methods ### `map(func, iterable, chunksize=None)` `map` is the **blocking, ordered, all-at-once** form. Two things happen up front that people forget: 1. First, if the iterable has no `__len__` it is converted with `list()`, because the pool needs the length to compute a chunk size — so `map()` over an unbounded generator does not stream, it hangs while trying to materialise infinity. 2. Second, results are accumulated into one list that is returned only when the final item is done. Peak memory therefore holds the whole input *and* the whole output. In exchange you get the simplest possible call site and results indexed exactly like the input. ### `imap(func, iterable, chunksize=1)` `imap` returns an `IMapIterator` immediately. Workers begin consuming right away, and you can start processing result 0 while result 500 is still being computed. The important subtlety is that ***ordered* streaming is not free**: the iterator must hand you item 0 before item 1, so if item 0 takes a minute and items 1-99 take a millisecond each, those ninety-nine finished results sit in a buffer in the parent until item 0 lands. The workers were never blocked — only your consumption was. That buffer is exactly the memory `imap` was supposed to save you. ### `imap_unordered(func, iterable, chunksize=1)` `imap_unordered` gives up input order. Each result is yielded as soon as it arrives, so a straggler blocks nothing and no reordering buffer builds up. - This is the right default whenever the consumer does not care which input produced which output, or when each result carries its own identity — a common pattern is to have the worker return `(key, value)` so the caller can reassemble order itself if it ever needs to. - What it does **not** give you is a speedup: the same tasks run on the same workers, so total wall time is essentially unchanged. - What changes is time-to-first-result and peak memory. ### `map_async(func, iterable, chunksize=None, callback=None)` `map_async` is `map` without the block: it returns an `AsyncResult` straight away, letting the parent do other work and later call `.get()`, `.wait()` or `.ready()`. `apply_async` is the single-call-per-submission sibling, useful when the work items are heterogeneous rather than one function over one iterable. ## Chunk-size defaults differ, and that difference is deliberate - `map` computes a chunk size from the input length and the worker count — roughly four chunks per worker — because it already knows the total and is optimising for **throughput**. - `imap` and `imap_unordered` default to `chunksize=1` because they are optimising for **latency-to-first-result**; batching would make you wait for a whole chunk before anything is yielded. Raising `chunksize` on `imap` trades that latency back for fewer dispatches. ## Failure behaviour is shared If the worker raises, the exception is pickled back and re-raised in the parent: - at the `map()` call for `map`; - and at the iteration step that would have produced that item for the `imap` family. With `imap`, an exception on item 5 surfaces when you reach position 5 in the iteration, which is why partially consumed results can already have been processed by the time you learn something failed. Design for that: - either wrap the worker so it returns a result-or-error value, - or accept that the first failure aborts the loop. ## Abandoning an `imap` iterator does not cancel the work The tasks already queued keep running in the workers; you simply stop reading their results. Stopping early genuinely means shutting the pool down, not breaking out of the loop. ## A practical decision rule - If the input is small and you want a list, use `map`. - If the input is large or the consumer can work incrementally, use `imap_unordered` and carry identity in the result. - Reach for `imap` only when you truly need input order *and* streaming, and then remember you are paying for the reorder buffer.
- Does imap_unordered finish the whole batch faster than imap?No. The same tasks run on the same workers, so total wall time is essentially identical. What improves is time-to-first-usable-result and peak memory: `imap` has to hold finished out-of-order results in a parent-side buffer until the item ahead of them arrives, while `imap_unordered` hands each one over immediately. Choose it for latency and memory, never as a throughput trick.
- What happens if you call Pool.map on an infinite generator?It hangs and eventually exhausts memory. `map` needs the input length to compute a chunk size, so an iterable without `__len__` is passed through `list()` before any work is dispatched. The `imap` family does not do this — it pulls lazily — which is the real reason to prefer it for large or unbounded inputs.
- How does a worker exception reach the parent with imap?The exception is pickled back and re-raised in the parent at the iteration step that would have yielded that item — not at the call that created the iterator. So earlier results may already have been consumed and acted on before you learn item 5 failed. If you need all results regardless, have the worker return a success-or-error value instead of raising.
- If you break out of an imap loop early, does the remaining work stop?No. Tasks already dispatched to workers keep running; abandoning the iterator only stops you from reading their results, and the pool stays busy. Genuinely cancelling means shutting the pool down with terminate(), accepting that in-flight work is killed mid-execution.
map is waiting for the whole delivery van to be unloaded before you touch anything; imap is taking parcels off the belt but only in label order, so one stuck parcel piles the rest up beside you; imap_unordered is taking whichever parcel lands next.
saying these in an interview costs you the question
- Claiming imap_unordered is faster overall than imap
- Believing Pool.map returns an iterator rather than a list
- Calling Pool.map on an unbounded generator to save memory
- Assuming imap avoids all parent-side buffering
- Expecting imap_unordered to keep input order when tasks take equal time
- Thinking breaking out of an imap loop cancels queued tasks