How do Executor.map and concurrent.futures.as_completed differ in ordering and error timing?
answer
- one preserves input order, one does not
- a slow first item stalls one of them
- which one hands you the Future itself
- the first error stops the ordered stream
- wait() returns done and not_done sets
basics
~20 sExecutor.map yields results in input order, so one slow early job holds back everything behind it and its exception surfaces only when iteration reaches that position. as_completed yields Futures in completion order, so whatever finishes or fails first reaches you first.
solid answer
~50 s`Executor.map(fn, iterable)` submits the work and returns a lazy iterator that yields results **in the order of the inputs**, re-raising a job's exception at the moment iteration reaches that input's slot — so a failure in item 0 blocks you from ever seeing items 1..n, and a slow item 0 stalls the whole stream even though later results are already sitting there. `concurrent.futures.as_completed(futures)` takes futures you submitted yourself and yields each one **as it finishes**, in completion order, so a fast failure surfaces as soon as you call `result()` on the yielded `Future`. Use `map` when you want ordered, positional results and simple code; use `submit` plus `as_completed` when you want streaming progress, per-item error handling, or the ability to react to the first failure. `concurrent.futures.wait` is the third option: it blocks and returns `(done, not_done)` sets and never raises the jobs' exceptions itself.
code
python · 11 linesimport time
from concurrent.futures import ThreadPoolExecutor, as_completed
def work(n):
time.sleep(0.3 - n * 0.1)
return n
with ThreadPoolExecutor(max_workers=3) as pool:
print("map:", list(pool.map(work, [0, 1, 2])))
futures = [pool.submit(work, n) for n in (0, 1, 2)]
print("as_completed:", [f.result() for f in as_completed(futures)])go deeper
Know that Executor.map looks like the built-in map and keeps input order, while as_completed gives you jobs as they finish. Be able to say which one you would reach for to print progress.
Explain the mechanics: map returns a lazy generator over per-input futures and re-raises at the failing position, as_completed yields the Future so you handle each failure yourself, and wait returns done/not-done sets without raising.
Demonstrate the operational choice on a long batch: streaming progress, per-item failure accounting, correlating futures back to inputs, and using FIRST_EXCEPTION plus a shutdown for fail-fast. Mention head-of-line blocking as a real latency bug.
Own the contract the batch exposes: ordered-results-or-abort versus stream-and-report. That choice drives retry policy, observability and whether a partially failed run is restartable, and it should be uniform across the codebase's batch jobs.
Both `Executor.map` and `concurrent.futures.as_completed` fan work out over the same pool. They differ in the order results arrive, in when an error hits you, and in how much you can do per item. ### `Executor.map` `pool.map(fn, iterable)` submits jobs and returns a **generator**, not a list. Under the hood it creates one `Future` per input and the generator pops them off the front in input order, calling `result()` on each. Three consequences follow directly: 1. **Ordered output.** Result *i* always corresponds to input *i*, regardless of which job actually finished first. That is exactly what you want when the results feed a positional structure. 2. **Head-of-line blocking.** If input 0 takes ten seconds and inputs 1..99 take ten milliseconds each, you get nothing for ten seconds even though 99 results are already computed and waiting in memory. 3. **Errors stop the stream.** When the generator reaches a `Future` that failed, `result()` re-raises there, and the generator's cleanup cancels the futures that have not started. You cannot see the results *after* the failing item, and you do not learn how many other items failed — only the first one in input order. `map` also has two parameters worth knowing. `chunksize` batches inputs into groups so a `ProcessPoolExecutor` pays the inter-process round trip once per batch rather than once per item; `ThreadPoolExecutor` ignores it entirely, since there is no serialization to amortize. And `buffersize`, **new in Python 3.14**, caps how far ahead `map` consumes its input iterable: before 3.14, `map` drained the whole iterable and submitted every job up front, which meant an infinite or huge input either hung or blew up memory. With `buffersize=n` it keeps roughly `n` items in flight and pulls more only as results are consumed, giving you backpressure. ### `as_completed` `as_completed(futures, timeout=None)` takes an iterable of futures you already submitted and yields them **in the order they finish**. Futures that are already done come out first, then the rest as they complete. It yields the `Future`, not the value, so you still call `result()` or `exception()` yourself — which is the whole point: you get to handle each item's failure individually and keep going. This is the shape you want for progress reporting (`for i, fut in enumerate(as_completed(futures), 1): ...`), for a first-result-wins race, and for collecting a full failure report instead of stopping at the first error. Its cost is that you must keep every `Future` alive to pass in, and to know *which* input a yielded `Future` belongs to you need a side mapping — the common idiom is a dict `{pool.submit(fn, item): item for item in items}`, then look up `futures[fut]` on the way out. The `timeout` argument applies to the whole iteration, not per future, and raises `TimeoutError` when it expires. ### `wait` `concurrent.futures.wait(futures, timeout=None, return_when=ALL_COMPLETED)` is the blocking, non-streaming sibling. It returns a named tuple of two sets, `done` and `not_done`, and never raises a job's exception itself. `return_when` accepts `FIRST_COMPLETED`, `FIRST_EXCEPTION` (return as soon as any job raises, or when all finish if none does) and `ALL_COMPLETED`. `FIRST_EXCEPTION` is the natural building block for fail-fast batch jobs: wait for it, then shut the pool down. ### Memory, correlation and cancellation on failure The two shapes also differ in what you must hold. `as_completed` needs the full list of futures in memory for the whole run, and a `Future` carries no reference to the arguments it was called with, so correlating a result back to its input needs a side mapping you build at submit time. `map` needs no such bookkeeping — position *is* the correlation — and on 3.14 with `buffersize` it does not even need the input list materialized. There is one behaviour people miss: when the `map` generator is closed, whether by an exception or by garbage collection, its cleanup cancels the futures that have not started yet. Breaking out of a `for` loop over `pool.map(...)` therefore drops the remaining queued work, while breaking out of a loop over `as_completed` leaves every outstanding job running. ### Choosing A useful decision rule: if you would have written a plain list comprehension over `map()` in single-threaded code and you want the same shape back, use `Executor.map`. The moment you need per-item error handling, progress as work lands, correlation back to the input, or a reaction to the first failure, drop to `submit` plus `as_completed` or `wait`. Rewriting `map` into `submit` + `as_completed` is a two-line change, so this is not a decision to agonize over up front. A final gotcha that shows up in reviews: `pool.map` is lazy, so calling it and never iterating the result does almost nothing visible — the work is submitted, but the exceptions stay unread. Wrapping it in `list(...)` inside the `with` block is what actually forces both the results and the errors out.
- With as_completed, how do you tell which input a yielded Future came from?Build the mapping when you submit: `futures = {pool.submit(fn, item): item for item in items}`, then `futures[fut]` on the way out. A `Future` carries no reference to its arguments, so without that dict you can only recover the input by wrapping the callable to return the input alongside the result.
- What does the chunksize argument to Executor.map actually change?It batches inputs so each unit of work sent to a worker covers several items. On a `ProcessPoolExecutor` that amortizes the pickling and pipe round trip, which can be a large win for many small items. On a `ThreadPoolExecutor` it is ignored, because threads share memory and there is nothing to serialize.
- How would you stop a batch as soon as any job raises, rather than at the first input position?Submit the jobs yourself and call `concurrent.futures.wait(futures, return_when=FIRST_EXCEPTION)`. It returns `(done, not_done)` the moment any job raises, without raising itself, and you then call `shutdown(cancel_futures=True)` to drop what is still queued. `Executor.map` cannot do this: it only notices failures in input order.
saying these in an interview costs you the question
- Says Executor.map returns results in completion order
- Thinks Executor.map returns a list rather than a lazy iterator
- Believes as_completed yields results instead of Futures
- Assumes map reports every failure, not just the first reached
- Expects chunksize to speed up a ThreadPoolExecutor
- Calls pool.map and never iterates it, then reports success