How does itertools.batched chunk a lazy record iterator, and what happens at the tail?
answer
- A stdlib chunker, not a hand-rolled one
- New in the 3.12 release
- Tuples out, and one may be short
- A keyword makes the short one an error
- The old zip trick drops the remainder
basics
~10 sitertools.batched, added in 3.12, pulls lazily from an iterable and yields tuples of length n; the final tuple may be shorter. Since 3.13, passing strict=True makes an incomplete final batch raise ValueError instead.
solid answer
~40 s`itertools.batched(iterable, n)` arrived in 3.12. It consumes the source lazily - just enough to fill each batch - and yields **tuples** of exactly `n` items, except the last, which holds whatever was left over. `n` must be at least one or it raises `ValueError` at call time. In 3.13 a keyword-only `strict` was added: `batched(src, n, strict=True)` raises `ValueError` if the final batch is short, which is what you want when a downstream call requires full-size groups. The tail is the whole story - the old `zip(*[iter(src)]*n)` idiom looks equivalent but silently *drops* the remainder, because zip stops at its shortest argument. Before 3.12 the correct idiom was a generator looping `list(itertools.islice(it, n))` until it came back empty. Never call `list()` on the source first: that defeats the point of batching a stream.
code
python · 8 linesfrom itertools import batched
records = (f"rec-{i}" for i in range(7))
for chunk in batched(records, 3):
print(chunk)
# ('rec-0', 'rec-1', 'rec-2')
# ('rec-3', 'rec-4', 'rec-5')
# ('rec-6',)go deeper
Know that itertools.batched exists from 3.12, that it groups any iterable into tuples of a given size, and that the last tuple can be smaller than the rest.
Explain the laziness - only one batch is held at a time - and the exact tail rule. Be ready to write the pre-3.12 islice-based chunker and to say why n must be at least one.
Diagnose the tail as the risk in a streaming pipeline: the dropped remainder from the zip idiom, a swallowed per-batch exception that turns a truncated run into a green one, and the reconciliation count that would have caught either. Say when strict=True is the correct contract.
Own the batch size as a capacity decision: per-call overhead against memory, retry blast radius and downstream limits, plus whether partial batches are legal in your data contract at all. Decide where that invariant is enforced so it is not re-litigated in every pipeline.
### The shape of the problem Batching shows up whenever a per-item cost is dominated by a per-call cost: sending records to a bulk endpoint, writing rows in transactions, or feeding a scoring service that accepts groups. Take a genome-annotation pipeline that streams variant records out of a lazy reader and hands them to an annotation call that takes up to 500 at a time. The source must stay lazy - the file is far larger than memory - so the chunker has to pull just enough for one batch and no more. ### itertools.batched `itertools.batched(iterable, n)`, added in **3.12**, is exactly that. Its contract: * It yields **tuples**, not lists. If a downstream call mutates the batch, convert it. * Every tuple has exactly `n` items **except possibly the last**, which carries the remainder. `batched(range(7), 3)` yields `(0, 1, 2)`, `(3, 4, 5)`, `(6,)`. * It is lazy in both directions: the source is consumed just enough to fill the batch being yielded, and nothing is buffered beyond one batch. * `n` must be at least one; `batched(src, 0)` raises `ValueError: n must be at least one` immediately, at call time rather than during iteration. * An empty source yields no batches at all - not one empty tuple. In **3.13** a keyword-only `strict` parameter was added. `batched(src, n, strict=True)` raises `ValueError: batched(): incomplete batch` when the final group is short. That is the right setting when a partial batch is a data error rather than a normal tail - fixed-width framing, a protocol that requires full groups, a paired-record format where an odd count means the input is truncated. It fails during iteration, when the tail is reached, not at call time. ### The idiom it replaced, and why that idiom was dangerous The classic pre-3.12 one-liner is `zip(*[iter(src)] * n)`. It works by putting the *same* iterator object into the argument list `n` times, so each round of `zip` pulls `n` consecutive items. It is clever, and it is a silent data-loss bug: `zip` stops at its shortest argument, so when the source runs out mid-group the partial group is thrown away. Over the seven-record example it yields `(0, 1, 2)` and `(3, 4, 5)` and quietly loses the `6`. On a pipeline with an arbitrary record count you lose up to `n - 1` records on every run and nothing anywhere reports it. The safe pre-3.12 idiom is an explicit generator: ```python def chunks(iterable, n): it = iter(iterable) while batch := list(islice(it, n)): yield batch ``` This relies on the property that successive `islice` calls over the same iterator continue where the previous one stopped, and that an empty list is falsy, which ends the loop. It yields lists rather than tuples. Write this only when you must support interpreters older than 3.12; on 3.14, `batched` is faster and clearer. ### The failure mode to watch for in production The tail interacts badly with over-broad error handling. Wrap the per-batch downstream call in `try: ... except Exception: pass` and a failure on a batch is swallowed: the loop keeps going, the run exits zero, and the pipeline reports success with records missing. Because the loss is at the boundary rather than at the start, a short smoke test frequently passes while a full run is wrong - and if verifying that costs a 27-minute end-to-end suite, nobody reruns it on a change that "only touched batching". Two habits defuse this: count what you send and assert the total against the source count at the end, and let batch errors propagate, or record them per-batch, rather than catching bare `Exception`. When a short final batch is genuinely invalid, `strict=True` turns a silent truncation into a loud `ValueError`. ### Choosing n `n` is a throughput knob, not a constant to guess once. It trades per-call overhead against per-batch memory and against the blast radius of a retry: a bigger batch amortises the round trip but holds more records in memory and, when it fails, redoes more work. Note also that a batch is materialised as a tuple, so peak memory is roughly `n` records plus whatever the downstream call builds from them. ### What an interviewer wants Name `batched`, date it to 3.12, state that the last tuple may be short, know that `strict=True` (3.13) turns that into an error, and identify the tail as the risk - including why the `zip(*[iter(x)]*n)` trick loses data. Volunteer that you would never materialise the source with `list()` first, because streaming was the point.
- Why does zip(*[iter(records)] * 3) lose data when the record count is not a multiple of three?The list holds three references to the *same* iterator, so each round of zip pulls three consecutive items - that part works. But zip stops as soon as any argument is exhausted, and it discards the partial row it had already pulled. When the source runs out mid-group, up to n-1 records disappear with no error and no warning. itertools.batched yields that remainder as a short tuple instead.
- When would you pass strict=True to itertools.batched rather than handling a short final batch?When a partial group means the input is wrong rather than merely finished: fixed-width framing, a protocol requiring full groups, or paired records where an odd count signals truncation. strict=True (3.13) raises ValueError at the tail so the run fails loudly. If a short final batch is normal - the last page of records - leave it off and let the downstream call accept a smaller group.
- How did you chunk a one-shot iterator before 3.12?With a generator that calls itertools.islice repeatedly against the same iterator: `while batch := list(islice(it, n)): yield batch`. Successive islice calls continue where the previous stopped, and the empty list at the end is falsy, which terminates the loop. It yields lists rather than tuples. It is still the right code when you must support 3.10 or 3.11, but on 3.12 and later batched is faster and clearer.
saying these in an interview costs you the question
- Assumes every batch has exactly n items
- Uses zip(*[iter(x)]*n) without noticing the dropped tail
- Calls list() on the source before batching a stream
- Claims itertools.batched exists on 3.10 or 3.11
- Says batched yields lists rather than tuples
- Swallows a per-batch exception and reports a clean run