When should you replay a one-shot iterator with itertools.tee() rather than list()?
answer
- It depends how the two passes interleave
- Cost tracks the gap, not the length
- Lock-step is cheap, sequential is not
- One branch drained first means use list()
- Never reuse the original after splitting
basics
~20 sUse itertools.tee() only when the branches are consumed roughly in step, so its internal buffer stays small. If one branch is drained before the other starts, tee buffers the entire stream anyway and list() is simpler and faster.
solid answer
~50 s`itertools.tee(source, n)` returns n independent iterators over one source, and it works by buffering every item the fastest branch has taken but the slowest has not yet seen. So the cost tracks the **gap between consumers**, not the length of the stream: nearly free for lock-step consumption, and no cheaper than a list when one branch runs to the end first — CPython's own docs say to prefer `list()` in that case. The decision rule is therefore about the consumption pattern. Lock-step over a huge or unbounded source → `tee`. Two full sequential passes over bounded data → `list()`. Source cheap to re-read → just read it again and hold nothing. Two rules come with `tee`: never touch the original iterator after splitting, because the branches own it now, and never share branches across threads, since the buffer is not thread-safe.
code
python · 7 linesimport itertools
source = iter(range(1_000_000))
left, right = itertools.tee(source)
next(left)
print(list(itertools.islice(left, 3)))
print(list(itertools.islice(right, 3)))go deeper
Recall that a one-shot iterator cannot be looped twice, and that the straightforward fix is to build a list from it once and then loop over that list as many times as you need.
Explain what itertools.tee actually does: several branches over one source, with an internal buffer holding items the fastest branch has read and the slowest has not. Then say why that makes two full sequential passes no cheaper than a list.
Show the judgement on a real pipeline: characterise the consumption pattern, choose between re-reading, tee and materializing, and state the memory bound your choice implies. Mention the original-iterator rule and the thread-safety caveat unprompted.
Own the pipeline contract: whether stages may demand multiple passes at all, where the replay boundary sits, and what it is budgeted to cost. Argue when spilling or re-querying beats holding a dataset in process memory across a whole fleet.
### What tee actually does `itertools.tee(iterable, n=2)` returns `n` independent iterators that all yield the items of one underlying source. It cannot conjure the items twice from a one-shot source, so it does the only thing available: it **buffers**. Internally it keeps the items that the *fastest* branch has already pulled but the *slowest* branch has not yet seen. That single implementation detail decides every question about when to use it. The memory cost is proportional to the **gap between the consumers**, not to the length of the stream. - Branches read almost in step → the buffer holds a handful of items → `tee` is nearly free, even over an enormous or unbounded source. - One branch is drained completely before the other starts → the buffer necessarily grows to the entire stream → you have paid for a list, plus per-item bookkeeping, and got a worse API. CPython's own documentation says as much: when one iterator uses most or all of the data before another starts, `list()` is faster. ```python import itertools source = iter(range(1_000_000)) left, right = itertools.tee(source) next(left) # one item now buffered for right print(list(itertools.islice(left, 3))) # [1, 2, 3] print(list(itertools.islice(right, 3))) # [0, 1, 2] ``` ### The decision, stated as a rule Ask *how the two passes interleave*, then choose: 1. **Lock-step or near-lock-step, and the source is huge, unbounded or expensive to reproduce** → `tee`. Comparing each item with its neighbour, feeding two aggregators that advance together, splitting a stream between a checksum and a writer. 2. **Two full sequential passes over a bounded stream** → `list()` (or `tuple()`). One obvious allocation, no hidden buffer, faster iteration, and the result is re-iterable by every caller thereafter. 3. **The source is cheap to re-read** → re-read it. If the stream comes from a cached fetch running an 83% cache-hit rate, a second traversal costs almost nothing and holds no memory at all. This third option is the one candidates most often forget. 4. **Unbounded and not lock-step** → neither works; redesign so one pass suffices, or bound the replay explicitly with `itertools.islice` and fail loudly on overflow rather than growing until the process is killed. ### A worked example A search-index rebuilder reads a stream of documents and needs two things from it: the newest timestamp in the batch, so it can detect a clock-skew artefact from a machine whose clock ran ahead, and the filtered set of documents to write. Those two needs interleave differently depending on how you frame them. Framed as "compute the maximum, then filter against it", the first pass must finish before the second begins — strictly sequential, so `tee` would buffer the whole batch and `list()` is the honest choice. Framed as "compare each document against its predecessor to flag out-of-order timestamps", the two passes advance together one item apart, and `tee` costs a single buffered item over a stream of any size. Same data, same source, opposite answers — which is why the interview question is about the consumption pattern, not about the function. For the lock-step framing specifically, `itertools.pairwise` (added in Python 3.10) does the job directly and is clearer than hand-rolling it with `tee`. ### The two rules that come with tee **Do not use the original iterator after splitting it.** The branches now own the source and pull from it on demand. Anything you consume directly is simply missing from every branch, with no error to tell you. If you need the raw stream too, make it one of the branches. **Do not share branches across threads.** The branches share one buffer and one source with no locking; concurrent `next()` calls on different branches can raise `RuntimeError` or lose items. The GIL does not save you here, because a single `next()` is not one atomic step. Give each thread its own materialized copy, or serialize access behind a lock you own. ### The judgement an interviewer is listening for The weak answer is "`tee` copies the iterator, so use it whenever you need two passes". The strong answer names the buffer, ties its size to the gap between consumers, and then reaches for the cheapest of re-read / `tee` / materialize based on the actual pattern — while stating the memory bound the choice implies. In a pipeline that anyone else will extend, the deeper point is a contract question: decide whether stages are permitted to demand multiple passes at all, and if they are, where the replay boundary lives and what it is allowed to cost.
- What happens if you keep using the original iterator after calling itertools.tee() on it?The branches lose whatever you take. `tee` pulls from the source on demand, so any item consumed directly is simply missing from every branch, and nothing reports it. The documented rule is that the original must not be used elsewhere once split. If you genuinely need the raw stream too, make it one of the branches instead of reaching past them.
- How do you bound the memory risk when the stream might not fit in memory?Decide the shape first. If both passes can advance in step, `tee` costs only the gap; if they cannot, re-read the source or spill the first pass to a temporary file rather than a list. Where a hard bound is required, cap the replay with `itertools.islice` and fail loudly on overflow instead of letting a materialized pass grow until the process is killed.
- Is it safe to hand two itertools.tee() branches to two threads?No. The branches share one buffer and one source with no locking, so concurrent `next()` calls on different branches can raise `RuntimeError` or lose items, and the GIL does not help because a single `next()` is not one atomic step. Give each thread its own materialized copy, or serialize access behind a lock you own.
saying these in an interview costs you the question
- Says tee() is always cheaper than building a list
- Thinks tee() copies items with no buffering at all
- Keeps consuming the original iterator after tee()
- Calls tee() on an unbounded stream then drains one branch
- Assumes tee() branches are safe to share across threads
- Never considers simply re-reading a cheap source