How much memory can itertools.tee hold when one forked iterator runs far ahead of another?
answer
- Nothing is copied up front
- The forks share one source iterator
- Cost follows the lead between branches
- Draining one fork buffers everything
- Two full passes: materialise instead
basics
~20 sitertools.tee holds every item until the slowest of its forked iterators has consumed it. If one fork is drained before the other starts, the buffer grows to the entire stream, and a plain list is then cheaper and clearer.
solid answer
~40 s`itertools.tee(source, n)` copies nothing up front: it returns n iterators sharing one internal buffer over a single underlying iterator. Every item pulled from the source is retained until all n branches have moved past it, so peak auxiliary storage tracks the **largest lead** between the fastest and slowest branch, not any fixed constant. Branches that advance in lockstep — the one-item lookahead that `itertools.pairwise` performs in C, for example — cost almost nothing. A pipeline that drains one branch fully before touching the second buffers the whole stream, paying list-sized memory plus per-item bookkeeping. Two operating rules follow: if you will traverse the data twice from start to finish, call `list()` instead of `tee`, and once an iterator has been teed, never advance the original — anything you pull yourself disappears from every branch.
code
python · 8 linesimport itertools
records = iter([("a", 1.5), ("b", 2.5), ("c", 3.5)])
left, right = itertools.tee(records, 2)
print(list(left))
print(list(right))
print(list(records))go deeper
Know what itertools.tee is for: it splits one iterator into several that each see the same items, which a bare iterator cannot offer because consuming it once leaves it empty.
Explain the shared buffer: an item is held until every branch has passed it, so the cost tracks how far apart the branches drift, and the original iterator must not be advanced afterwards.
Demonstrate the operating judgement — spot the pipeline whose branches drift, predict the memory before it shows up as climbing RSS, and choose list(), a second read of the source, or one fused loop instead.
Own the pattern-level call: when a pipeline should stream at all, when materialising is the honest answer, and when two consumers at different speeds should be decoupled by a bounded queue with real back-pressure rather than an invisible in-process buffer.
### What tee returns `itertools.tee(iterable, n=2)` returns a tuple of `n` independent iterators that each yield the same items in the same order. That is genuinely useful, because a plain iterator is one-shot: consuming it to answer one question leaves nothing for the second question. tee lets two consumers each believe they have their own stream. The important part is what tee does **not** do. It does not read the source ahead of time, and it does not build `n` copies of the data. It wraps the single underlying iterator in a shared buffer. When any branch asks for an item that has not been fetched yet, tee pulls one item from the source, appends it to the buffer, and hands it out. When a branch asks for an item that another branch has already caused to be fetched, it is served from the buffer. An item can only be discarded once **every** branch has moved past it. ### The cost model, stated precisely Peak auxiliary storage is proportional to the maximum gap, measured in items, between the furthest-ahead branch and the furthest-behind one, over the whole life of the tee. Not the stream length; not a constant. That single sentence answers most interview follow-ups: - Branches consumed in lockstep, one item apart: the buffer holds one or two items. Free. - Branches that drift by a few thousand items: the buffer holds a few thousand references. Usually fine. - One branch drained completely, then the other started: the buffer holds the entire stream. Worse than a list — you pay for the same object references a list would hold, plus the linked chunks and per-branch position bookkeeping tee maintains, plus a Python-level pull per item, and you still cannot index or re-iterate. ### Where this actually bites Picture a geocoding batch owned by an 11-person team: a generator streams address records out of a large export, and someone tees it so one branch can compute summary statistics while the other writes the enriched rows out. The statistics branch is a tight `sum`-style loop and finishes in seconds; the writer branch does network-bound work and lags by the whole file. Nothing is obviously wrong in the code, no exception is ever raised, and the process quietly holds every record in the tee buffer for the duration of the run. It shows up as steadily climbing RSS with the number of live objects growing in lockstep, and the misleading part is that the generator was introduced *to keep memory flat*. The fix is to stop teeing. If the batch fits in memory, `rows = list(source)` and iterate it twice — explicit, indexable, and no worse in memory than the buffer was. If it does not fit, either read the source twice (re-open the export), or fuse the two consumers into one loop that hands each record to both and then drops it. When the consumers genuinely must run concurrently, decouple them with a bounded `queue.Queue`, whose maximum size gives you back-pressure and a hard memory ceiling that a tee buffer never had. ### The other rule: do not touch the original Both branches read from the same underlying iterator, so any item you pull from that original yourself is consumed and never reaches either branch. The documented contract is that once an iterable has been passed to tee, the original should not be used elsewhere. If some third consumer needs the stream too, ask tee for three branches rather than mixing paradigms. ### Where tee is exactly right Lookahead and windowing: fork the iterator, advance one branch by a step, and zip the two together so each yields the current and the following item. The branches stay a fixed distance apart, so the buffer never grows past that distance. That pattern is common enough that Python 3.10 added `itertools.pairwise`, which does the adjacent-pair case in C without the tee bookkeeping — prefer it for pairs and keep the tee construction for wider or asymmetric windows. tee is also right when two consumers are interleaved on purpose — both driven inside the same loop, or both feeding a merge — because interleaving keeps the lead small by construction. ### How to answer this in an interview Say what tee is, then immediately give the cost model in terms of the lead between branches, then name the two failure shapes: draining one branch first (unbounded buffer) and touching the original (silently missing items). Finishing with "if I need two full passes I materialise a list, and if the data cannot be materialised I read the source twice or fuse the consumers" is what separates a description from operating judgement.
- What happens to the forked iterators if you keep advancing the original after calling itertools.tee?The branches share that original iterator, so every item you pull yourself is consumed and simply never appears in any branch — silently, with no error. The documented contract is that the original must not be used elsewhere once it has been teed. If a third consumer needs the data, ask tee for a third branch instead.
- Two consumers of one large batch run at wildly different speeds. How do you cap the memory?Do not tee. Materialise once with `list()` if the batch fits, re-read the source a second time if it does not, or fuse the consumers into a single loop that hands each record to both and then drops it. If they must run concurrently, put a bounded `queue.Queue` between them so the fast side blocks and back-pressure caps the buffer.
- Why can a fully drifted tee cost more than a list holding the same items?The buffer stores the same object references a list would, and adds the chunked internal storage plus per-branch position bookkeeping. You also keep the per-item Python-level pull cost and lose indexing, slicing and `len()`. So you pay list-scale memory for iterator-scale ergonomics — strictly the worse trade for a full second pass.
tee is two readers sharing one scroll being copied line by line: the scribe cannot discard a line until the slower reader has passed it, so a reader who sprints to the end forces the whole scroll to be kept.
saying these in an interview costs you the question
- Says tee eagerly copies the data into n separate lists
- Claims tee is lazy so it costs no extra memory
- Keeps consuming the original iterator after teeing it
- Uses tee to make an exhausted stream restartable
- Assumes the tee buffer has a fixed maximum size
- Prefers tee over list() for two full passes