skip to content

What does itertools.tee cost in memory when one of its branches runs far ahead?

level: seniorimportance: nice to knowfreq 22%

answer

  1. It remembers what one branch already passed
  2. The buffer is the gap between consumers
  3. Lockstep is cheap, sequential is not
  4. Drain one first and you have a list
  5. The original iterator is no longer yours

basics

~20 s

itertools.tee buffers every item the fastest branch has consumed but the slowest has not. If one branch is drained before the other starts, that buffer holds the entire stream, so tee costs as much as building a list and adds indirection.

solid answer

~50 s

`itertools.tee(iterable, n)` returns n independent iterators over one source. It can only do that by remembering: it pulls each item from the source once and keeps it in an internal buffer until every branch has seen it. The buffer therefore holds the *gap* between the leading and trailing branch. Consumed in lockstep, that gap is a couple of items and tee is nearly free. Drain one branch completely and then start the second, and the gap is the whole stream — at which point `list(iterable)` is simpler, faster and no more expensive, which is what the standard library documentation itself advises. Two further rules: do not touch the original iterator after teeing it, because tee has taken ownership and advancing it yourself makes branches skip items; and the returned iterators are not safe to consume from multiple threads without your own lock.

code

python · 7 lines
python
import itertools

source = iter(range(6))
left, right = itertools.tee(source, 2)

print(list(left))
print(list(right))

go deeper

for a junior

Know only that an iterator is one-shot and that iterating it a second time yields nothing. The named helper for splitting a stream is something you can look up when you first need it.

for a middle

Be able to say what tee returns and that it must buffer to do its job. Explaining that the buffer grows with the distance between the two consumers is the level-appropriate depth here.

for a senior

Show the judgement: recognise the drain-one-then-the-other pattern as a hidden materialization, prefer a single restructured pass, and know the ownership rule that consuming the original silently skips items in every branch.

for a principal

Frame it as a resource decision. Say when re-reading the source, buffering, or materializing is the right trade for a given data size and I/O cost, and how you would keep unbounded fan-out buffers out of long-running services.

## What tee actually does `itertools.tee(iterable, n=2)` answers the question "how do I iterate this one-shot stream more than once?" — and it answers it the only way possible for a source that cannot be rewound: by keeping a copy of anything not yet consumed by every consumer. Internally each returned iterator has a position in a shared, linked FIFO of already-fetched items. When a branch needs the next value and it is already in the structure, tee hands it over. When it is not, tee pulls one item from the underlying source, appends it, and hands it over. Items are dropped once every branch has moved past them. The consequence is one sentence long: **tee's memory is the distance between the furthest-ahead branch and the furthest-behind one.** ## The two regimes *Lockstep consumption* — a `zip` over both branches, or a loop that pulls one item from each per round — keeps the gap at a handful of items. Here tee is genuinely a streaming tool and costs almost nothing. *Sequential consumption* — `list(left)` and only then `list(right)` — forces the gap to the full length of the stream. Every item is retained until the second branch catches up, which happens only at the very end. You have materialized the entire source, but inside an opaque structure, with per-item overhead, instead of in a plain list you could measure with `len()` and index into. The standard library documentation states this directly: when one iterator will consume most or all of the data before another starts, `list()` is the faster choice. So the interview answer is not "tee is a memory leak" and not "tee lets you iterate twice for free". It is: tee converts a re-iteration requirement into a buffering cost proportional to how far apart the consumers drift. ## Ownership and the surprise After `left, right = itertools.tee(source)`, the name `source` must be considered dead. Tee assumes it is the only consumer; if you advance the original yourself, those items are never seen by either branch, and the branches silently skip values. This is a genuine production bug — a helper that tees a stream for a side channel and also keeps consuming the original produces gaps that look like data loss upstream. Relatedly, teeing a tee branch is fine, but each additional branch is another position in the same buffer, so the retained window is governed by the slowest of all of them. ## Thread safety The iterators returned by tee share mutable state and are not safe to consume from several threads at once; you must serialize access with your own lock. Under free threading (officially supported in 3.14, PEP 779) this stops being a theoretical concern, because two threads really can be inside the same object concurrently rather than being separated by the interpreter's own scheduling. ## What to do instead When you find yourself reaching for tee, ask which of these actually applies: 1. **Restructure into one pass.** Most two-pass problems (a count and a sum, a maximum and a filter against it) collapse into a single loop that maintains both accumulators. This is almost always the right answer, and interviewers are listening for it. 2. **Re-create the source.** If the source is a file or a query you can run again, opening it twice costs I/O but keeps memory flat — often the correct trade for very large inputs. 3. **Materialize deliberately.** If the data fits, `rows = list(source)` is honest, debuggable, indexable and measurable. Choosing it on purpose is not a defeat. 4. **Use tee only for genuine lockstep fan-out**, where two consumers advance together — for example feeding a running aggregate and a writer from one pass. ## Measuring rather than arguing `tracemalloc.start()` followed by `tracemalloc.get_traced_memory()` around each shape settles it empirically: peak allocation for the lockstep version is flat, and for the drain-one-then-the-other version it tracks the size of the source. That measurement is the difference between quoting a rule and knowing why the rule exists.

  • When is itertools.tee genuinely the right tool rather than a disguised list?
    When the branches advance together. Feeding two consumers in lockstep — a running aggregate and a writer, say, pulled one item at a time from each — keeps the retained gap at a couple of items, so tee streams properly and never materializes. The moment one consumer can outrun the other by an unbounded distance, the buffer becomes the whole stream and an explicit list is the clearer choice.
  • You need both the mean and the maximum of a huge stream. What do you do?
    One pass that maintains all three accumulators: a running total, a count and a running maximum. No tee, no list, constant memory. Reaching for tee or a materialized list here is a sign of not having considered restructuring, which is exactly what the question is probing.
  • Can you keep using the original iterator after passing it to itertools.tee?
    No. tee assumes exclusive ownership of the source: any item you pull yourself is never placed in the shared buffer and so is invisible to every branch. The result is silently skipped values rather than an exception, which makes it a nasty bug. Treat the original name as dead after the call.

Two people reading one printout as it comes off the printer. If they read side by side, one page is enough. If the first races to the end while the second waits, someone has to keep every page they have both not finished — which is just the whole stack, in a more awkward pile.

saying these in an interview costs you the question

  • Claims tee lets you re-iterate with no memory cost
  • Uses tee then drains one branch completely first
  • Keeps consuming the original iterator after teeing
  • Thinks tee re-reads the underlying source twice
  • Shares tee branches across threads without a lock

context