skip to content

A nightly report job's custom iterable caches every row so callers can loop it twice, and memory grows unbounded. How would you redesign the class?

level: seniorimportance: should knowfreq 32%

answer

  1. Re-iterable does not have to mean retained
  2. Keep the recipe, not the rows
  3. The method rebuilds the pass each call
  4. Open the source inside the generator
  5. A with block frees the handle early

basics

~20 s

Make it a restartable iterable: store only how to reach the source, and have iter open a fresh handle inside a with block and yield rows. Nothing is retained between passes, so memory stays flat.

solid answer

~40 s

The class caches because it wants to be re-iterable, but re-iterability does not require retention -- it requires that `__iter__` can *rebuild* the pass. Store the parameters (a factory, a path, a query) instead of the rows, and write `__iter__` as a generator that opens the source under a `with` block, yields row by row, and closes on the way out. Each call is then an independent pass with its own handle, memory is O(1) in the number of rows, and nested passes no longer share state. The cost moves into the open: a second pass re-pays retrieval, so if that is a 45-second cold start, make it visible rather than hiding it in a cache. Callers needing two passes then choose their own remedy -- `list()`, `itertools.tee`, or one pass computing everything.

code

python · 18 lines
python
import io


class Rows:
    def __init__(self, open_source):
        self._open_source = open_source

    def __iter__(self):
        with self._open_source() as stream:
            for line in stream:
                yield line.rstrip("\n")


data = "a,1\nb,2\nc,3\n"
rows = Rows(lambda: io.StringIO(data))
print(list(rows))
print(list(rows))
print(sum(1 for _ in rows))

go deeper

for a junior

Take away the core idea: a class can be looped twice without storing anything, if __iter__ re-opens the source and yields. Storing the rows is only one way to be re-iterable, and the expensive one.

for a middle

Explain the mechanics: __iter__ as a generator with a with block, one fresh handle per pass, memory flat in the number of rows, and the fact that the body does not run until the first element is pulled.

for a senior

Show production judgement -- confirm the leak with tracemalloc snapshot diffs before redesigning, reason about handle lifetime under an early break, count how many connections two concurrent passes open, and weigh a second slow pass against materialising once.

for a principal

Own where cost lives. Decide whether re-reading, caching or a single-pass restructure is the standard for this pipeline, who pays the cold start, and how the single-pass versus restartable contract is expressed in types so the next person does not rediscover it by leaking memory.

The symptom -- resident memory climbing through the night until the job is killed -- and the cause are one design decision apart. Somebody needed to loop the rows twice, discovered that the object was single-pass, and fixed it by keeping the rows. That works until the report grows. ### The distinction that solves it Re-iterable does not mean *retained*. It means every `__iter__()` call can produce a fresh, independent pass. There are exactly two ways to achieve that: keep the data, or keep the *recipe* for getting it. The first is O(n) memory forever; the second is O(1), and for anything sourced from a file, a query, a socket or an API it is almost always the right one. So invert what the object holds. It stores a factory -- a callable that opens the source, or the path and parameters needed to do so -- and never the rows. `__iter__` is a generator function that opens the source, yields each row, and closes it when the pass ends. ```python class Rows: def __init__(self, open_source): self._open_source = open_source def __iter__(self): with self._open_source() as stream: for line in stream: yield line.rstrip("\n") ``` That class is re-iterable, nested-loop safe, and flat in memory. This is usually called a *restartable iterable*, and it is the standard shape for a lazy view over an external source. ### Resource lifetime, which is where this pattern goes wrong A generator `__iter__` that opens a handle owns that handle for the life of the generator, not the life of the call. Three consequences to be explicit about. **Use `with`, not a bare open.** If the consumer breaks out of the loop early, the generator is left suspended inside the `with`. When it is garbage-collected -- or when someone calls `close()` on it -- `GeneratorExit` is thrown in at the `yield`, the `with` block unwinds, and the handle is released. Without the `with` there is nothing to unwind and the descriptor leaks until process exit. Relying on collection is still nondeterministic under reference cycles, so a consumer holding the generator in a long-lived structure can hold the handle open for a long time; `contextlib.closing` around an explicit iterator, or simply consuming the pass fully, removes the doubt. **Each concurrent pass opens its own handle.** That is the price of independence, and it is the right trade -- but a nested loop over the same object now holds two connections, and a threaded consumer holds one per thread. Bound it: if the source is a connection pool, iterating in a loop over another iteration of the same object can exhaust it. **Nothing runs until the first `next()`.** A generator body is lazy, so the `open` inside `__iter__` does not happen at `__iter__()` time. `obj.__iter__()` succeeding tells you nothing about whether the source exists; the error surfaces at the first element. If early validation matters, do it in `__init__` or in a non-generator `__iter__` that opens the source and returns an inner generator. ### Where materialisation should live Pushing the cache out of the class does not make the second pass free. Name the three options and let the call site choose. *Two passes over the source.* Simplest, and correct when the source is cheap. When it is not -- a 45-second cold start before the first row arrives -- doubling that is a real cost, but at least it is a cost the caller opted into and can measure. *`list()` at the call site.* One line, obvious in review, and the memory is visible where the decision is made rather than hidden inside a class everybody uses. *`itertools.tee`.* Two independent iterators from one pass, but tee buffers every item in the gap between the fastest and slowest consumer. If one consumer runs to completion first, tee has buffered the entire source -- you have reinvented the leak with more machinery. It pays off only when the consumers advance roughly together. Often the best answer is none of them: restructure the two passes into one. A total and a normalisation, or a count and a filter, can usually be computed in a single walk with a couple of accumulators, and then the question disappears. ### Making the contract explicit Whichever way it lands, say so in the type. A restartable view should subclass `collections.abc.Iterable`, be named like a collection, and document *each iteration re-reads the source*. A genuinely single-pass reader should subclass `collections.abc.Iterator`, be named like a cursor, and document that a second pass yields nothing. Callers can then test with `isinstance(obj, collections.abc.Iterator)` before deciding whether they must materialise. Most of these bugs come from a type that never said which it was. ### Confirming the diagnosis first Before redesigning, prove it. Take `tracemalloc` snapshots at intervals during a run and diff them; a retained list of row objects shows up as a single allocation site holding a growing share. That distinguishes the buffering iterable from the other usual suspects -- an unbounded cache, an accumulating log handler, or reference cycles keeping objects alive.

  • The consumer breaks out of the loop halfway. What closes the file the generator opened?
    Closing the generator does. `GeneratorExit` is thrown in at the suspended `yield` -- either by an explicit `close()` or when the generator is collected -- and the `with` block unwinds and releases the handle. Without a `with` there is nothing to unwind. Because collection timing is not guaranteed, especially under reference cycles, prefer consuming the pass fully or closing the iterator explicitly when the handle is scarce.
  • Why is `itertools.tee` not a general fix for needing two passes?
    `tee` buffers every element in the gap between its fastest and slowest consumer. If one consumer runs to the end before the other starts, the buffer holds the whole source and you have recreated the memory problem with extra indirection. It is the right tool only when the two consumers advance roughly in step; otherwise materialise explicitly or restructure the work into a single pass.
  • How would you confirm the buffering iterable is the leak rather than something else?
    Take `tracemalloc` snapshots at intervals during a run and diff them, grouped by allocation site: a retained row buffer appears as one site whose size grows monotonically with rows processed. Compare that against the other candidates -- an unbounded cache, an accumulating handler, reference cycles keeping objects alive -- and confirm by running with the caching removed and watching resident memory flatten.
  • How do you make the single-pass versus restartable contract visible to callers?
    Put it in the type. A restartable view subclasses `collections.abc.Iterable`, is named like a collection, and documents that each pass re-reads the source; a genuine cursor subclasses `collections.abc.Iterator`, is named like a reader, and documents that a second pass yields nothing. Callers can then branch on `isinstance(obj, collections.abc.Iterator)` instead of guessing, and reviewers can see the intent.

A restartable iterable is a library card rather than a photocopy: each reader fetches the book again instead of everyone sharing one ever-growing stack of copies on your desk.

saying these in an interview costs you the question

  • Says re-iterable requires keeping the data in memory
  • Opens the source in __init__ and reuses one handle
  • Reaches for itertools.tee without mentioning its buffer
  • Omits the with block inside a generator __iter__
  • Assumes the generator body runs when __iter__ is called
  • Hides materialisation inside the class instead of at the call site

context