An importer's helper returns a lazy line generator built inside `with open(...)`; why does the caller get ValueError?
answer
- The block ends the moment you return
- Nothing has been read yet
- Laziness outlives the resource
- Move the yield inside the block
- Or let the caller own the with
basics
~20 sThe with block ends when the helper returns, so the file is closed before the caller consumes the lazy generator, and the first read raises ValueError. Fix it by materialising inside the block or yielding from a generator function.
solid answer
~50 sA `with` statement closes the file when its block exits, which for a plain function is the moment it returns. If the value returned is a generator expression over the file, nothing has been read yet; the caller iterates a stream that is already closed and gets `ValueError: I/O operation on closed file`. It looks intermittent only because call sites that consume the result inside the block never trip it. Three fixes are honest: read eagerly inside the block when the data is small; make the helper a **generator function** with `yield` inside the `with`, so the block stays open until iteration finishes; or hand the caller the open file object and let it own the `with`. The lazy shapes matter for the same reason `for line in f` beats reading everything: only one line is resident at a time.
code
python · 17 linesimport io
def broken(stream):
with stream:
return (line.strip() for line in stream)
def fixed(stream):
with stream:
for line in stream:
yield line.strip()
try:
print(list(broken(io.StringIO("a\nb\n"))))
except ValueError as exc:
print("broken:", exc)
print("fixed:", list(fixed(io.StringIO("a\nb\n"))))go deeper
Remember that a with block closes the file as soon as control leaves it, including on a return, and that reading must happen inside the block. Prefer 'for line in f' over reading the whole file.
Explain why a generator expression built inside the block reads nothing before the close, and how moving the yield into a generator function ties the block's lifetime to the iteration instead of to the call.
Diagnose it from the symptom: a ValueError that follows the call site rather than the data. Weigh the three fixes on memory and handle lifetime, and know that relying on refcounting to close files is an implementation detail, not a guarantee.
Set the ownership convention: helpers take open streams, callers own the with, and stand-in streams make the code testable without the filesystem. Decide where holding handles open across slow consumers is acceptable at all.
This is the failure that teaches what `with` actually guarantees. Consider a museum-catalogue importer whose helper opens an export of about 6,800 object records and returns rows to a caller. ```python def rows(path): with open(path, encoding="utf-8") as f: return (line.rstrip("\n").split("\t") for line in f) ``` Every call site that does `for row in rows(path)` raises `ValueError: I/O operation on closed file` on the first iteration. ## Why `with` binds the file to a context manager: `__enter__` runs at the top of the block and `__exit__` runs when control leaves it, *for any reason* — falling off the end, returning, raising. `__exit__` on a file closes it. A generator expression, meanwhile, is lazy: constructing it reads nothing. So `return` triggers `__exit__`, the stream closes, and the object you handed back still holds a reference to a closed file. The reference keeps the object alive; it does not keep it usable. The reason this reaches production looking flaky is that it depends on the call site rather than on the data. A caller that wraps the result in `list()` *inside* the same block works. A caller that only checks whether the result is falsy never touches the file. A test that passes a small fixture but consumes it immediately passes. The one call site that stores the generator and iterates later is the one that fails, and it fails identically every time — it is the code path that is intermittent, not the bug. ## The three fixes, and when each is right **1. Materialise inside the block.** `return [line.rstrip("\n").split("\t") for line in f]` reads everything before `__exit__` runs. Simple and correct, and fine for a configuration file. For a large export it means holding every row in memory at once, which is the same objection as `readlines()`. **2. Make the helper a generator function.** Move the `yield` inside the `with`: ```python def rows(path): with open(path, encoding="utf-8") as f: for line in f: yield line.rstrip("\n").split("\t") ``` Now calling `rows(path)` runs no body at all; the file is not even opened until the first `next()`. The `with` block belongs to the generator's own frame, so it stays open across the whole iteration and closes when the loop ends — or when the generator is closed or garbage-collected, which is what happens if the caller abandons it half way through. This keeps memory flat at one row and is usually the right answer for streaming work. It does come with a caveat worth stating in an interview: the file is now open for as long as the caller takes to consume the generator. If a consumer holds it across a slow network call, you are holding a file handle for that whole time. That is a deliberate trade, not an accident, and it is the reason the third option exists. **3. Give the caller the handle.** Have the helper take an already-open file object, and let the caller write the `with`. The helper becomes trivially testable — anything with the same reading protocol works, including an `io.StringIO` built from a literal string in a test, or an `io.BytesIO` for the binary case — and ownership of the resource sits with the code that knows how long it is needed. Where several resources must be opened conditionally, `contextlib.ExitStack` lets the caller collect them in one block. ## The related discipline: iterate, do not slurp The same laziness question shows up without any `with` bug. Iterating a text file yields one line at a time and holds only that line. `readlines()` builds a list of every line as a separate `str`, and `read()` builds one enormous `str`; both are proportional to the file. For a 6,800-row catalogue export none of this matters. For a multi-gigabyte one it is the difference between a steady process and an out-of-memory kill, and the habit of writing `for line in f` costs nothing on the small files either. `readlines()` earns its place when you genuinely need random access or a length up front. ## And do not rely on the garbage collector A final point that this bug's mirror image raises: if you drop `with` and just let the file object go out of scope, CPython's reference counting usually closes it promptly — but that is an implementation detail, not a language guarantee. If the file is caught in a reference cycle, captured by a traceback held for later, or the code runs on an implementation without reference counting, the close is deferred to an unpredictable moment. Under development mode you will see a `ResourceWarning` telling you so. `with` is what makes the close deterministic; the whole point of this failure is that the guarantee is real and it fires exactly when the block ends, including on the `return` you did not think of as leaving the block.
- What is the cost of the generator-function fix in a long-running process?The file stays open for as long as the caller takes to consume the generator, so a consumer that iterates slowly — or abandons the generator without closing it — holds a file handle that whole time. Descriptor exhaustion under many concurrent consumers is the failure mode. Where the consumption window is unpredictable, prefer handing the caller the open file so it owns the lifetime explicitly.
- How would you exercise that helper without touching the filesystem?Take an already-open file object as a parameter rather than a path, then pass `io.StringIO("...")` for the text case or `io.BytesIO(b"...")` for the binary case. Both implement the same reading protocol, including iteration and the context-manager methods, so the helper cannot tell the difference. That also removes the encoding question from the helper, since the stand-in already holds decoded text.
- If you removed the with statement entirely, when would the file be closed?In CPython, when the last reference goes away, because reference counting collects it immediately — which is why the sloppy version usually appears to work. That is an implementation detail: a reference cycle, a traceback that captured the frame, or an interpreter without reference counting all defer the close indefinitely. Development mode reports the leak as a `ResourceWarning`. `with` is what makes it deterministic.
It is like photographing a library shelf label and leaving the building: the note tells you where the books are, but the doors locked behind you the moment you walked out.
saying these in an interview costs you the question
- Thinks returning a generator keeps the with block alive
- Says readlines() streams the file lazily
- Assumes garbage collection always closes files promptly
- Believes 'for line in f' reads the whole file first
- Tries to fix it by reopening the file in the caller
- Calls the failure random rather than call-site dependent