A geocoding batch streams a 6 GB CSV with csv.DictReader and writes each result as its lookup finishes - how do you keep memory flat and the output rows correctly matched?
answer
- The reader is already lazy, you are not
- list() over the reader is the leak
- The file must stay open while iterating
- Completion order is not input order
- Carry a key column, never a position
basics
~20 sIterate the reader lazily and never materialise it with list(), keeping one record alive at a time. Never rely on output position matching input position: carry a stable key column from each input record into the output row you write.
solid answer
~50 sMemory is the easy half. `csv.DictReader` is a lazy iterator over an open file object, so a plain `for row in reader:` loop holds one record at a time regardless of file size; the leak is almost always your own code accumulating rows, or a `list(reader)` written for convenience. Keep the file open for the whole iteration - returning the reader out of a `with` block closes it and the next step raises `ValueError`. The correctness half is the ordering assumption. When a warm cache returns some lookups instantly while others wait on the network, results arrive in an order unrelated to input order, so writing them positionally silently pairs each address with the wrong coordinates. Treat a column as the record's identity instead: carry it into every row you write with `csv.DictWriter`, and rejoin on that key, never on position.
code
python · 11 linesimport csv
import io
src = io.StringIO("id,address\n7,1 Rue A\n8,2 Str B\n9,3 Way C\n", newline="")
out = io.StringIO(newline="")
writer = csv.DictWriter(out, fieldnames=["id", "lat", "lon"])
writer.writeheader()
for row in csv.DictReader(src): # one row alive at a time
writer.writerow({"id": row["id"], "lat": "48.85", "lon": "2.35"})
print(out.getvalue().splitlines())go deeper
Know that csv.reader and csv.DictReader are lazy iterators, so a for loop over a huge file is fine while calling list() on it is not. Keep the file open for as long as you are reading it.
Explain where the memory actually goes: accumulator lists, unbounded caches, or a reader materialised during development. Describe how the reader ties to the open file object and why a closed file raises on the next step.
Demonstrate the ordering judgement. Say out loud that completion order stops matching input order the moment lookups take different paths, and design the output row so identity travels in a key column, making the batch both correct and restartable.
Own the pipeline contract: what identifies a record, who guarantees ordering, how bad records are routed rather than fatal, and whether a flat file is the right interchange at all once the batch needs restartability and keyed rejoin.
### The two independent failures A batch over a file that does not fit in memory has one obvious risk and one that is far more expensive, because it does not raise. ### Memory: the module is already streaming `csv.reader` and `csv.DictReader` are iterators over the file object, not loaders. Each step consumes only as many lines as the current record needs, and once you drop the reference the record is collectable. A loop over a 6 GB file holds one record. So when RSS climbs, the module is virtually never the cause. Look for: - `rows = list(reader)` or a comprehension over the reader - written during development on a small sample, never revisited. - A results list, dict or set accumulated for a summary at the end. A per-key accumulator is bounded by the number of distinct keys, not by rows, and that is usually fine; a per-row accumulator is not. - A cache with no eviction. If the batch memoises lookups and the key space is large, that cache is the file. - Objects retained by an exception traceback or a logging handler that captured the record. The writing side streams too: `csv.DictWriter` formats and writes each row as it is given, so producing a 6 GB output is equally flat, as long as you do not buffer results to sort them later. One structural rule: the reader is only valid while its file is open. A helper that opens a file with `with`, builds a reader and returns it hands back an iterator over a closed file, and the caller gets `ValueError` on the first step. Either do the work inside the `with`, or make the helper a generator so that the `with` block stays alive across yields. ### Ordering: the failure that stays silent The expensive bug is the assumption that output row *n* corresponds to input row *n*. It holds only while every lookup takes the same path. Put a cache in front of the geocoder with a hit rate around 83% and it stops holding immediately: the hits return in microseconds while the misses wait on the network, so if any concurrency exists the completion order is close to random with respect to input order. Write those results positionally and every address is paired with someone else's coordinates. Nothing raises. The output file has the right number of rows, the right columns and plausible values, and the defect is found weeks later - if at all. The fix is to stop encoding identity in position. Choose a column that identifies the record - a supplied id, or a hash of the address if there is no id - and carry it through the whole pipeline: into the work item, back with the result, and out as a column of the row you write with `csv.DictWriter`. Then position carries no meaning, out-of-order writing is harmless, and rejoining is a key-based operation. If a consumer really needs input order, sort by that key afterwards, or write to a keyed store and emit in order at the end - either way, the ordering is restored explicitly rather than assumed. That property is also what makes the batch restartable. If the output row carries the input key, a crashed run can be resumed by reading which keys are already done and skipping them, instead of counting rows and hoping. ### Parser limits worth knowing `csv.field_size_limit()` returns the maximum field size the parser will accept and, called with an argument, sets it and returns the previous value. Exceed it and the parser raises `csv.Error` - not a data error you can skip past, but the whole record failing. Files with an embedded blob or a long free-text column hit this, and raising the limit deliberately at start-up beats discovering it mid-run. Ragged records are the other routine defect. `csv.DictReader` with `restkey=` and `restval=` turns them into something inspectable rather than something silently truncated, which lets the batch count and route bad records instead of dying on one. ### Operating the run Open source and destination together, write as you go, and let the OS buffer handle throughput. Flush at a checkpoint interval if a crash must not lose the last chunk. Log progress by record count so a stall is visible. And test the ordering logic honestly: a test where every lookup is instant will pass on a pipeline that is positionally broken, so the test that earns its keep is the one that returns results deliberately out of order and asserts that each output row still carries the coordinates its own key asked for.
- Why does returning a csv.reader from a helper that used a with block fail?The reader is a lazy iterator over the file object, so closing the file leaves the caller with an iterator that raises `ValueError` on the first step. Either do the work inside the `with`, or make the helper a generator function that yields rows - the `with` block then stays alive for the life of the iteration.
- What raises csv.Error on an otherwise well-formed record?A field larger than `csv.field_size_limit()`. The parser refuses the record rather than truncating it, which is correct but fatal mid-run. Files with an embedded blob or a long free-text column hit this, so raise the limit explicitly at start-up if you expect large fields, and remember the call returns the previous value.
- How would you test that the batch is not relying on row order?Drive it with a stub whose completion order is deliberately shuffled relative to input, then assert per key that each output row holds the result its own key asked for. A test where every lookup returns instantly preserves input order by accident and will pass on a pipeline that is positionally broken.
- How do you make such a batch resumable after a crash?Because each output row carries the input key, a restart can read the keys already written and skip them, then append. That works precisely because identity lives in a column rather than in row position - a run that counted rows to find its place would be wrong the moment any record was written out of order or skipped.
It is a coat check. Give every coat a numbered ticket and the order they come back stops mattering; without tickets you are trusting that everyone returns in the order they arrived.
saying these in an interview costs you the question
- Calls list() on the reader to get a row count first
- Assumes output row n matches input row n
- Says csv.DictReader loads the whole file into memory
- Returns a reader from a closed with block
- Buffers all results in memory to restore order
- Cannot say what csv.field_size_limit protects against