skip to content

How do you stream a 40 GB log file in Python without loading it into memory?

level: middleimportance: must knowfreq 70%

answer

  1. The file object is already an iterator
  2. Never read() or readlines() the whole thing
  3. Chain generator expressions, keep every stage lazy
  4. The consumer decides whether it streams
  5. The handle must outlive the pipeline

basics

~20 s

Iterate the open file object directly with a for loop: it is its own iterator and hands back one buffered line at a time. Never call read() or readlines(), and keep every downstream stage lazy so nothing collects the whole file.

solid answer

~50 s

A file object opened with `open()` is an iterator over its own lines, so `for line in handle:` reads a buffered chunk from disk and yields one line at a time — peak memory is the buffer, not the file. The mistakes are `handle.read()` and `handle.readlines()`, which both pull all 40 GB in. Layer transformations as generator expressions so the pipeline stays lazy, and use `itertools.islice` to take a sample or `itertools.chain` to treat several files as one stream. Two traps matter in production: the consumer must also be streaming, because `sorted()`, `list()` or `len()` collapses the pipeline back into memory; and the file must stay open while the lazy pipeline is consumed, so returning a generator out of a `with` block gives you a closed file. For binary work, read fixed-size chunks with the two-argument form of `iter()`.

code

python · 11 lines
python
import itertools
import tempfile
from pathlib import Path

path = Path(tempfile.mkdtemp()) / "picks.log"
path.write_text("".join(f"sku-{i},qty=2\n" for i in range(1000)), encoding="utf-8")

with path.open(encoding="utf-8") as handle:
    stripped = (line.rstrip("\n") for line in handle)
    wanted = (line for line in stripped if "qty=" in line)
    print(list(itertools.islice(wanted, 3)))

go deeper

for a junior

Be ready to write the four-line answer without hesitating: open the file, loop over the handle, process each line, and never call read() or readlines() on something large. Say why the naive version dies.

for a middle

Explain the mechanics beneath the idiom — buffered reads rather than a syscall per line, generator expressions chained into a pipeline, islice and chain for shaping it, and which consumers silently collapse the laziness.

for a senior

Demonstrate the production judgement: file-handle lifetime across a lazy pipeline, decode errors surfacing hours in, a single unbounded line, compressed and memory-mapped sources, and proving peak memory with tracemalloc rather than asserting it.

for a principal

Own the boundary. Say when a single-pass stream is the right architecture and when the job genuinely needs chunking, an external sort or a database, and how you would keep a team from spreading lazy pipelines whose resource lifetimes nobody can reason about.

## The file object is already lazy The answer an interviewer wants first is one line: ```python with open(path, encoding="utf-8") as handle: for line in handle: process(line) ``` A text file object is its own iterator. Iterating it reads a buffered block from the operating system (tens of kilobytes, not one syscall per line), splits off the next line, and yields it. Peak memory is the buffer plus the longest single line, whether the file is 40 megabytes or 40 gigabytes. The two ways to get this wrong are `handle.read()`, which returns the entire contents as one `str`, and `handle.readlines()`, which returns a list of every line — both allocate the whole file, and on a 40 GB input both simply die. `for line in handle.readlines()` is the classic near-miss: it looks like the streaming version and behaves like the eager one. ## Build a pipeline, keep it lazy Real work is rarely one loop. Layer generator expressions so each stage transforms one item at a time: ```python stripped = (line.rstrip("\n") for line in handle) records = (line.split(",") for line in stripped if line) totals = sum(int(row[2]) for row in records) ``` Nothing between `handle` and `sum()` holds more than one row. `itertools` is the toolbox for shaping such a stream without breaking it: `islice` takes the first N items (ideal for sampling a huge file while developing), `chain` and `chain.from_iterable` concatenate several sources into one continuous stream, and `batched` (added in 3.12) groups items into fixed-size tuples when a downstream API prefers chunks. Since 3.12 `batched` returns tuples and, from 3.13, accepts `strict` to reject a short final chunk. The crucial rule is that **the consumer decides**. A lazy source feeding `sorted()`, `list()`, `tuple()`, `len()` or `str.join` is not streaming at all — those drain their input completely before producing anything. Streaming end to end means finishing with `sum()`, `min()`, `max()`, `any()`, a `for` loop, or a writer that emits each result as it is produced. ## Lifetime: the trap that bites in production Because the pipeline is lazy, the file must still be open when the last item is pulled. This is broken: ```python def rows(path): with open(path) as handle: return (line.split(",") for line in handle) # closed on return ``` The `with` block exits at `return`, so the caller gets a pipeline over a closed file and a `ValueError` on the first item. Either consume inside the `with` block, or make the function itself a generator function so the block stays open across suspensions, or pass an already-open handle in and let the caller own its lifetime. ## Details that decide whether this survives real data - **Encoding and newlines.** Decoding happens as you iterate, so a bad byte halfway through a 40 GB file raises `UnicodeDecodeError` after hours of work. Choose the encoding explicitly and consider `errors="replace"` for logs of unknown provenance. - **One enormous line.** Line iteration bounds memory per *line*, not absolutely. A file with no newlines is read entirely into one string. If that is possible, read fixed-size chunks instead. - **Binary chunking.** The two-argument form of `iter()` turns any zero-argument callable into an iterator that stops at a sentinel: `iter(functools.partial(handle.read, 1 << 16), b"")` yields 64 KiB blocks until end of file. - **Compression.** `gzip`, `bz2`, `lzma` and (since 3.14) `compression.zstd` all return file objects that iterate lazily, so a compressed log streams exactly the same way. - **Memory mapping.** `mmap` gives random access to a large file without reading it into the heap, which is the right tool when you need to seek around rather than sweep forward. - **Measuring, not guessing.** `tracemalloc.start()` and `tracemalloc.get_traced_memory()` around the loop report peak allocation and settle arguments about whether a pipeline is really streaming. ## Where laziness stops being enough If the job needs a global sort, a join against another large source, or several passes, no amount of laziness saves you: you need bounded buffering, an external sort, chunked processing, or a database. Streaming solves *peak memory for a single forward pass*, and knowing that boundary is the difference between reciting the idiom and understanding it.

  • A function opens a file in a with block and returns a generator expression over it. What goes wrong?
    The `with` block exits at `return`, closing the file before anything has been read. The caller holds a lazy pipeline over a closed handle and gets `ValueError: I/O operation on closed file` on the first item. Fix it by consuming inside the block, by making the function a generator function so the block stays open across suspensions, or by accepting an already-open handle and leaving the lifetime to the caller.
  • Your streaming pipeline still exhausts memory. What is the most likely cause?
    An eager consumer at the end. `sorted()`, `list()`, `tuple()`, `len()` and `str.join` all drain their input completely, so a perfectly lazy chain in front of them buys nothing. The other usual causes are an accumulating dict or set built per item, and a source with no newlines so one line is the whole file. `tracemalloc` peak figures identify which.
  • How would you process a 40 GB file that is gzip-compressed?
    Identically. `gzip.open()` returns a file object that iterates lazily, decompressing a buffer at a time, so the same `for line in handle:` loop works and peak memory stays flat. The same is true of `bz2`, `lzma` and, since 3.14, `compression.zstd`. What you cannot do is seek cheaply inside a compressed stream, so random access needs the uncompressed file or an index.

saying these in an interview costs you the question

  • Reaches for readlines() and calls it streaming
  • Reads the file into a string and splits on newlines
  • Ends the lazy pipeline with sorted() or list()
  • Returns a generator out of a closed with block
  • Assumes line iteration bounds memory with no newlines
  • Claims a compressed file must be decompressed to disk first

context