skip to content

Why can zlib.compress exhaust memory on a multi-gigabyte file, and what replaces it?

level: middleimportance: must knowfreq 50%

answer

  1. The whole input must already be in hand
  2. Two API shapes per codec, not one
  3. Feed it, then drain it
  4. Bounded by chunk plus window, not file size
  5. compressobj carries per-stream state

basics

~10 s

zlib.compress takes one bytes object and returns another, so both the whole input and the whole output sit in memory at once. Stream instead: zlib.compressobj(), feed it chunks with compress(), and finish with flush().

solid answer

~40 s

`zlib.compress(data)` is a one-shot API — it needs the entire input as a single `bytes` object and builds the entire output as another, so peak memory is roughly input plus output. For a multi-gigabyte capture that is fatal. The incremental API exists for exactly this: `co = zlib.compressobj(level=6)` gives you a stateful compressor; you loop over fixed-size reads calling `co.compress(chunk)` (which returns whatever bytes are ready, often `b''`), then call `co.flush()` once at the end to drain the encoder. Memory is then the chunk size plus deflate's own window and buffers, regardless of file size. `zlib.decompressobj()` mirrors it, with `unused_data` and `eof` telling you where the stream ended. If you just want a gzip file, `gzip.open(path, 'wb')` plus `shutil.copyfileobj` does the same streaming for you.

code

python · 13 lines
python
import zlib

def compress_stream(src, dst, chunk=1 << 20):
    co = zlib.compressobj(level=6, wbits=zlib.MAX_WBITS | 16)  # gzip container
    with open(src, "rb") as fin, open(dst, "wb") as fout:
        while block := fin.read(chunk):
            fout.write(co.compress(block))
        fout.write(co.flush(zlib.Z_FINISH))

with open("capture.log", "wb") as f:
    f.write(b"sensor=7 celsius=21.5\n" * 50_000)

compress_stream("capture.log", "capture.log.gz")

go deeper

for a junior

Recall that zlib.compress and gzip.compress want the whole payload as one bytes object, and that the stdlib offers an incremental alternative for files too big to hold. Knowing the split exists is the bar here.

for a middle

Explain the mechanics: compressobj(), chunked compress() calls that may return b'', a terminal flush(), and what the bounded working set actually consists of. Be ready to say why the object cannot be shared between streams.

for a senior

Show the operational judgement — bounded memory as a deployment constraint, wbits chosen so operators can read the output with standard tools, max_length on decompression of data you did not produce, and a per-request compressor rather than a module-level one.

for a principal

Own the pipeline shape: where compression happens relative to batching and upload, whether a mid-stream flush is worth the ratio loss for freshness, and how bounded memory per worker sets the concurrency your collector fleet can run.

### One-shot versus incremental Every compression module in the stdlib ships two shapes of API. The one-shot pair — `zlib.compress`/`zlib.decompress`, `gzip.compress`/`gzip.decompress`, `bz2.compress`, `lzma.compress`, and on 3.14 `compression.zstd.compress` — takes a complete `bytes` object and returns a complete `bytes` object. That is perfect for a config blob or an HTTP body of a few hundred kilobytes, and it is a memory bug waiting to happen for a 4 GB telemetry capture: you must already hold the input, and the interpreter must allocate the output beside it. The incremental pair is the answer. `zlib.compressobj(level=-1, method=8, wbits=15, memLevel=8, strategy=0, zdict=None)` returns a compressor object with three methods: `compress(data)`, `flush([mode])` and `copy()`. You feed it as much or as little as you like; it returns the bytes that are ready *so far*, which is frequently `b''` because deflate buffers until it has enough to emit. When the input is exhausted you call `flush()` — by default `zlib.Z_FINISH` — and it returns the tail of the stream, including the closing bits. After a `Z_FINISH` flush the object is finished and cannot be reused. ### What the memory ceiling actually is Streaming does not make compression free; it makes it *bounded*. The working set is your read chunk (pick something like 1 MiB), plus deflate's sliding window — 32 KiB at the default `wbits=15` — plus internal hash and output buffers sized by `memLevel`. A few hundred kilobytes, constant, whether the file is 4 MB or 400 GB. That is the property that lets a collector process run in a small memory limit. ### Choosing the container with wbits `wbits` does double duty: it sets the window size *and* selects the wrapper. `15` gives a zlib stream (the 2-byte zlib header plus an Adler-32 trailer), `-15` gives a raw deflate stream with no wrapper, and `15 | 16` gives a gzip stream — header, CRC32 and length trailer — which is what you want if the file should be readable by ordinary gzip tools. That last one is the difference between a file your operators can open and a blob only your code understands. ### The stateful-object trap A `compressobj` is not a pure function; it carries the encoder state for one stream. Hoisting it to module level and reusing it across calls — the compression flavour of any shared mutable default — produces output that decompresses correctly the first time and then fails: ```python import zlib _shared = zlib.compressobj() # module-level: wrong def bad_compress(payload): return _shared.compress(payload) + _shared.flush(zlib.Z_FINISH) ``` The first call finishes the stream; the second raises `zlib.error: Error -2 while compressing data: inconsistent stream state`. Create a fresh compressor per stream. The same rule applies to `zlib.decompressobj`, `bz2.BZ2Compressor` and `lzma.LZMACompressor`, which answer a post-flush `compress()` with `ValueError: Compressor has been flushed`. On 3.14 `compression.zstd.ZstdCompressor` is the one exception: flushing with `FLUSH_FRAME` closes a frame and the same object may begin the next one. `copy()` exists precisely because branching a partially-fed stream otherwise has no safe form. ### Decompressing incrementally `zlib.decompressobj()` is the mirror image, with two attributes worth knowing. `unused_data` holds whatever followed the end of the compressed stream — useful when a compressed member is embedded in a larger frame — and `eof` becomes `True` once the stream's own terminator has been seen, which is how you tell a *complete* stream from a *truncated* one. `unconsumed_tail` holds input you fed that could not be processed because you capped the output with the `max_length` argument, which is the standard defence when decompressing data whose expansion ratio you do not trust. ### The higher-level shortcut Most of the time you do not need `compressobj` at all. `gzip.open(path, 'wb')` returns a file object that streams, so: ```python import gzip, shutil with open('capture.log', 'rb') as src, gzip.open('capture.log.gz', 'wb') as dst: shutil.copyfileobj(src, dst, length=1 << 20) ``` is fully streaming, bounded, and readable by any gzip tool. Reach for `compressobj` when you need something the file wrapper does not give you: raw deflate, a preset dictionary via `zdict`, `Z_SYNC_FLUSH` to emit a decodable boundary mid-stream for a live feed, or compression of a stream you are producing rather than a file you can open. ### What the interviewer is listening for Three things: that you name the one-shot/incremental split rather than reaching for a bigger machine; that you can say what stays bounded and what does not; and that you know the compressor object is per-stream state, not a reusable service object.

  • How do you produce a real gzip file with zlib.compressobj rather than a bare zlib stream?
    Pass `wbits=zlib.MAX_WBITS | 16`. `wbits` picks both the window size and the wrapper: `15` is a zlib stream with an Adler-32 trailer, `-15` is raw deflate with no wrapper at all, and `15 | 16` writes the gzip header and the CRC32-plus-length trailer that ordinary gzip tools expect. Without it you get bytes that only your own code can read.
  • What do unused_data and eof on a zlib.decompressobj tell you?
    `eof` becomes `True` once the decompressor has seen the stream's own end marker, which is how you distinguish a complete stream from a truncated one — a short read alone cannot tell you. `unused_data` holds any bytes that followed that end marker, so when a compressed member is embedded in a larger frame you can pick up the remainder and hand it to the next parser.
  • When would you flush mid-stream instead of only at the end?
    For a live feed. `flush(zlib.Z_SYNC_FLUSH)` emits everything buffered so far and ends it at a byte boundary the decompressor can act on, so a reader sees your data without waiting for the stream to close — at the cost of a worse ratio, since the encoder loses buffered context. `Z_FINISH` is terminal: after it the object cannot accept more input.

saying these in an interview costs you the question

  • Says just read the file in chunks and call zlib.compress on each
  • Reuses one compressobj across independent streams
  • Forgets flush(), producing a truncated stream
  • Assumes compress() returns bytes on every call
  • Thinks streaming reduces total memory rather than bounding it
  • Believes a raw zlib stream is a gzip file

context