Compressing sensor telemetry with gzip level 9 saturates the collector's CPU — how do you pick a codec and level?
answer
- Answer with a table, not a preference
- The top level's return is tiny
- Read cost and write cost are different questions
- Data temperature drives the choice
- These C extensions release the GIL
basics
~20 sMeasure ratio and CPU on your own data. Level 9 typically buys under a percent over level 6 for several times the CPU, and codecs differ sharply: lzma smallest, zstd fast at a good ratio, gzip most portable.
solid answer
~50 sTreat it as a measurement, not a preference. Take a representative sample of the real telemetry and record, per candidate, three numbers: compressed size, compression time and *decompression* time — the last one matters because archives are usually read more than once. On typical log-shaped text, deflate levels 6 and 9 land within about half a percent of each other while level 9 costs several times the CPU, so the first fix is usually dropping to 6. Then pick by role: `gzip` for maximum tool compatibility, `bz2` for a middling ratio, `lzma` when the file is cold and size dominates, and on 3.14 `compression.zstd` when you want a good ratio at a fraction of the CPU. Stream in every case, and remember these codecs are C extensions that release the GIL, so a thread pool gives real parallelism across the collector's shards.
code
python · 15 linesimport bz2, gzip, lzma, time
from compression import zstd
data = b"".join(b"sensor=%d celsius=21.5\n" % i for i in range(200_000))
candidates = [
("gzip-6", lambda d: gzip.compress(d, 6)),
("gzip-9", lambda d: gzip.compress(d, 9)),
("bz2-9", bz2.compress),
("lzma", lzma.compress),
("zstd-3", lambda d: zstd.compress(d, 3)),
]
for name, fn in candidates:
start = time.perf_counter()
out = fn(data)
print(f"{name:8} {len(out):8} bytes {time.perf_counter() - start:6.2f}s")go deeper
Recall that compression level is a dial between CPU and file size, and that the codecs differ: gzip is the portable default, lzma makes the smallest files, and 3.14 adds zstd. You are not expected to run the benchmark yet.
Explain the diminishing returns of the top deflate levels, and that decompression speed for deflate barely depends on the level used to write. Be ready to describe how you would benchmark candidates on real sample data.
Demonstrate production judgement: tier by data temperature, weight read cost by how often archives are queried, keep memory bounded with streaming, and parallelise independent shards across threads because these codecs release the GIL.
Own the policy: one format for interchange and another for cold storage, who pays the re-compression cost when data ages, whether a 3.14-only codec is worth the interpreter floor, and how the choice interacts with retention and egress budgets.
### Reframe it as a measurement "Which compression should we use" has no context-free answer; it has a benchmark. Take a genuinely representative sample of the telemetry — not lorem ipsum, not a synthetic repeat, because compressibility is a property of *your* data — and for each candidate record compressed size, compression wall time, and decompression wall time. Three numbers, one table. Everything below is how to read that table. ### The level-9 reflex Deflate's `compresslevel` runs 0-9 with 9 as `gzip.compress`'s default, and the curve is steeply diminishing. On log-shaped text, going from 6 to 9 typically shaves a fraction of a percent off the output while multiplying compression CPU several times over, because the encoder spends far longer searching for marginally better matches. If a collector is CPU-bound in compression, dropping to level 6 — or lower — is usually the single change with the best return, and it costs nothing at read time: decompression speed for deflate is essentially independent of the level the data was written at. That asymmetry is worth stating out loud, because it is the one people forget. ### The codecs, by role * **`gzip`** — deflate in a gzip container. Middling ratio, fast, and readable by every tool and language on earth. The default choice when anything outside your code will open the file. * **`bz2`** — better ratio than deflate, notably slower both ways. Rarely the right answer today; it sits in a niche both neighbours have eaten into. * **`lzma`** — the best ratio of the classic three by a wide margin, and by far the slowest to compress; decompression is slow too, though much less so. The right choice for cold data: written once, read rarely, size is the cost that matters. * **`compression.zstd`** (3.14) — fast at a ratio close to or better than lzma's at comparable settings, with very fast decompression, plus trained dictionaries for small records. Where you are free to choose the format, this is usually where the measurement points. ### Read amplification decides more than write cost A file compressed once and read a hundred times has different economics from one written continuously and read on an incident. If the archive is queried, weight decompression speed heavily — a codec that saves 20% of bytes but doubles read time is a bad trade for a hot tier. If it is a legal-retention dump nobody touches, spend the CPU once at write time and take the smallest file. That naturally produces a tiered policy: a fast codec at a low level for the recent window, a re-compression pass into a high-ratio codec when data ages out. ### Where the CPU actually goes, and how to spend less of it Three levers, in order of return: 1. **Lower the level.** Free, immediate, usually near-lossless in ratio. 2. **Compress less.** Deduplicate, drop fields nobody queries, and batch small records so the encoder has enough context to work with — a stream of 200-byte records compresses far worse than the same bytes batched, which is also exactly what a trained zstd dictionary addresses. 3. **Use more cores.** The stdlib compression modules are C extensions that release the GIL around the actual work, so a `concurrent.futures.ThreadPoolExecutor` over independent shards genuinely parallelises — you do not need processes here, and you avoid paying to pickle the payloads across a process boundary. Alternatively `compression.zstd` can multi-thread inside a single call via `CompressionParameter.nb_workers`. With a collector fanning in from 17 upstream services, the natural unit is one shard per service: independent streams, one compressor object each, mapped across a small thread pool. ### Keep memory flat while you do it Whatever you choose, use the streaming API — `gzip.open(path, 'wb')` with `shutil.copyfileobj`, or the incremental compressor object — so the working set is a chunk plus the codec's window rather than the whole batch. A codec decision that quietly turns a bounded process into one that holds a gigabyte is not a win. ### Integrity, briefly gzip already carries a CRC32 and raises on a mismatch; `zlib.crc32` lets you record the same checksum for the *uncompressed* bytes in a manifest, so you can prove a round trip rather than only that the container is intact. zstd can write a content checksum into the frame via `CompressionParameter.checksum_flag`. Neither is a cryptographic guarantee — a CRC detects accidental corruption, not tampering. ### What a strong answer sounds like It names the measurement first, names level 9's poor return second, distinguishes write-side from read-side cost, and ends with a tiering policy rather than a single winner. An answer that just says "use zstd" has skipped every part the interviewer was asking about.
- Why do threads help here when the GIL usually blocks CPU-bound parallelism?Because the work is not happening in Python bytecode. `zlib`, `bz2`, `lzma` and `compression.zstd` are C extensions that release the GIL around the compression call itself, so several threads genuinely run on several cores. That makes a `ThreadPoolExecutor` the cheap choice for independent shards: no pickling of payloads across a process boundary, no child interpreters, and the memory of one process.
- How does the ratio change if you compress tiny records individually instead of batching them?It collapses. Compression exploits redundancy within the data it can see, and a 200-byte record gives the encoder almost no history, while every output still pays for a header and trailer — small enough inputs can grow. Batching records into larger blocks before compressing is the fix, and a trained zstd dictionary is the alternative when records must stay individually addressable.
- Does the compression level a file was written at slow down reading it?For the deflate codecs, essentially no: decompression cost tracks the data, not the effort spent searching for matches, so a level-9 gzip file reads about as fast as a level-6 one. That is why lowering the level is such a cheap fix on the write side. Across *codecs* it is a different story — lzma decompresses far more slowly than gzip or zstd, and that cost is paid on every read.
- What does zlib.crc32 add if gzip already checksums its own contents?It checksums the payload independently of the container. gzip's CRC32 proves the compressed member was not corrupted in transit; recording `zlib.crc32` of the *uncompressed* bytes in a manifest lets you verify that what came out of a decompress equals what went in, across a re-compression or a format change. Neither detects deliberate tampering — a CRC is an accident detector, not a signature.
saying these in an interview costs you the question
- Names a favourite codec without measuring the real data
- Assumes level 9 is meaningfully smaller than level 6
- Ignores decompression cost for frequently read archives
- Thinks a higher level also slows down reading gzip data
- Reaches for processes when the codecs release the GIL
- Compresses tiny records individually and expects a good ratio