How do you cap a decompression bomb when a worker extracts uploaded archives?
answer
- Path checks answer a different question
- Headers state a size, never prove it
- Count as you go, abort mid-stream
- Several caps, not one: count, member, total
- A killable child on a disposable directory
basics
~20 sExtraction filters guard paths, not volume, so enforce your own limits: a member count cap, a total decompressed-byte budget counted as you stream members out, and a per-member cap. Treat the sizes in archive headers as hints, and unpack inside a disposable, bounded scratch directory.
solid answer
~40 sNeither `tarfile`'s data filter nor `zipfile`'s name sanitising bounds output size, so a few hundred kilobytes of deflated zeros can still fill a disk. The cap has to be yours, and it has to count **bytes you actually decompressed**: iterate `ZipFile.infolist()` or `TarFile.getmembers()`, open each member as a stream with `ZipFile.open()` or `TarFile.extractfile()`, read in fixed chunks, add each chunk's length to a running total, and abort the moment the total crosses the budget. `ZipInfo.file_size` comes from the archive's own headers and can lie, so use it only as a cheap early reject. Add a member-count cap and a nesting rule, run the work in a child process you can kill, and unpack into a fresh temporary directory that is deleted on any failure.
code
python · 31 linesimport io
import zipfile
MAX_TOTAL = 50 * 1024 * 1024
CHUNK = 64 * 1024
def measure(zf: zipfile.ZipFile) -> int:
total = 0
for info in zf.infolist():
if info.file_size > MAX_TOTAL: # cheap early reject; the header may lie
raise ValueError(f"declared size too large: {info.filename}")
with zf.open(info) as src:
while chunk := src.read(CHUNK):
total += len(chunk) # count what really came out
if total > MAX_TOTAL:
raise ValueError("uncompressed size cap exceeded")
return total
buf = io.BytesIO()
with zipfile.ZipFile(buf, "w", zipfile.ZIP_DEFLATED) as zf:
zf.writestr("frames.raw", b"\0" * (200 * 1024 * 1024))
print("archive bytes:", buf.tell())
with zipfile.ZipFile(buf) as zf:
print("declared file_size:", zf.infolist()[0].file_size)
try:
measure(zf)
except ValueError as exc:
print("refused:", exc)go deeper
Know the phrase and the shape: a small archive can expand enormously, and neither tarfile nor zipfile stops it. Recall that the sizes recorded in an archive are claims, and that untrusted archives belong in a temporary directory.
Explain the streaming check - open each member, read in chunks, keep a running total, abort at the budget - and why the header size is only a first-pass filter. Be able to name the other caps: member count, per-member size, nesting.
Demonstrate operating this: caps derived from measured legitimate jobs, a killable child process with an operating-system file-size limit, a disposable scratch directory, and metrics on rejections so a badly chosen cap is visible before users report it.
Own it as a platform concern - one shared extraction helper rather than per-service loops, a documented budget policy, capacity thinking about how many concurrent jobs a node can absorb at the cap, and the decision about whether untrusted extraction belongs on shared workers at all.
## The filter and the budget are different problems Path containment and volume containment get conflated constantly. `tarfile`'s `data` filter, which Python 3.14 applies by default, decides *where* a member may land and *what kind of thing* it may be. `zipfile`'s extractor cleans member names. Neither of them asks how many bytes come out. A compressed archive of a few hundred kilobytes expands to hundreds of megabytes of zeros without anything unusual in its structure, and nesting one archive inside another multiplies that again. Concretely: a worker that generates thumbnails from uploaded image batches accepts an archive per job. Every path check in the standard library passes. The archive contains one member, `frames.raw`, that decompresses to eighty gigabytes, and the box the worker shares with everything else runs out of disk. Nothing was exploited; the service simply did what it was asked. ## Count what you decompressed, not what you were told The single load-bearing habit is that the running total must come from bytes you have actually produced: ```python def read_capped(src, budget, chunk=65536): total = 0 while data := src.read(chunk): total += len(data) if total > budget: raise ValueError("uncompressed size cap exceeded") yield data ``` `ZipFile.open()` and `TarFile.extractfile()` both give you a binary stream you can drive this way, so you never hand the whole member to the filesystem in one call. The declared sizes — `ZipInfo.file_size`, `ZipInfo.compress_size`, `TarInfo.size` — are attacker-controlled header fields, useful as a cheap early rejection ("this member claims more than the whole job budget, stop now") and useless as a guarantee. The two formats differ slightly here: for tar, the header size is what extraction will read, so a per-member `TarInfo.size` check is a real bound on that member; for zip, the central-directory size is just a claim, and only the read loop bounds it. Neither bounds the *total*, which is what actually fills the disk. ## The set of limits worth enforcing * **Total decompressed bytes** per job, checked incrementally. This is the one that matters. * **Member count**, so a million-empty-file archive cannot exhaust inodes or spend an hour in `os.makedirs()`. * **Per-member size**, so one pathological member cannot consume the entire job budget. * **Nesting depth** — either refuse archives inside archives outright, or count the inner extraction against the same outer budget. * **Wall-clock and CPU**, because decompression can be slow rather than large. Express the cap as an integer comparison on a byte count. Deriving a compression *ratio* in floating point and testing it against a threshold invites rounding drift exactly at the boundary, so an archive tuned to sit on the line lands on either side depending on chunk boundaries. Keep ratio as a secondary signal for logging and alerting, never as the primary gate: legitimate data compresses extremely well too, so a ratio trigger alone both rejects real work and misses a bomb that stays under the ratio while blowing the absolute budget. ## Bound the blast radius, not just the number Limits inside the process are the first line, not the only one. Extract in a **child process** so you have a kill switch that does not depend on your own loop noticing anything, and give that process an operating-system file-size limit so a bug in the accounting still cannot fill the volume. Unpack into a fresh temporary directory on a volume sized for one job — `tempfile.TemporaryDirectory` as a context manager gives you deletion on any exception for free — validate there, then move only the artefacts you wanted. If the job fails, the whole directory goes, so a partially written extraction is never mistaken for a finished one. ## The organisational half of the answer The reason this shows up in reviews is that the limits get re-implemented per call site and drift. On a team of eleven engineers, the durable fix is one reviewed extraction helper that owns the budget, the streaming loop, the filter choice and the scratch directory, plus a lint rule or review habit that flags a bare `extractall()` on an uploaded path. The numbers themselves belong in configuration, chosen from what real jobs measure rather than guessed: instrument the decompressed size of legitimate work first, set the cap above the observed maximum with headroom, and emit a metric on every rejection so a cap that is too tight shows up as a spike rather than as a support ticket.
- Why is ZipInfo.file_size not a trustworthy cap on its own?It is a field in the archive's own headers, written by whoever produced the archive, so a malicious builder can understate it. It earns its place as a cheap first pass that rejects an obviously oversized member before you spend any CPU, but the enforcing check has to be the running total of bytes read out of ZipFile.open, because that number cannot be forged.
- Where does a compression-ratio threshold mislead you?In both directions. Legitimate data - sparse frames, logs, repeated padding - reaches ratios of a thousand to one, so a ratio gate rejects real jobs; and an attacker can stay under the ratio while still exceeding your absolute budget by shipping a larger archive. Cap absolute decompressed bytes as the gate and keep ratio as a signal you log and alert on.
- Why run the extraction in a separate process rather than a thread?Because you get a kill switch that does not depend on the extraction loop cooperating, an operating-system file-size and CPU limit scoped to that process, and a crash boundary: a native decompressor that dies takes the child with it rather than the worker. It also lets you bound memory per job, which a thread inside the same interpreter cannot give you.
saying these in an interview costs you the question
- Trusts ZipInfo.file_size as the true uncompressed size
- Says the tarfile data filter also protects against bombs
- Checks only the archive's size on disk
- Extracts everything first, then measures the directory
- Assumes a nested archive cannot multiply the output
- Relies on the disk filling up as the safety net