skip to content

A service extracts user-uploaded archives to disk, using the member names stored inside the archive. Explain why that hands an attacker a file-write primitive, and what a safe extractor verifies that a naive one does not — beyond rejecting names that contain `..`.

level: middleimportance: should knowfreq 42%

answer

  1. member names are attacker-authored, format guarantees nothing
  2. write primitive → interpreted directory → code execution
  3. link member first, then a name that resolves through it — no `..`
  4. containment per member, after link resolution, component-wise
  5. best: code-generated names, archived name kept only as metadata

basics

~20 s

Member names are attacker-authored strings that the extractor resolves under a destination, so each one is a write of chosen content to a chosen name. Beyond ..: absolute and drive-qualified names, link members that redirect later members, and name collisions after case folding or Unicode normalisation.

solid answer

~60 s

An archive stores member names as arbitrary byte strings; the format never promised they were relative. The extractor is a resolver that joins each name under a destination, so every member is a write of attacker-chosen bytes to an attacker-influenced name — a general write primitive, which is far more dangerous than a read. Beyond rejecting `..`, a safe extractor: - refuses absolute and drive-qualified names, and platform-specific forms such as `\` separators or `\\?\` prefixes; - **never restores link members** — a member `logs` that is a symlink to `/etc`, followed by `logs/cron.d/job`, escapes with no `..` in either name, because the escape comes from resolution state an earlier member created; - resolves each destination and containment-tests it component-wise, per member, after link resolution; - treats colliding names — case folding, Unicode normalisation, trailing dots — as a rejection, since two members can silently overwrite one file; - drops archived modes and ownership. Best of all: write to code-generated names and keep the archived names only as metadata.

code

text · 5 lines
text
member 1: name="logs"           type=symlink  target="/etc"
member 2: name="logs/cron.d/x"  type=file     content=<attacker>

neither name contains ".." or a leading "/"
member 1 installs the operator; member 2 resolves through it

go deeper

for a junior

Know that names inside an archive are attacker-controlled, that this is a write rather than a read, and that containment must be checked for every member.

for a middle

Describe the link-member sequence that escapes without any .., and list the checks beyond traversal sequences: absolute names, duplicate names after case folding or normalisation, restored metadata.

for a senior

Lead with the code-generated-names design, then give the fallback containment algorithm with its assumptions, and separate resource limits from traversal as distinct obligations.

for a principal

Treat untrusted extraction as an isolation decision — a confined resolver or a sacrificial namespace owned by a single component — rather than a validation routine each service reimplements.

## Why extraction is a write primitive A read-only traversal lets an attacker choose which of your files they see. Extraction lets them choose the name **and** the content, for as many files as they can fit in the upload. That converts routinely into code execution: drop into a directory that something later interprets — a startup or scheduled-job directory, a web root, a plugin or template directory, a shell profile, a configuration file that a privileged process reads — and the write becomes an execution at whatever privilege that consumer runs with. Rank archive extraction above ordinary file reads for this reason alone. The structural cause is simple and worth saying plainly: **archive member names are attacker-authored strings in formats that never guaranteed they were relative or benign.** A zip's central directory stores a name field; a tar header stores a name field. Nothing in either format constrains those to be relative, to lack `..`, or to be distinct from one another. The extractor is a resolver, and the archive is a program written for it. ## What a naive extractor checks, and what it misses A naive extractor rejects names containing `..` and joins the rest under a destination. Here is what escapes that. **Absolute and platform-qualified names.** A member named `/etc/cron.d/job` is well-formed. So are drive-relative and drive-absolute Windows names, and names using `\` as a separator — which a POSIX-oriented check treats as an ordinary character while a Windows extractor treats as a directory separator. `\\?\` and UNC prefixes change parsing mode entirely. **Link members — the case that defeats per-name checking.** Archive formats can carry symbolic and hard links as members. Extract a member named `logs` whose type is *symlink* and whose target is `/etc`; then extract a member named `logs/cron.d/job`. Neither name contains `..`, neither is absolute, and both pass a purely lexical check. The escape emerges from the **sequence**: the first member installed an operator into the namespace, and the second member's name resolved through it. This is the sharpest illustration that traversal is about resolution, not about strings, and it is the single thing most worth being able to describe in an interview. The same trick with a hard link to a sensitive file turns a write into a write *into* that file. **Collisions that are not visible in the raw bytes.** Two members named `Config` and `config` are distinct in the archive and the same file on a case-insensitive filesystem. Two members differing only by Unicode normalisation form can collapse to one name. A trailing dot or space can be stripped by the resolver. In each case the second write silently replaces the first, which defeats any scheme that validated file *contents* on the first write. **Restored metadata.** Modes and ownership recorded in the archive are attacker-controlled. Restoring a setuid bit or a world-writable mode from an untrusted archive is a privilege problem that has nothing to do with the name at all. ## What a safe extractor does Order the controls the same way as any other resolution defence. *Strongest — do not use the archived names.* Write each entry to a name your code generates and keep the archived name as metadata in a database or manifest. This is closed-world enumeration applied to extraction: the set of names your process can create is finite and owned by you, so the entire class disappears rather than being detected. It is available far more often than people assume; most systems only need the original name for display or for a later download filename. *Next — extract through a resolver that cannot escape.* Extract inside a jail or namespace rooted at the destination, or open each target relative to a directory handle with a resolve-beneath constraint and without following links. Then the hostile name is inexpressible as an escape. *Otherwise — resolve and containment-test, per member, after link resolution*, with the comparison done component-wise against a destination captured once. Refuse to create link members at all; refuse duplicate names after case folding and normalisation; refuse absolute and platform-qualified forms; do not restore modes or ownership. Extract into a fresh directory that no other principal can write to, so the parent chain itself cannot be redirected mid-extraction. *And separately* — resource limits. Declared versus actual size, total entry count, nesting depth. A decompression bomb is an availability problem rather than a traversal, but it lands on the same code path and any real extractor needs both. ## The one-sentence version An archive is a list of writes whose destinations the attacker wrote down; per-name string checks fail because one member can change how the next member's name resolves, so containment has to be tested on each member's resolved destination — or, better, the names should never have been used.

  • Your extractor checks containment on every member's resolved destination. Why still refuse to create link members?
    Because containment is checked at extraction time and the link persists afterwards. A link that points outside the destination is itself inside the destination, so it can pass containment and then be traversed by any later consumer of that directory — your own code, a backup job, a static file server. Refusing to materialise links removes an operator from your namespace instead of arguing about it once.
  • Why does an archive-extraction traversal usually rate higher than a read-only path traversal at the same endpoint?
    Because it is a write of chosen content to a chosen name, and writes convert into execution far more readily than reads convert into anything. Landing a file in a scheduled-job directory, a web root, a plugin directory or a configuration file gives execution at the privilege of whatever consumes it, whereas a read is bounded by what the process can already see.

saying these in an interview costs you the question

  • Believing that rejecting `..` and absolute names is sufficient, which misses link members entirely.
  • Checking each member's name in isolation, when the escape is produced by the sequence of members and the resolution state earlier ones created.
  • Restoring modes, ownership or timestamps from an untrusted archive.
  • Treating decompression-bomb limits as the security control for extraction, or conversely ignoring them because traversal was handled.
  • Extracting into a directory that other principals can write to, so the parent chain can be redirected during extraction.

context