Path traversal is usually taught as "strip ../ out of the filename". Give the definition of the bug class that also explains why an archive extractor, a URL path passing through a reverse proxy, and a symlink planted in an upload directory all exhibit it — and say why filtering the incoming string is the wrong frame.
answer
- name = program, resolver = interpreter
- authorise the resolved identity, not the submitted string
- operator set is open-world and resolver-defined
- symlink = operator stored in the namespace, not in the string
- checked artifact ≠ used artifact
basics
~20 sA name is a small program in a resolver's language, and .., absolute roots and symlinks are its operators. The bug is authorising the submitted string while operating on whatever object the resolver returns. Filtering guesses at an open-world, resolver-defined operator set.
solid answer
~50 sDefine the class by the resolver, not by a character sequence. A name is a short program in a naming language, and something evaluates it: a filesystem, a URL router, an archive extractor, a class-path or template loader. `..`, a leading `/` or a drive letter, a symlink, `%2e%2e` after one more decode pass — all operators of that language. The vulnerability is that the code authorises the **submitted string** and then acts on the **object the resolver returned**, with nothing forcing those to be the same object. Filtering is the wrong frame because the operator set is open-world and resolver-defined. POSIX treats `\` as an ordinary filename character while Win32 accepts it as a separator and also honours `C:file` and `\\?\`; a proxy and an origin can disagree about whether `%2f` was a separator; an extractor can be steered by a symlink member containing no `..` at all. The invariant: authorise the resolved identity, produced by the resolver that will perform the operation.
go deeper
Be able to state the one-line invariant: the decision must be about the object the resolver returns, not about the string that arrived. Know that .., a leading / and symlinks are all ways to change which object a name denotes.
Explain the check-versus-use gap concretely and give at least two resolvers besides the filesystem that have the same defect. Say why rejecting .. is a heuristic rather than a guarantee.
Drive the definition to design consequences: enforcement lives adjacent to the sink in the sink's grammar, resolution should happen once, and a name's safety cannot be carried across a boundary as an attribute.
Frame it as an interpreter-boundary class whose interpreter has external mutable state, and use that to argue where in an architecture a namespace decision may legitimately be made — and where a shared edge sanitiser is structurally the wrong place.
## The class, defined A path is not data in the way a person's surname is data. It is a tiny program written in a naming language, and some component — call it the **resolver** — evaluates that program to produce an object: a file, a route, a stored blob, a template. A filesystem is a resolver. A URL router is a resolver. An archive extractor that joins each member name onto a destination directory is a resolver. So is a class-path loader, a template `include` resolver, or an object-store client that later materialises keys onto a disk. Every naming language has **operators** — tokens that change which object the name denotes rather than contributing literal characters to it. The obvious ones are `..` (move up one level) and a leading `/` or `C:` (restart from a root). The less obvious ones are just as real: a **symlink** is an operator stored in the namespace rather than in the string; `%2e%2e` becomes an operator the moment another decode pass runs; on some resolvers a trailing dot or space is silently stripped, changing which name you actually addressed. So the definition: **path traversal is the failure to make the security decision about the object the resolver will actually return, made instead about the name the caller submitted.** It is the same shape as every injection class — a boundary where something the developer treated as inert data is evaluated as instructions — with one twist that matters: in SQL the interpreter is a query parser, and here the interpreter is a *name resolver whose state lives outside your process*. ## Why the naive frame fails, in three independent ways **1. You check a different artifact than the one you use.** Between the validator and the syscall there is usually at least one transformation: a decode, a join, a normalisation, a case fold. Each transformation can reintroduce an operator into a string that had none when you looked at it. A check that runs before a transformation is a check on a different string. **2. The operator set is open-world and belongs to the resolver, not to you.** You cannot enumerate what you must reject, because "what counts as an operator" is a property of a component you did not write and may swap out. Three specific divergences make this concrete. POSIX filesystems accept `\` as an ordinary character in a filename; Win32 accepts it as a path separator, and additionally honours drive-relative names like `C:file`, UNC and `\\?\` prefixes, and reserved device names. HTTP intermediaries differ on whether an encoded separator is decoded before routing, so a proxy and an origin can disagree about which resource was requested by the same request line. Archive formats store member names as arbitrary byte strings and never promised they were relative, so an absolute member name is well-formed to the format and hostile to the extractor. **3. Resolution is dynamic.** The same string can denote different objects at different moments, because the namespace is mutable and shared. A symlink created between your check and your open changes the answer without changing the name. This is why "the string is clean" can never be a durable property. ## The same defect across four resolvers - *Filesystem:* `../../` climbs out of the intended directory. - *Archive extractor:* a member named `logs` that is a symlink to `/etc`, followed by a member `logs/cron.d/job`, escapes with no `..` anywhere. The escape emerges from the sequence and from state the earlier member created. - *HTTP path in front of a proxy:* two resolvers see the same request line and disagree about its decoded form, so the authorisation decision is made about one resource and the fetch happens on another. - *Symlink in a user-writable upload directory:* the attacker never supplies a traversal string at all; they supply an operator that lives in the namespace, and your later "safe" name resolves through it. That these four look like different bugs to a scanner and identical bugs under this definition is the whole value of the definition. ## What the invariant buys you Once you state it as *authorise the resolved identity, produced by the resolver that will perform the operation*, the design consequences fall out. Enforcement has to sit adjacent to the sink, in the sink's own grammar — a shared `sanitizePath()` at the HTTP edge is validating a language that is not the one the syscall speaks. There should be exactly one resolution, and the decision and the operation should be made about the same artifact; the strongest form is to resolve once to a handle and then never name the object again. And where the platform can constrain the resolver itself so the escape is inexpressible, you stop needing to model the operator set at all. Sibling sinks — a shell's argument parser, a SQL parser, an HTML context — share the abstract shape but not the operator set, and each needs its own sink-local enforcement. Naming that shared shape is fine; assuming one filter covers all of them is exactly the mistake this class punishes.
- If `..` is not really the operator, what is?The resolver is. `..` is only an operator because some resolver assigns it that meaning; in an object store's key namespace the same two characters are a literal component and denote nothing special. Conversely a symlink is an operator that never appears in the caller's string at all. That is why safety is a relation between a name and a specific resolver, and never a property of a string on its own.
- Does that definition also cover the case where the attacker stays inside the intended directory?Yes, and it exposes a second gap. Containment says the resolved object lies under your base; it says nothing about whether this caller may have that object. A multi-tenant store where every tenant's files sit under one base directory can be perfectly traversal-free and still leak across tenants. Containment is a namespace property; authorisation is a per-principal decision and has to be made separately.
Filtering .. is like vetting a mailing address by scanning it for the word "forward". The post office, not your scanner, decides where the letter ends up — and it can be redirected by instructions stored at the destination rather than written on the envelope.
saying these in an interview costs you the question
- "We strip `../` from the input, so we're safe" — the strip runs on a pre-transformation string, misses alternate separators and encodings, and cannot see symlinks.
- Treating traversal as a string-scanning problem rather than a resolution problem, so symlink and archive variants read as unrelated bugs.
- Assuming an input that was cleaned once stays clean across joins, decodes and normalisations further down the call path.
- Believing that if the resolved path is inside the base directory, access is therefore authorised.