skip to content

How does the gap between validating a path with Path.resolve and opening it allow an escape?

level: seniorimportance: should knowfreq 32%

answer

  1. The check and the open are two separate lookups
  2. The kernel does not remember your earlier question
  3. Names can be re-pointed; handles cannot
  4. A descriptor for the directory, not its name
  5. os.O_NOFOLLOW covers only the last component

basics

~20 s

The check and the open are two independent name lookups. Between them an attacker can replace a component with a symlink or rename a directory, so the second lookup reaches a different file than the one that was validated. Close it by opening through a directory file descriptor.

solid answer

~40 s

`Path.resolve` answers a question about the filesystem at the instant it runs; the later `open` performs a *fresh* walk of the same name. Anything that can write into a component of that path — a shared upload directory, a scratch mount, a co-tenant process — can swap a component for a symlink in between, and the open follows the new mapping. This is a time-of-check to time-of-use race, and no amount of re-checking the string removes it. The fix is to stop re-walking names: open the base directory once with `os.open`, then open members relative to it using the `dir_fd` argument together with `os.O_NOFOLLOW`, so the kernel refuses a symlinked final component. Verify support via `os.supports_dir_fd`, and where the race cannot be closed, remove the shared writability instead.

code

python · 15 lines
python
import errno, os, tempfile

with tempfile.TemporaryDirectory() as base:
    dir_fd = os.open(base, os.O_RDONLY)
    try:
        os.symlink('/etc/hostname', os.path.join(base, 'escape.txt'))
        try:
            os.open('escape.txt', os.O_RDONLY | os.O_NOFOLLOW, dir_fd=dir_fd)
            print('opened the symlink')
        except OSError as exc:
            print('refused:', exc.errno == errno.ELOOP)
    finally:
        os.close(dir_fd)

print('dir_fd supported:', os.open in os.supports_dir_fd)

go deeper

for a junior

Know the idea rather than the syscalls: a path check describes the filesystem at that moment, and the file the program opens a moment later may no longer be the same one. Name the containment check as necessary but not the whole story.

for a middle

Explain that the check and the open are two independent name lookups, and that a symlink swapped in between redirects the second. Be able to say why re-checking does not help and what a file descriptor gives you that a name does not.

for a senior

Diagnose it as TOCTOU, show the dir_fd plus os.O_NOFOLLOW pattern guarded by os.supports_dir_fd, note that O_NOFOLLOW covers only the final component, and be candid that removing shared write access to the directory is usually the cheaper and stronger fix.

for a principal

Frame it as a trust-boundary decision rather than a coding one: who can write into the base tree, whether services should accept paths from callers at all, and whether an opaque-handle indirection should be the platform default so no team has to get descriptor-walking right.

## The shape of the bug Consider a feature-flag service whose per-tenant override files live under one directory, and whose loader validates a tenant-supplied filename before reading it. The code looks correct: join, resolve, prove containment, then open. But the validation and the open are two separate journeys through the same *name*. The kernel does not remember that you asked a moment ago; it walks the components again, from scratch, using whatever they point at now. So an attacker with any write access to a component of that path — the upload directory itself, a scratch mount, a co-tenant workload sharing the volume — replaces a component between the two walks. A file that resolved inside the base becomes a symlink to `/etc` and the open reads a host file. Nothing in your Python is wrong at the language level: this is a **time-of-check to time-of-use (TOCTOU)** race, and it is a property of addressing files by name twice. The window is not narrow enough to dismiss. A 4-person team debugging this the first time invariably reasons "the two lines are adjacent, nobody can hit that" — but the attacker does not need to win one race, only to keep flipping a symlink in a loop while re-requesting, and the scheduler eventually cooperates. Assume the window is always losable. ## What does not fix it - **Re-running the check just before the open.** You have added a third name walk; the window between check two and the open is the same kind of window. - **Checking the file after reading it.** The bytes have already been read; if the response leaves the process, the disclosure has happened. - **`os.path.islink` on the candidate.** It is itself a check that can be raced, and it inspects only the final component while an intermediate directory can be the swapped one. ## What does fix it: stop using names A file descriptor is a handle to an object, not to a name. Once you hold a descriptor for the base directory, that descriptor keeps pointing at the same directory even if the *name* is renamed or replaced. `os.open` accepts a `dir_fd` argument on platforms that support the `openat` family, and the path is then resolved **relative to that descriptor**: ```python base_fd = os.open('/srv/flags', os.O_RDONLY) fd = os.open(name, os.O_RDONLY | os.O_NOFOLLOW, dir_fd=base_fd) ``` `os.O_NOFOLLOW` makes the kernel refuse — with `ELOOP` — if the **final** component is a symlink, which removes the classic swap. Because it covers only the last component, a path with subdirectories needs the same treatment per level: open each directory component in turn with `dir_fd` and `O_NOFOLLOW`, carrying the descriptor forward and closing the previous one. That is verbose, and its verbosity is a good argument for the simpler design below. Check `os.open in os.supports_dir_fd` before relying on any of it; the builtin `open` accepts an `opener` callable if you want the file-object API on top. A second, weaker pattern is **open first, verify second**: open the file, `os.fstat` the descriptor, and compare the result against the base's `os.stat` chain with `os.path.samestat`. You still read through a descriptor that cannot be swapped after the fact, and unlike a name check the verification applies to the object you are actually holding. ## The design-level answer The most durable fix is to make the race unreachable: - **Remove shared writability.** If nothing less trusted than the service can write into any component of the base path, the swap has no author. This is usually a deployment change — ownership, mode bits, a dedicated volume — and it is cheaper than perfect descriptor code. - **Do not let user text become path components at all.** Store an opaque identifier and map it, through a database or a manifest, to a path the service composed itself. There is no traversal problem when there is no user-controlled path. - **Serve through a copy.** Materialize the file into a private directory the request owns, and serve from there. ## What a senior candidate should say out loud That containment checks are still worth writing — they catch the overwhelming majority of real traversal, they are cheap, and they fail loudly — but that they are *point-in-time assertions*, and any claim that a path was validated is only as strong as the trust boundary around the directory. Name the mechanism (TOCTOU), name the tool (`dir_fd` plus `O_NOFOLLOW`, `os.supports_dir_fd`), and name the cheaper structural fix. A candidate who insists a re-check closes the window has not understood which of the two operations the attacker is racing.

  • Why does re-running the containment check immediately before the open not close the window?
    Because the open is still a separate name lookup. You have moved the window, not removed it: the attacker now races the gap between the second check and the open, which is the same kind of gap. Any defence built on checking a *name* and then using that name again is racing by construction; the fix is to stop resolving the name twice.
  • What does os.O_NOFOLLOW actually protect, and what does it miss?
    It makes the open fail with `ELOOP` when the **final** component is a symlink. It says nothing about intermediate directory components, so `a/b/c.txt` is still exposed if `b` can be swapped. Covering the whole path means opening each component in turn with `dir_fd` and `O_NOFOLLOW`, or ensuring no component of the base tree is writable by anything less trusted than the service.
  • The flag files arrive with names encoded by another service, and an encoding mismatch means the validated string and the opened string differ. Why is that the same class of bug?
    Because the value that was proven safe is not the value handed to the kernel. Whether the divergence comes from a symlink swap or from decoding a name twice with different assumptions, the check applied to one object and the open to another. Keep a single canonical value — decode once with `os.fsdecode`, validate that value, open that value — and never re-derive the path downstream.
  • Is there a fix that removes the race rather than narrowing it?
    Yes, and it is usually the right one: remove the attacker's write access to every component of the base path, or stop letting user text become a path at all. Map an opaque identifier to a path the service composes from trusted data. A race needs a second writer; a directory only the service can write into has none.

Checking the nameplate on a door and then walking through it later is safe only if nobody can move the nameplate; holding the door open is the version that cannot be swapped behind your back.

saying these in an interview costs you the question

  • Says re-checking just before the open closes the window
  • Calls the window too short to be exploitable in practice
  • Uses os.path.islink as the defence against a swap
  • Thinks O_NOFOLLOW protects intermediate path components
  • Treats a containment check as a permanent guarantee
  • Ignores who else can write into the base directory

context