skip to content

A service that accepts uploads slowly fills its disk with temporary files. How does multipart temp-file cleanup work, and where does it fail?

level: seniorimportance: should knowfreq 44%

answer

  1. created on spill, deleted at request end
  2. unusual endings skip the hook
  3. aborts and tripped limits leak
  4. crashes orphan everything in flight
  5. age-based reaper catches the rest

basics

~20 s

A spilled part's temporary file is deleted by a hook tied to the end of the request. Disk fills on the endings that skip it: client aborts, parses stopped by a limit, handlers that move the file, and crashes.

solid answer

~50 s

When a part is spilled to disk, the parser creates a temporary file and registers its deletion against request completion, so the file normally disappears when the response finishes. The leak paths are the ones that miss that hook. A client that aborts mid-upload, or a parse stopped by a size limit, can leave a partially written file created before the request reached the stage that arranges cleanup. A handler that renames or moves the file out from under the framework leaves a deletion targeting a path that no longer exists, or worse, a permanent file nobody tracks. A crash or a forced restart abandons everything in flight. The fixes are operational: point the temporary directory at a known, monitored volume, alert on free space rather than on upload errors, and run a reaper that removes files older than the longest request you allow.

go deeper

for a junior

Know that an upload spilled to disk leaves a real file, and that it is deleted when the request finishes rather than lingering by design.

for a middle

Explain the hook tied to request completion, and name the endings that bypass it: client aborts, a tripped size limit, a handler that moves the file, a crash.

for a senior

Answer operationally — a dedicated monitored volume, free-space and file-age alerts, an age-based reaper, and load tests that abort and over-send on purpose.

for a principal

Decide it once for the platform: whether upload paths get a temporary volume with a quota, what the reaper guarantees, and how that bound relates to the size limits every service is allowed to set.

## The normal lifecycle When a part crosses the in-memory threshold, the parser needs somewhere to put the remaining bytes. It creates a file in a configured temporary directory and streams into it, handing the handler a handle backed by that file. From that instant the file is a resource with an owner and an expected end: 1. The parser creates the file and begins writing. 2. It registers a cleanup action tied to the completion of the request. 3. The handler reads the handle, and typically copies or moves the content to permanent storage. 4. The response completes, the registered cleanup runs, the file is deleted. That path works. Every leak is a path that skipped a step, and the useful skill is enumerating them. ## Where it actually fails - **The upload aborts mid-parse.** The client disconnects, or the network drops, after the file was created and before the request completed normally. Whether cleanup runs depends on whether the abort path is treated as request completion — for many stacks it is, but it is precisely the path least likely to be exercised in testing. - **A limit trips.** Parsing stops at an over-size part, and the partially written file is exactly the thing the limit was protecting against. Rejecting a request and then keeping its bytes on disk is an unhappy combination. - **The handler moves or renames the file.** Cleanup now points at a path that no longer exists — usually harmless — but if the move was a copy, or if the destination is inside the same temporary directory, the result is a file nothing will delete. - **The handler keeps the handle for later.** Reading the file after the response has completed races the cleanup, and disabling the cleanup to make that work leaves the file behind. Anything genuinely outliving the response needs its own storage and lifecycle. - **The process dies.** A crash, a deploy, a container replacement: every in-flight temporary file is orphaned, and nothing in the request lifecycle will ever revisit it. - **The directory is shared.** Several instances, or several services, writing into one volume means one service's leak becomes everyone's outage, and nobody's monitoring names the culprit. - **Deletion silently fails.** Permissions, a read-only mount, or a full filesystem can make the delete fail. If the cleanup swallows that error, the leak is invisible until the volume is full. ## Why the symptom is confusing The reported symptom — disk fills although uploads succeed — usually means the *successful* path is fine and one of the failure paths above is not. That is why upload error rates look healthy while the volume drains. Three properties make it hard to notice early: - Growth is proportional to traffic, so it is slow at first and then suddenly not. - The temporary directory is frequently a default nobody chose, often on the same volume as logs or the application itself, so the eventual failure appears somewhere unrelated. - The first visible failure is often something else entirely — a write that cannot complete, a log that cannot be appended — and the upload path is the last place anyone looks. ## What to do about it | Action | What it buys | |---|---| | Set the temporary directory explicitly | You know which volume is at risk and can size it | | Give it its own volume or quota | A leak degrades uploads instead of the whole instance | | Alert on free space and on file age | Detection before exhaustion, and a signal that cleanup is not running | | Run a reaper for files older than the maximum request duration | Catches aborts, crashes and every path you did not enumerate | | Log cleanup failures rather than swallowing them | Turns an invisible leak into a visible error | | Load-test with aborted and over-size uploads | Exercises exactly the paths that leak, which happy-path tests never touch | The reaper deserves emphasis: it is the one mitigation that does not depend on enumerating the failure modes correctly. Any file older than the longest request the service permits cannot belong to a live request, so deleting it is safe by construction — which is why it also survives a crash. ## Design choices that reduce the exposure - **Stream past the disk entirely** where the handler does not need to re-read the content — no temporary file is created, so none can leak. - **Raise the in-memory threshold** so ordinary form values do not reach the filesystem at all, leaving only genuine file parts on the risky path. - **Move rather than copy** into permanent storage where the two live on the same volume, so there is no window in which two copies exist. - **Keep the limits tight**, since the maximum possible leak is bounded by the maximum accepted request size multiplied by the requests in flight. ## The short version Deletion is tied to the end of the request, so anything that ends a request unusually — an abort, a tripped limit, a crash — is a candidate leak, and anything the handler does to the file behind the framework's back is another. Bound the blast radius with a dedicated, monitored temporary volume, and put an age-based reaper behind the whole thing.

  • Why is an age-based reaper worth running even when cleanup appears to work?
    Because it does not depend on having enumerated the failure paths. Any file older than the longest permitted request cannot belong to a live request, so removing it is safe by construction — which is also why it keeps working across crashes and restarts, where a per-request cleanup hook cannot run at all.
  • What is the worst-case disk a service can leak from uploads?
    Bounded by the maximum accepted request size multiplied by the requests in flight, accumulated over however long leaked files survive. Tight limits and a reaper shrink both factors, which is why the limits and the temporary volume should be sized together.
  • Why does the disk keep filling even though the upload error rate looks healthy?
    Because the successful path is usually the one that cleans up correctly. The leak lives on abnormal endings — client aborts, over-size rejections, crashes — which are counted elsewhere or not counted at all, so the upload success metric stays flat while the volume drains.

saying these in an interview costs you the question

  • Assumes the framework deletes temporary files on every path
  • Tests only successful uploads and never an aborted one
  • Leaves the temporary directory at whatever the default is
  • Keeps a temp-file handle to read after the response completes
  • Monitors upload errors but never the free space on the volume