skip to content

A worker unpacks each upload into a temporary directory inside the container — why declare a scratch area for that instead?

level: middleimportance: should knowfreq 55%

answer

  1. where do the intermediate files go?
  2. the image is read-only; the layer is not
  3. declare it rather than inherit it
  4. named, sized, typed, mounted
  5. bounded, but still instance-lifetime

basics

~20 s

A declared scratch area gives the intermediate files a named path, an explicit size ceiling and a chosen backing — memory or disk — that the platform can account for. The writable layer gives none of that: it is unbounded and drawn silently from shared host storage.

solid answer

~50 s

Both are thrown away with the instance, so this is not about durability — it is about **bounds and visibility**. Writing intermediates into the writable layer means an unbounded amount of shared host storage is consumed by an instance whose spec says nothing about it, and nobody finds out until the host is under pressure. A declared scratch area is written into the workload's spec: a name, a path the process is pointed at, a **backing** (memory or disk) and a **size ceiling** the workload cannot exceed. That turns an invisible, unbounded appetite into a stated requirement the platform can account for at placement and enforce at runtime, and it fails the one greedy job instead of the whole host. Size it from the largest single unit of work times the concurrency you allow, plus headroom.

code

yaml · 9 lines
yaml
workload: thumbnail-worker
scratch:
  - name: unpack
    mountedAt: /scratch/unpack
    backing: disk        # alternative: memory
    sizeCeiling: 4GiB
    lifetime: instance   # discarded when the instance is
environment:
  TEMP_DIR: /scratch/unpack

go deeper

for a junior

Know that temporary files still consume real storage on the machine, and that the right move is to write them to a path the workload asked for rather than wherever the process happens to default to.

for a middle

Explain the four things a declaration carries — path, backing, size ceiling, lifetime — and be precise that the lifetime is unchanged. This question is testing whether you confuse a bound with durability.

for a senior

Show how you derive the ceiling from the largest work unit and the concurrency, and why a bounded failure on one job beats an unbounded appetite that surfaces as a host-level problem hours later.

for a principal

The call you own is whether undeclared temporary storage is allowed at all in your estate. Making scratch a declared, reviewed resource like memory converts a recurring shared-host incident into a per-workload capacity conversation.

## The default is scratch space nobody declared A worker that unpacks an upload, writes intermediate files beside it and produces an output needs somewhere to put the middle of that pipeline. If it just writes to a temporary directory at an ordinary path, those files go into the instance's **writable layer** — the thin private area every container gets over its read-only image. That works, and it is exactly why it is a trap. The writable layer accepts writes silently, has no size of its own, and draws from storage the host shares with every other workload on it. A spec that declares two CPUs and four gigabytes of memory can, at runtime, be consuming eighty gigabytes of host disk, and nothing in the spec, the image or the workload's declared shape says so. ## What "declared scratch" means A scratch area is short-lived storage the workload **asks for**, rather than inherits. Across platforms the vocabulary differs but the declaration carries the same four things: - **A name and a mount path**, so the process can be pointed at it deliberately instead of defaulting to whatever temporary directory it finds. - **A backing**: ordinary disk on the host, or memory. The choice is a real trade-off, because memory-backed space is fast and is charged against the workload's memory budget. - **A size ceiling**, so the area cannot grow without limit. - **A lifetime**, which is still the instance's lifetime. This is the part people misread. ## The comparison, line by line | | writable layer | declared scratch area | |---|---|---| | how you get it | automatic, for every container | written into the workload's spec | | path | wherever the process happens to write | a path you choose and point the process at | | size | none of its own; shared host storage | an explicit ceiling | | backing | host storage | host storage or memory, your choice | | visible in the spec | no | yes, and can be accounted for at placement | | lifetime | the instance's | the instance's — identical | The last row is the one to say out loud in an interview. Declaring scratch buys **bounds, a path and visibility**; it buys **no durability at all**. Storage meant to outlive the instance is a different mechanism with a different question behind it. ## Sizing it The ceiling is a capacity decision, not a guess: 1. Measure the largest single unit of work — the biggest upload times its expansion factor, plus whatever intermediate forms exist at once. 2. Multiply by the number of units the instance processes concurrently. 3. Add headroom for the moment an old file has not been deleted yet and a new one is already being written. 4. Make the process **delete as it goes**, so the ceiling covers work in flight rather than the whole day's volume. A ceiling that is too low turns a rare large upload into a failed job; a ceiling far too high is a declaration that buys nothing back. Both are better than no ceiling, because a bounded failure lands on the job that caused it. ## What it does not buy you - **It does not persist anything.** A replaced instance starts with an empty scratch area, exactly as it starts with an empty writable layer. - **It does not clean up during the instance's life.** Nothing prunes files inside a running instance; if the worker leaks temporary files, a ceiling turns a slow host-wide problem into a fast local failure, which is an improvement but not a fix. - **It does not make writes atomic.** A half-written intermediate file is still observable by anything that looks; that is the worker's problem to solve with a write-then-rename discipline. - **It does not decide the backing for you.** Memory-backed scratch is fast and spends memory; disk-backed scratch is slower and spends disk. ## The habit behind the answer The reason this question is asked is that it separates engineers who think of a container as a small machine from engineers who think of it as a declared shape. In the second view, everything a workload consumes — CPU, memory, and the space its intermediate files need — is stated where a scheduler and a reviewer can both see it. Temporary files are the last resource most teams leave undeclared, and they are the one that shows up as somebody else's outage.

  • The scratch area still dies with the instance — so what did declaring it actually buy?
    Three things the writable layer cannot give: a ceiling, so one greedy job fails instead of the host filling up; a chosen backing, so you decide whether the speed is paid for in memory or disk; and a stated requirement the platform can account for when it places the workload. Durability is a separate mechanism and a separate question.
  • How would you pick the ceiling for the thumbnail worker?
    From the largest single upload times its expansion factor, times how many uploads the instance handles at once, plus headroom for the overlap between deleting one job's files and writing the next. Then make the worker delete intermediates as it finishes each job, so the ceiling covers work in flight rather than a day's accumulation.
  • The worker still writes some files to its default temporary directory. Does the declared area help there?
    Only if the process is actually pointed at it — through configuration, an environment value, or the temporary-directory setting the runtime honours. A declared area mounted at a path nothing writes to is bookkeeping; the writes keep landing in the unbounded writable layer.

saying these in an interview costs you the question

  • Thinks a declared scratch area survives replacement of the instance
  • Picks a size ceiling by guessing rather than from the largest work unit
  • Leaves intermediates in the writable layer because it works locally
  • Believes temporary files cost nothing because they get deleted later
  • Assumes the platform prunes temporary files while the instance runs
  • Declares the area but never points the process at its path