A long-running worker's writable layer has grown to tens of gigabytes of temporary and rotated files — what actually clears it?
answer
- the image's size tells you nothing
- nothing inside prunes it for you
- same container restarted keeps its layer
- replacement discards; restart in place does not
- deleting image-provided files frees nothing
basics
~20 sTwo things clear it: the process deleting files it wrote, which releases those bytes immediately, and replacing the instance, which discards the layer whole. Restarting the same container in place keeps it, and nothing inside prunes it on a schedule.
solid answer
~50 sOnly two mechanisms return that space. The process can **delete the files it wrote**, which releases the bytes back to the host's shared pool right away — though deleting a file that came from the read-only image only hides it and frees nothing. Or the **instance can be replaced**, which discards the entire writable layer. A restart of the *same* container keeps the same layer and its contents, which is why "we restart it nightly" is often not the cleanup people think it is. Nothing inside a running instance prunes temporary files for you; the image's size tells you nothing about the layer's; and unless the platform imposes an ephemeral-storage quota, there is no ceiling on it at all. The real fix is upstream: intermediates into a sized scratch area, rotation capped by size as well as count, and deletion as each unit of work finishes.
go deeper
Remember that temporary files written inside a container keep using real space on the machine until something deletes them, and that the size of the image says nothing about how much that is.
Explain the two mechanisms that return the space — deletion by the process, and replacement of the instance — and why restarting the same container in place returns none of it.
Triage it properly: find what is consuming space inside the instance, separate files the process wrote from files that came from the image, and recognise long uptime as the reason this never appeared in testing.
The pattern to eliminate is workloads whose storage appetite is undeclared. Requiring a stated ceiling for short-lived storage, as for memory, converts a recurring shared-host incident into a per-workload capacity decision someone owns.
## Why it grows in the first place The writable layer accepts everything the process writes and has **no size of its own**. It is not derived from the image, it is not bounded by it, and it draws from storage the host shares with every other workload on that machine. Two sources dominate: - **Temporary files that are never deleted** — a unit of work that unpacks, transforms and writes an output, then moves on without cleaning up. Each one is small; a month of them is not. - **Files the application writes itself and rotates** — rotation capped by file count but not by file size is unbounded, and rotation that compresses in place still keeps every generation it was told to keep. The reason this reaches production undetected is uptime. In a test environment instances are replaced constantly, and every replacement silently resets the growth to zero. The same image on a host where an instance runs for weeks accumulates for weeks. ## What does not clear it - **A restart of the same container.** The instance is the same instance; it keeps its layer and everything in it. Platform designs differ here — a single-host runtime typically restarts in place, while a cluster scheduler usually replaces the instance somewhere else — which is why the same nightly-restart habit genuinely cleans up on one setup and does nothing on another. - **Rebuilding the image without the temporary directory.** The image was never the problem; the bytes are in a running instance's layer. - **Pulling the image again.** Content that is already present is not re-fetched, and the pull does not touch the running instance's layer at all. - **Waiting.** No timer inside the instance sweeps temporary files. ## What does clear it 1. **The process deletes files it wrote.** Those bytes are released back to the host's shared pool immediately. This is the everyday fix and the one that should be in the worker's code. 2. **The instance is replaced.** The layer is discarded whole, and the new instance starts empty. This is the blunt fix, and the one a rollout performs for free. ## Deleting from inside: two cases that behave differently This distinction separates senior answers from confident wrong ones. | the file being deleted | what deletion does | space returned | |---|---|---| | written by the process into the writable layer | removes it from the layer | yes, immediately | | came from the image's read-only content | records that it is gone, so the container stops seeing it | no — the image's bytes stay on the host for every other instance | So "I deleted the big files and disk did not drop" has two very different explanations depending on which kind of file was deleted, and asking which one is the right first question. ## Why the image's size tells you nothing A small image is a statement about what is distributed, not about what a running instance consumes. A workload built on a few tens of megabytes of read-only content can be sitting on tens of gigabytes of writable layer, because those bytes were produced at runtime and were never part of anything anybody reviewed. Teams that track image size as a proxy for footprint are measuring the half that does not vary. ## Is there any bound? Not inherently. Platforms differ in what they offer: some let a workload declare a request and a ceiling for short-lived storage the way it declares memory, some impose a quota on the layer, and some leave it entirely unbounded against the host's shared pool. Where a ceiling exists, exceeding it is a bounded failure for that workload — which is the outcome you want, because the alternative is a shared resource being consumed by whichever workload happens to be the greediest. What the platform does once the host's storage genuinely runs short is a separate subject with its own mechanics. ## Designing the problem out - Point the process at a **declared, sized scratch area** for intermediates, so the appetite is stated and bounded. - Make the worker **delete as it goes**, at the end of each unit of work rather than at shutdown — a worker that only cleans up on exit never cleans up if it is killed. - Cap application log rotation by **size as well as count**, and prefer not writing large files inside the instance at all. - Treat **uptime as a risk factor**: anything that only stays healthy because it is replaced often is a latent problem waiting for the first long-lived instance. - When triaging, ask what is actually consuming the space *inside* the instance before touching the image, because the image is very rarely the answer.
- Does anything cap how large a container's writable layer can get?Nothing inherently — it has no size of its own and draws on the host's shared storage. Platforms differ: some let a workload declare a request and ceiling for short-lived storage, some impose a quota, some leave it unbounded. Where a ceiling exists the greedy workload fails instead of the host, which is the outcome worth having.
- The same worker never showed this in the test environment. Why not?Because instances there are replaced constantly, and every replacement discards the layer and resets the growth to zero. Unbounded accumulation only becomes visible with uptime, so a test environment with frequent deploys systematically hides it. Treat long instance lifetime as its own risk factor.
- The team deleted a multi-gigabyte file from inside the container and the host's free space did not move. What happened?Most likely the file came from the image's read-only content rather than being written by the process: deleting it records that the container should no longer see it, but the underlying bytes remain on the host for every other instance. Files the process itself wrote into the layer do release their space when deleted.
saying these in an interview costs you the question
- Believes restarting the same container empties its writable layer
- Thinks a small image implies a small runtime footprint on the host
- Assumes the platform prunes temporary files inside a running instance
- Caps rotated log files by count while leaving each file unbounded
- Thinks deleting from inside the instance can never return space
- Cleans up only at shutdown, which a killed worker never reaches