A node runs out of disk and starts evicting workloads that barely write anything — where did the space go?
answer
- the disk is the node's, not yours
- cache from workloads long gone
- dead artifacts before live workloads
- unreferenced layers reclaimed first
- the next start pays a cold pull
basics
~20 sMost of a node's disk is consumed by things no running workload owns: image layers cached from everything ever placed there, and artifacts left by containers already removed. The node reclaims those first, and only then evicts.
solid answer
~40 sThe disk that ran short is the **node's own**, shared by every workload on it, and most of what fills it is not attributable to anything currently running. Two consumers dominate: image layers pulled and kept as a cache, including layers nothing on the node refers to any more, and the artifacts of containers that have already been removed — their writable layers and their retained output. So the node reclaims before it evicts: delete unreferenced layers, then the leftovers of removed containers, and only if that is still not enough does it start stopping live workloads by rank. That ordering is why the workloads that get stopped are so often the ones that wrote nothing: by the time live workloads are touched, everything with no live owner is already gone.
code
pseudocode · 13 lineson nodeFreeDiskBelow(reclaimLine):
deleteImageLayersNothingRefers()
if nodeFreeDisk() >= targetFree:
return # recovered without touching anything live
deleteArtifactsOfRemovedContainers() # writable layers and retained output
if nodeFreeDisk() >= targetFree:
return # still nothing live was stopped
for each w in evictionOrder(liveWorkloads):
stop(w)
if nodeFreeDisk() >= targetFree:
returngo deeper
Remember the node has one shared disk and most of what fills it is cached images, not your writes. Being stopped for space is usually not a verdict on your workload.
Explain the reclaim stages in order — unreferenced layers, then the leftovers of removed containers, then live workloads — and why each stage exists before the next.
Diagnose it: reconstruct where the space went from the gap between the node's free-space figure and what the workloads report, and name the cost the reclaim just bought you on the next start.
Decide how much cache a node is allowed to keep and who pays for the cold starts when it is dropped. That is a density-versus-start-latency trade across the fleet, not a per-workload setting.
## Which disk is short The shortage is on the **node's own storage** — the filesystem the runtime uses to hold images and each container's writable layer, and where the platform retains captured output. It is shared by every workload placed there. It is not any one workload's durable volume: storage provisioned off the node consumes the backing system's capacity, not the node's. That distinction is the whole diagnosis. A node can be out of space while every workload on it truthfully reports tiny storage usage, because the node's figure and the sum of the workloads' figures are measuring different things. ## Where the space actually went - **Cached image layers.** Every image pulled to that node is kept so the next start is fast. That cache accumulates across everything ever placed there, including layers that nothing running refers to any more. - **Artifacts of containers already removed.** A replaced container leaves its writable layer and its retained output behind until something reclaims them. - **Live containers' writable layers.** Anything a running workload writes to a path that is not a volume lands on the node's disk. - **Retained captured output.** A long-lived, chatty workload's output is held on the node for reading, and the platform applies a cap to keep it bounded. Where a workload is chatty enough, that bounded figure is still meaningful at node scale once multiplied by everything running. The first two have no live owner, which is exactly why they grow unnoticed: no workload's own reporting attributes them to anything. ## Reclaim before eviction The node agent works in stages, cheapest and least disruptive first: 1. Delete image layers nothing on the node refers to. This is free in the sense that nothing running is affected. 2. Delete the leftovers of containers that have already been removed — their writable layers and retained output. 3. Only if the node is still short, stop live workloads by rank and reclaim what they were holding. Each stage stops as soon as free space recovers, so most disk shortages never reach stage three at all. ## Why the victim is not the filler When stage three does run, the workloads stopped are ranked on what they are using against what they declared — the same ranking used for memory. That ranking has nothing to say about the cache, because the cache belongs to no one. The workload that filled the node was, in effect, the node's whole history of placements, and history cannot be evicted. So the stopped workload is genuinely innocent, and looking for a storage bug in it is looking in the wrong place. ## Disk pressure is not memory pressure | | Memory shortage | Disk shortage | |---|---|---| | Reclaimable without stopping anything | very little | often most of it | | Typical largest consumer | one live workload's working set | artifacts with no live owner | | First action | rank and stop | delete dead artifacts | | If that is not enough | the kernel kills at host scope | live workloads are evicted by rank | | Cost after recovery | none beyond the restart | the next start pays a cold pull | The difference comes from reclaimability. Memory held by a running process cannot be taken back, so pressure resolves only by stopping something. Disk is full of things nobody is using, so a node can very often recover without disturbing a single workload. ## What to do about it - Treat an eviction for space on a node as a **node** finding first. Check what reclaim freed before you look at any workload's writes. - Expect the reclaim to have a **cost on the next start**: layers deleted from the cache are fetched again, so first starts are slower and traffic to the registry rises. The cache was doing real work; the node was sized as though it were free. - Watch the gap between the node's own free-space figure and the sum of what the workloads report. A widening gap is accumulating cache, not a leaking service. - Remember that nothing prevents the same node from filling again along the same path, because the cause is placement history rather than any workload's behaviour.
- Reclaim deletes the unreferenced image layers and the node recovers. What is the cost next time?Cold starts. Every workload placed there afterwards fetches layers it would have found locally, so first starts are slower and traffic to the registry rises — and a registry that rate-limits makes that visible as failed starts. The cache was doing real work, and the node's sizing quietly assumed it.
- Why can a node be short of disk while every workload on it reports low storage usage?Because most of what is consumed is attributed to nothing running: cached layers from every image ever pulled there, and the leftovers of containers already removed. The node's figure covers the whole filesystem; each workload's figure covers only its own writes, and the two never have to agree.
- What makes a disk shortage easier to survive than a memory shortage on the same node?Reclaimability. Most of the disk is held by artifacts with no live owner, so the node can delete its way back to health without disturbing any workload. Memory held by a running process cannot be taken back, so the only lever there is to stop something.
saying these in an interview costs you the question
- Assumes the evicted workload is the one that filled the node's disk.
- Thinks cached image layers are deleted when the container using them ends.
- Counts data on an externally provisioned volume against the node's disk.
- Expects live workloads to be stopped before unused layers are deleted.
- Believes retained output stops consuming space once nobody reads it.