A long-lived host runs out of disk after months of pulls — what does an image reclamation pass free?
answer
- unused first, oldest first
- high mark starts it, low mark stops it
- shared blobs survive the deletion
- frees only what nothing else references
- reclaiming discards the warm cache
basics
~20 sIt removes images nothing is using, usually least-recently-used first, until free space passes a low-water mark. It frees only the blobs no remaining image still references, so deleting a large image can reclaim surprisingly little.
solid answer
~40 sImages accumulate: every pull a host has ever done stays in its content store until something removes it. A reclamation pass watches disk against a high-water mark, and when it is crossed it lists images no container still references, ranks them by how long since they were used, and deletes until usage drops under a low-water mark. The part people get wrong is what deletion frees. Blobs are shared by digest, so removing an image releases only the blobs that no remaining image references — delete an image whose base is shared with three others and you free its top layers and nothing else. The second-order cost matters too: reclamation throws away the warm cache, so the next start of that image is a full cold transfer against the source.
code
pseudocode · 16 linesif diskUsed(store) < highWaterMark:
return # pressure has not been reached yet
candidates = []
for image in store.images:
if image.referencedByAnyContainer: # running, or stopped and retained
continue
if image.pinned: # protected base the fleet depends on
continue
candidates.add(image)
sortByLastUsedAscending(candidates)
for image in candidates:
store.delete(image) # drops only blobs nothing else references
if diskUsed(store) <= lowWaterMark:
breakgo deeper
Know that pulled images stay on the host until something deletes them, and that a reclamation pass removes only images nothing is currently using.
Explain the high- and low-water marks and why shared blobs mean freed space is decided by references, not by the deleted image's size.
Show the operating judgment: size the store for the working set, protect the shared base, and read repeated reclamation as an undersized store paid for in pull traffic.
Treat host disk as a funded cache of the registry and decide the estate-wide balance between disk cost and cold-start latency under burst.
## When a reclamation pass runs Nothing removes an image because it stopped being used. Unless a host is rebuilt, its content store only grows, and a host that has served many revisions over months holds every one of them. Reclamation is a deliberate pass, typically driven by disk pressure: 1. Disk usage crosses a **high-water mark**. 2. The pass lists the images in the store and excludes anything still referenced. 3. It ranks the remainder, usually by how long since each was last used. 4. It deletes from the oldest end until usage falls below a **low-water mark**. 5. It stops — it does not empty the store, because leaving the recently used images in place is the whole point. The two marks matter separately: the high one decides *when* the pass runs, the low one decides *how much* it takes. A gap that is too narrow makes the pass run constantly; too wide and each pass evicts images you were about to need. ## What it will not touch - **An image any running container references.** Deleting the layers a live process is reading from is not an option. - **An image a stopped-but-retained container still needs**, where the platform keeps the stopped instance around for inspection. - **A deliberately protected set**, where the platform supports pinning — typically a base image the whole fleet depends on for fast scale-out. Everything else — old revisions, one-off debugging images, images from a workload that has since moved elsewhere — is fair game, and on a long-lived host that is usually most of the store. ## Why deleting a large image frees little Blobs are stored once per digest and referenced by every image that lists them. Deletion is therefore reference-counted at blob level, not image level. Take a host holding an image of **800 MB expanded**, of which **750 MB** is a base also listed by another image present on the same host. Reclaiming the first image removes its manifest and drops its references, but the 750 MB base is still referenced by the second image, so it stays. **Freed: about 50 MB from an 800 MB deletion.** The practical consequences: - Freed space is not predictable from image sizes; it depends on what else the host holds. - Two hosts can free wildly different amounts deleting the same image. - A pass can appear to have done nothing, run again immediately, and delete far more the second time, once the last referencing image goes. - Conversely, deleting one small unique image can release a large chain if it was the only thing referencing it. ## The trade the thresholds encode | Setting | Effect on disk | Effect on pulls | |---|---|---| | aggressive marks, small store | plenty of headroom, little risk of a full disk | more cold pulls, more source load, slower starts | | generous marks, large store | disk pressure, and failed transfers when it fills | most starts are warm; scale-out is cheap | There is no free setting. Disk on a host is a **cache of the registry**, and every megabyte reclaimed is a megabyte that may have to cross the network again. Two other facts sharpen the choice: expanded on-disk size exceeds the compressed transfer figure, so the store grows faster than download numbers suggest; and a full disk does not merely block pulls, it degrades everything else the host is doing. ## How it interacts with a scale-out This is where the two halves of distribution meet. A pass that runs during a quiet period, sized to free a lot, can evict the shared base every workload on that host stands on. The fleet notices nothing until the next burst — at which point hosts that would have been partly warm are fully cold, and all of that transfer lands on the source at once. So the operational rules are worth stating plainly: - **Protect the base** the fleet shares, by pinning it or by re-warming after a pass, because it is the layer that makes everything else cheap. - **Size the store for the working set**, not for one image: enough for the images a host genuinely cycles through, plus the shared base. - **Watch the pass, not just the disk.** A host reclaiming repeatedly is telling you its store is undersized for what it runs, and the cost is being paid in pull traffic and start-up latency somewhere else.
- Why can two hosts free very different amounts by deleting the same image?Because each frees only the blobs nothing else on that host references. A host whose other images share the base releases just the unique top layers, while a host where that image was the sole user of its chain releases the whole thing. The image's size is the same; the overlap is not.
- What should a reclamation pass never remove?Anything a running container references, and anything a retained stopped container still needs for inspection. A sensible pass also protects a pinned base that the fleet's scale-out depends on, since reclaiming it converts cheap partly warm starts into full cold transfers.
- How does reclamation interact with a burst of scale-out later?It sets how cold the burst is. A pass that evicted the shared base leaves every host fetching the full image from the source at once, so a quiet-hours cleanup can be the direct cause of a slow, throttled scale-out the next morning.
Clearing a shared parts drawer: scrapping one machine only frees the parts nothing else still needs, so the space you get back is far smaller than the machine you removed.
saying these in an interview costs you the question
- Expects deleting a large image to free its full expanded size
- Thinks reclamation removes images containers are still using
- Believes an unused image disappears on its own after it stops
- Sizes host disk from compressed download figures
- Treats a small image store as free rather than as repeated transfers