skip to content

Ten images on one host each report 800 MB, yet the host stores far less — why?

level: seniorimportance: nice to knowfreq 34%

answer

  1. the host stores layers, not images
  2. identical content is stored once
  3. per-image totals include shared layers
  4. summing images double-counts the base
  5. no sharing without identical content

basics

~20 s

Layers are stored once per host and referenced by every image that uses them. A per-image size is the sum of that image's own layers, so adding the ten figures counts every shared layer ten times over.

solid answer

~40 s

A host does not store images; it stores **layers**, each identified by its content. When several images reference the same lower layers, those layers exist once on disk and every image points at the same copy. A per-image size, though, is reported as the total of all the layers that image references — shared ones included — because that is a property of the image, not of the host. Summing per-image sizes therefore double-counts everything held in common: ten images at 800 MB sharing a 700 MB base occupy roughly 1700 MB, not 8000 MB. Sharing is also all-or-nothing per layer and depends on the content being identical, so two images built from "the same" base share nothing if those base layers came out byte-different.

go deeper

for a junior

Take away the headline: a host keeps one copy of a layer no matter how many images use it, so several similar images cost far less together than their individual sizes suggest.

for a middle

Explain why the reported size is defined the way it is, and be able to work the arithmetic: a shared base counted once plus each image's own layers, rather than the sum of the totals.

for a senior

Show the operating judgment — count distinct layers for a host, count missing layers for a deployment, and notice when sharing has silently stopped because two base layers are no longer byte-identical.

for a principal

Make it a standard: storage and fetch cost across an estate are set by how many distinct base layers the organisation allows to exist, which is a platform decision rather than a per-team one.

## A host stores layers, not images An image is, for storage purposes, an ordered **list of references to layers** plus the metadata that describes it. The layers are the things that occupy disk. Each one is identified by its content, so a host that already holds a layer recognises it when another image references it and keeps a single copy. Ten images that begin from the same base are ten lists that happen to name the same first few entries. That is the whole economy of the format. Lower layers are read-only, which makes them safe to share; they are identified by content, which makes sharing automatic rather than something anyone configures. ## Why per-image sizes double-count A reported image size answers the question "how big is this image?", which has to mean "the total of every layer it references" — the image is not smaller because something else on this host happens to reference the same base. That number is useful for one image and misleading in aggregate, because the moment two images share a layer, both sizes include it. Worked through, with ten images of 800 MB each that all sit on one 700 MB base layer and add 100 MB of their own: | Way of counting | Figure | What it means | |---|---|---| | Sum of reported image sizes | 8000 MB | Every shared layer counted ten times | | Actual host storage | 1700 MB | 700 MB base once, plus ten × 100 MB | | Marginal cost of the eleventh image | 100 MB | Only the layers the host lacks | The gap widens with the number of images and the size of the shared base — which is exactly the direction that makes capacity forecasts built on per-image sizes wrong by a large factor. ## Two different sizes for the same layer A second, independent source of confusion sits underneath the first. Each layer has: - a **compressed transfer size** — what moves across the network when a host fetches it; - an **expanded on-disk size** — what it occupies once unpacked and ready to be stacked. These can differ by several times for compressible content, and platforms differ over which one they report where. So "the image is 800 MB" may be a transfer figure, a disk figure, or a sum of layer figures in one of those two units. Before using any such number for capacity planning, establish which of the three it is. ## When sharing silently stops Sharing is per layer and it is all-or-nothing: two layers are the same layer only when their content is identical. The consequences are sharper than they first look. - Two images described as "built from the same base" share that base only if the base layers they actually reference are byte-identical. If the base was rebuilt in between so its contents differ at all, they are different layers, and the host now stores both. - A change to a lower layer forces every layer above it to be reproduced, so a small edit deep in the stack can end sharing for the whole upper part of the image. - Because a layer is the unit of sharing, a shared layer's hidden contents are also shared: a large object buried by a later layer is at least stored only once across all the images that inherit it. The practical rule is that a fleet's storage depends on how many **distinct** layers exist across the images it runs, not on how many images it runs. Estates that pin every workload to a small set of common base layers store a fraction of what estates with a base layer per team store, at identical per-image sizes. ## How to reason about the numbers 1. For **one** image, the reported size is the right number: it is what a host with nothing needs to obtain it. 2. For a **host or a fleet**, count distinct layers once and ignore per-image totals entirely. 3. For the cost of **adding** an image to a host that already runs related ones, count only the layers that host does not already hold — often a small fraction of the image's reported size. ## Answering it in an interview Say that a host stores content-identified layers once and that images are lists of references to them, so per-image sizes include shared layers and summing them double-counts. Then show the judgment: the number that matters for a host is distinct layers, the number that matters for a deployment is the layers that host lacks, and sharing quietly disappears the moment two supposedly identical base layers are not byte-identical.

  • Two images were built from the same base, yet the host stores it twice — how?
    Because the two base layers are not byte-identical. Sharing is decided by content, not by the name the base was referenced under, so a base that was rebuilt or produced slightly differently yields a different layer. The images look related and share nothing, and every layer stacked above the differing one is distinct as well.
  • Why can a fleet's average image size badly overstate what each host needs?
    Because the average includes every image's share of the layers it has in common with the others. A host running related images stores each shared layer once, so its real usage tracks the number of distinct layers it holds. The overstatement grows with the size of the common base and the number of images per host.
  • Which number should you use to estimate the cost of adding one more workload to a host?
    Only the layers that host does not already hold. If the new image shares its base with what is already running, the marginal cost is its own upper layers alone, which is frequently a small fraction of its reported size — and the reason a second workload from the same base lands far faster than the first.

saying these in an interview costs you the question

  • Adds per-image sizes to estimate a host's disk use
  • Thinks every image keeps a private copy of its base layers
  • Treats a reported image size as what the host actually stores
  • Assumes two images from the same base always share layers
  • Believes the compressed transfer size equals the expanded on-disk size