skip to content

Image Format & Distribution

What an image is before anything runs it: read-only layers, a digest naming the exact bytes, and a document listing them. Most 'it worked on my machine' failures are really failures here.

on this pageshow

questions

18

Why can a deployment that names an image by tag start different bytes next week, when a digest reference cannot?

level: juniorimportance: must knowfreq 74%

answer

  1. one name is chosen, one is derived
  2. labels move, content does not
  3. hash taken over the bytes themselves
  4. one byte differs, the name differs
  5. record the digest, not the label

basics

~20 s

A tag is a label a publisher attaches and can re-attach to different content at any time, while a digest is computed from the image's own bytes — change one byte and it is a different digest, so nothing can republish under the old one.

solid answer

~40 s

A tag is a mutable pointer. The publisher decides what it points at and can repoint it later without anyone downstream editing a line of configuration, so `ledger:1.4` is a statement of intent, not of content. A digest is the opposite: it is a cryptographic hash taken over the image's own bytes, so nobody assigns it — it falls out of the content. Any change at all, whether a new file, a rebuilt base or a different layer order, produces a different digest, and the old digest still names the old bytes. That is why a reference by tag can resolve to different content next week, while a digest reference either resolves to exactly those bytes or does not resolve at all.

go deeper

for a junior

Be able to say which of the two names a person chose and which one the bytes produced, and give one sentence on why only one of them can ever drift.

for a middle

Explain the mechanism itself: the hash is taken over the image's own content, so different bytes necessarily yield a different name and cannot reuse the old one.

for a senior

Show where this bites in production. A deployment record that names only a label cannot say what was running during last night's incident, and cannot be rolled back to it.

for a principal

Weigh the cost you take on. Recording and pinning content identities everywhere means fixes stop arriving by themselves, so the standard you set has to include how those identities get updated.

## Two ways to name one image An image is a set of read-only layers plus a small document that lists them and records the defaults a runtime should use. Once it is published, people need a way to refer to it: in a deployment file, in a change ticket, in a conversation between two teams. Two kinds of name are in circulation, and they behave nothing alike. A **tag** is a label. Somebody chose the string, attached it to a particular image, and published that association. Nothing about the image's content produced the string, and nothing prevents the publisher from attaching the same string to different content an hour later. A **digest** is a **cryptographic hash computed over the image's own bytes** — over the document that lists the image's layers and its runtime defaults, which in turn names each layer by that layer's own hash. Nobody assigns it. It is the output of a function whose only input is the content. Throughout this explanation, *digest* means the one that names the whole image, not an individual layer's digest. ## Why a content-derived name cannot drift Content addressing simply means the name is computed from the thing named. Three consequences follow, and together they are the whole argument for recording a digest: 1. **Change anything and the name changes.** Rebuild on a patched base, edit one line of a start-up script, add a file: the bytes differ, so the hash differs, so it is a different digest. There is no operation that publishes different content under an existing digest. 2. **The old name keeps working.** Publishing new content does not disturb any existing digest; it adds one. A digest written down last March still names exactly what was written down, whatever has been published since. 3. **A receiver can check it locally.** A host that asked for content by digest can hash what arrived and compare the result. If they match, it holds the bytes that name refers to. That third point is about **integrity only**. It establishes that these are those bytes. It says nothing about who built them, or whether they are fit or permitted to run — those are separate questions with separate owners. ## What a tag is actually for Tags are not a design error. They carry the meaning a digest cannot: - a tag is readable and memorable, while a digest is a long opaque string nobody quotes from memory - a tag expresses intent — the release a team calls `1.4`, or whichever build is currently current - a tag survives a rebuild, which is precisely the moment the content changes - a tag lets a publisher ship a fix without every consumer editing configuration - several tags may point at one image, and a perfectly good image may carry no tag at all - a tag tells you nothing about its content until somebody resolves it, and the answer is only true at that instant Every one of those is a feature for publishing and a hazard for deployment. The property that lets a publisher push a fix to everyone without coordinating is the same property that lets two machines reading one string end up running different software. ## The two names side by side | | tag | digest | |---|---|---| | who chooses it | a publisher | nobody — it falls out of the bytes | | can it mean something else later | yes, at any moment | no | | what it identifies | whatever it points at right now | one exact set of bytes | | what it tells you a year later | which label was in use | exactly what ran | | readable by a person | yes | not usefully | | checkable by the receiver | no, there is nothing to compare against | yes, by hashing what arrived | ## Why this matters to a deployment Picture a payments ledger running in two data centres, deployed from configuration that names a tag. Each site turns that label into content at the moment it starts its workload. If the tag was repointed between those two moments, the sites run different builds from byte-identical configuration — and nothing in the system disagrees, because the configuration, the change ticket and the dashboards all record the label, and none of them records the content. Three practical losses follow, and they are why recording a digest is the standard advice for anything you might later have to account for: - **You cannot say what ran.** "We were on `1.4`" names a label whose meaning at the time is no longer recoverable. - **You cannot roll back to it.** "Restore what we had" needs a target, and the label now points somewhere else. - **You cannot reproduce it.** Fetching the same tag on a laptop gets whatever it means today. Pinning is not free: a frozen reference also stops absorbing fixes on its own, and something has to propose new digests. But the guarantee it buys is unusually clean — the reference resolves to exactly those bytes, or it does not resolve.

  • If a publisher copies the identical image into a second registry, does the digest change?
    No. The digest is computed from the content, so the same bytes carry the same name wherever they are stored, and a host can confirm it by hashing what it received. What does change it is a rebuild: building again from the same source usually produces different bytes — timestamps, ordering, a refreshed base — and therefore a different digest, even though nothing anyone wrote changed.
  • Can one image carry several tags at once, and does that create several images?
    One image, several labels. Tags are pointers into a set of content-addressed objects, so many of them can point at the same digest — a release label, a rolling `current` label and a date label may all resolve to identical bytes. The image is not duplicated and its digest is unchanged; only the number of names pointing at it differs.
  • Does a moving tag ever make sense in a deployment reference?
    Yes, where picking up the newest build automatically is the point and the blast radius is small: a development environment, a scratch cluster, an internal tool nobody depends on. It stops making sense the moment you need to state what ran, or need two places to be running the same thing.

A tag is a label taped to a shelf: someone can move it to a different box tonight and it still reads the same. A digest is a fingerprint of what is in the box — you cannot move it onto other contents, and it changes the moment the contents do.

saying these in an interview costs you the question

  • Thinks a version-numbered tag is immutable because it looks specific.
  • Says a digest is just a shorter generated alias the publisher assigns.
  • Believes fetching the same tag twice always yields the same image.
  • Thinks copying identical bytes to another registry changes the digest.
  • Claims a digest proves the image is safe and approved to run.
open as a page

Why does the same image reach one host in seconds and take minutes on a host that never ran it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A pull moves only the layer blobs a host is missing. A host that already holds most of them reads a manifest and little else, while a host with an empty content store has to transfer every layer and expand it to disk.

open as a page

An image is a stack of read-only layers merged into one view — what does a process see when two layers hold the same path?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A union filesystem stacks the layers and serves one merged tree: for any path the highest layer that mentions it wins, so an upper layer's copy hides the identical path in every layer beneath it. The lower bytes remain in the image.

open as a page

What happens between naming an image reference and having its layers on a host that never held them?

level: middleimportance: must knowfreq 62%

basics

~20 s

The reference is parsed and resolved at the registry to a manifest and its digest, the client proves it may read that repository, the manifest is read for the blobs it lists, and only the digests the host lacks are fetched, checked and expanded.

open as a page

A large asset bundle is added in one layer and deleted in a later one — why doesn't the image shrink?

level: middleimportance: must knowfreq 68%

basics

~20 s

Layers are append-only and immutable. The later layer records a deletion marker that hides the path, but the bundle's bytes stay in the earlier layer, so they are still stored, still transferred to every host that fetches the image, and still counted in its size.

open as a page

What does an image manifest list, and what is the configuration blob it names?

level: middleimportance: must knowfreq 62%

basics

~20 s

A manifest is a small document listing an image's layer blobs in order, each by digest and size, plus a descriptor for one configuration blob. The configuration blob is the document holding the image's runtime defaults and the identities of its unpacked layers.

open as a page

A telemetry ingester pulls fine on build workstations but fails on fleet machines of another processor architecture — why?

level: middleimportance: must knowfreq 58%

basics

~20 s

The reference almost certainly resolves to a single-platform artifact. It matches the workstations it was built on, and the fleet machines find no manifest for their platform, so the pull is rejected before any layer transfers. The fix is coverage, published under the same reference.

open as a page

What does an image index contain, and how does a host choose which manifest inside it to pull?

level: juniorimportance: should knowfreq 48%

basics

~20 s

An image index is a small document listing one entry per platform, where a platform is a processor-architecture and operating-system pair. Each entry names a manifest by digest. The pulling host matches its own platform against those entries and fetches only that manifest.

open as a page

Two data centres run the same service from the same image tag yet behave differently — how is that possible?

level: middleimportance: should knowfreq 62%

basics

~20 s

A tag is a movable label, so the two sites resolved it at different moments and the publisher repointed it in between — identical configuration, two different images, and nothing in the deployment record disagrees.

open as a page

When a host fetches an image, who presents the read credential, and what breaks first once it expires?

level: middleimportance: should knowfreq 48%

basics

~20 s

The host-side component performing the pull presents it — not the workload, which does not exist until the image is on disk. When it expires, running copies keep serving and every cold host, replacement or new revision fails to obtain the image.

open as a page

Why is a container's first write to a large inherited file costlier than the writes that follow it?

level: middleimportance: should knowfreq 52%

basics

~20 s

The image's layers are read-only, so the union filesystem cannot modify the file in place. On the first modification it copies the file up into the container's writable top layer and applies the write there; every later write hits that copy directly.

open as a page

What runtime defaults does an image's configuration blob carry, and what happens when a workload spec sets the same fields?

level: middleimportance: should knowfreq 45%

basics

~20 s

The configuration blob declares defaults for the process: the command and its arguments, the user, the working directory, environment entries, and declared ports as metadata. A workload spec that sets the same field wins at start-up, and the image is unchanged — same digest, same defaults for the next caller.

open as a page

If every deployment pins its image by digest, what does that guarantee and what does it now cost the team?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Pinning by digest guarantees every host starts exactly the reviewed bytes and that the deployment record doubles as a record of what ran. The cost is that fixes no longer arrive on their own — something must propose and roll new digests.

open as a page

Sixty cold hosts pull the same large image at once and the source starts throttling — what do you change?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A burst of identical cold pulls is being rate-limited at the source, so scale-out stalls behind the slowest transfer. Cut demand rather than retry harder: a nearer cache, pre-warmed hosts, a smaller image, capped concurrency, and backoff with jitter.

open as a page

A long-lived host runs out of disk after months of pulls — what does an image reclamation pass free?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It removes images nothing is using, usually least-recently-used first, until free space passes a low-water mark. It frees only the blobs no remaining image still references, so deleting a large image can reclaim surprisingly little.

open as a page

After an incident, how do you establish exactly which image content each replica was serving?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Read the content digest back from each running replica rather than from the configuration that started it. If the deployment recorded only a label, the answer exists only for as long as those hosts still hold what they resolved.

open as a page

Ten images on one host each report 800 MB, yet the host stores far less — why?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Layers are stored once per host and referenced by every image that uses them. A per-image size is the sum of that image's own layers, so adding the ten figures counts every shared layer ten times over.

open as a page

An index entry claims the fleet's platform, the pull succeeds, yet the process dies instantly — what happened?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The platform fields on an index entry are a publisher's claim, and nothing verifies them against the blobs. Selection matched a false claim, so the layers that arrived hold binaries the machine cannot execute. The image's own configuration blob declares the truth.

open as a page