skip to content

After an incident, how do you establish exactly which image content each replica was serving?

level: seniorimportance: nice to knowfreq 32%

answer

  1. configuration records intent, not content
  2. ask the running host, not the file
  3. the answer expires with the replicas
  4. record the resolved identity at deploy
  5. keep the superseded identity too

basics

~20 s

Read the content digest back from each running replica rather than from the configuration that started it. If the deployment recorded only a label, the answer exists only for as long as those hosts still hold what they resolved.

solid answer

~40 s

Configuration tells you what was asked for, and after an incident that is usually a tag — which may well have moved since. The answer lives in two other places. First, each host knows the content identity it resolved and started, so reading it back from the affected replicas is authoritative while they are still up. Second, a deploy-time record that stores the resolved digest beside the reference answers the question permanently, including for replicas that have already been replaced. With neither, the honest report says the reference was `ledger:1.4` and the content is unknown — and that sentence is usually what finally motivates a team to start recording identities.

code

json · 9 lines
json
{
  "workload": "ledger-api",
  "site": "dc-west",
  "requestedReference": "ledger:1.4",
  "resolvedDigest": "hash-of-content:9f2c4e1b7a05d3",
  "resolvedAt": "2026-04-17T09:04:12Z",
  "replicasStarted": 6,
  "supersededDigest": "hash-of-content:1a7b90c4de62f8"
}

go deeper

for a junior

Know that the file naming the image states intent; what actually ran is known by the host that started it, and only while that host is still there.

for a middle

Explain why resolving the label after the fact cannot recover the answer, since the label may point somewhere else by the time you ask.

for a senior

Get the order right under pressure: capture the content identity from the affected replicas before anything restarts, replaces or drains them.

for a principal

Make the answer a property of the delivery path rather than of whoever is on call — deploy records carrying the resolved identity, retained as long as your reviews reach back.

## What the configuration can and cannot tell you The deployment configuration is a statement of **intent**: this workload should run the image named by this reference. When the reference names a content digest, intent and content are the same thing and the question is already answered. When it names a tag, the configuration records a label, and a label is not content — it is a pointer whose meaning at any past moment was never written down. That is why the instinct to "just look at what we deployed" fails here. The file is accurate. It simply does not contain the fact you need. ## The two places the answer actually lives **The running host.** A host resolved the reference once, at start, and knows the content identity it ended up with. Reading that back is authoritative and immediate — and it is per-replica, which matters, because replicas started at different times may well disagree. **A deploy-time record.** If the delivery path stored the resolved digest alongside the requested reference when it deployed, the question has a durable answer that outlives the replicas entirely. This is the only source that still works once the failing instances are gone. ## Why the answer expires Three things destroy it, in roughly this order of speed: 1. **A restart or replacement.** Whatever replaces a failing replica may resolve the reference again, and if the label has moved it starts different content — overwriting the one copy of the evidence that existed. 2. **A later publish.** Resolving the label after the fact reports where it points *now*, which is precisely the value that may have changed. The label cannot be rewound. 3. **Cleanup on the host.** Once a host stops holding content nothing is running, reclaiming disk removes the last local trace of what was there. The practical consequence is an ordering rule: during an incident, capture content identities from the affected replicas **before** anything restarts, replaces or drains them. It costs seconds, and it is the difference between a review that names a build and one that speculates about one. ## What to record at deploy time - the **requested reference** exactly as it appeared in configuration, label and all - the **resolved content digest**, which is the fact the label cannot carry - **when** it was resolved, so a divergence between two sites is visible as a timeline rather than a mystery - **which site or cluster** resolved it, because the same reference resolved twice is the failure you are guarding against - the **superseded digest** it replaced, which gives every deployment an exact rollback target - **how many replicas** that deployment actually started, so a partial rollout is distinguishable from a complete one ## Which evidence proves what | source | what it establishes | how long it survives | |---|---|---| | deployment configuration | the reference that was requested | indefinitely, but it is a label | | resolving that label now | where the label points today | not evidence about the past at all | | content digest read from a live replica | exactly what that replica is running | until that replica is replaced | | deploy record holding the resolved digest | what each site started, and when | as long as you retain the record | | a version string the application prints | what the build claims about itself | indefinitely, but it is a claim, not an identity | The last row is worth dwelling on. A build number in a start-up log line is useful and usually right, but it is produced by the software describing itself. Two different builds can print the same string, and a rebuild that changed only the base image will often print an identical one. A content digest is derived from the bytes, so it cannot be wrong in that way. ## The order of operations during the incident The reasoning chain is short and worth having ready, because it is executed under time pressure: 1. Ask each affected replica what content it is running, and write the answers down before touching anything. 2. Compare them to each other. Identical identities rule the image out and send you to configuration, data or a dependency. Differing identities mean you are looking at drift, and the timeline of resolutions explains it. 3. Compare them to what the deploy record says was intended. A mismatch there means the fleet is not running what anyone approved. 4. Choose the recovery target as an explicit content identity — the superseded digest, or a known-good one — rather than as a label whose meaning you would have to trust twice. Done in that order, the incident review can name two builds and diff them. Done in the opposite order, it can name a label and argue.

  • Why is capturing this the first move rather than restarting the failing replicas?
    Because a restart or a replacement can resolve the reference again and start different content, destroying the only remaining evidence of what was serving. Capture the content identity from each affected replica first; it takes seconds, and it decides whether the review names a build or speculates about one.
  • What does recording the superseded identity alongside the new one buy you?
    A rollback target and a diff. Knowing exactly which content the deployment replaced makes the recovery step an explicit reference rather than a guess at "the previous release", and it lets the investigation compare two known builds instead of two labels that may both have moved since.

saying these in an interview costs you the question

  • Says the deployment configuration proves what was running.
  • Assumes the label can be re-resolved later to recover the answer.
  • Assumes every registry keeps a history of where a label pointed.
  • Reads the content identity only after replacing the failed replicas.
  • Treats a version string the application prints as the content identity.