Why should the build itself emit a component inventory and a record of how the image was made, rather than deriving both later from the finished image?
answer
- only the builder saw the inputs
- the image forgets its own run
- inspection finds what it recognises
- compiled-in code leaves no record
- key both records to the digest
basics
~20 sBecause the builder is the only place the inputs are visible. Finished bytes show what ended up in the image, not which source revision, base bytes, steps and parameters produced it, or what was used during the build and discarded.
solid answer
~40 sTwo different records come out of a build: an **inventory** of the components that went in, and a **record of how the image was made** — which source revision, which base bytes, which steps and parameters, on what kind of builder. The second is not recoverable afterwards at all, because an image carries no memory of the run that produced it. The first is only partly recoverable: an inspection of a finished image enumerates what it can recognise, but code that was vendored, compiled in or copied over from an earlier build stage leaves no package record to find, and anything consumed during the build and dropped is invisible. Emitting both from the build, keyed to the image's digest, makes them travel with the artifact instead of being reconstructed later as guesswork.
code
pseudocode · 15 linesimage = build(source = revision, base = baseDigest, parameters = params)
inventory = collect_components(
declared = resolved_dependencies(revision),
observed = files_added_by(image.layers))
howItWasMade = record(
subject = image.digest,
sourceRev = revision,
baseDigest = baseDigest,
steps = params.steps,
startedAt = clock.now(),
builderKind = runner.kind)
publish_beside(image.digest, inventory, howItWasMade)go deeper
Know that a build can publish two records beside the image it produced: what went into it, and how it was made.
Explain what is unrecoverable later — the source revision, the base bytes, the parameters and steps of the run — and why inspecting the finished image cannot supply any of it.
Argue for producing both inside the build and binding them to the image's digest, and name what an after-the-fact inspection misses: vendored or compiled-in code, and inputs dropped during the build.
Set the floor every build in the estate must meet, and accept the cost of failing a build that cannot emit its own records rather than allowing a synthesised one.
## Two records, not one A build that wants its output to be trustable later emits two things beside the image: - a **component inventory** — what is in this artifact, at what versions, from where; - a **record of how it was made** — which source revision, which base bytes, which steps and parameters, what kind of machine ran it, when. They answer different questions and have different recoverability, and conflating them is the most common confusion in this material. "What is inside" is partly observable from the artifact. "Where did it come from" is not observable from the artifact **at all**. ## What the finished bytes can and cannot tell you | Question | Recoverable from the image? | Why | |---|---|---| | Which files are in each layer | yes | that is exactly what an inspection reads | | Which digest names this image | yes | it is computed from the bytes | | Which packaged system libraries are installed | mostly | package metadata is written into the image and can be read | | Which vendored or compiled-in code is present | no | it is ordinary files with no metadata identifying it | | Which source revision produced this | no | the image records nothing about its own build | | Which base bytes it started from, which parameters ran | no | same — the run left no trace inside the artifact | | What was used during the build and discarded | no | it never reached the final image | The last four rows are the argument. An inspection of a finished image is a **downstream vantage point**: it sees the result and infers backwards. The build is the **upstream** one: it *has* the inputs in hand, because it just used them. ## Why "we will reconstruct it at release time" fails Three ways, all of them quiet: 1. **The run is gone.** By release time the machine, the working copy and the resolved references no longer exist. Whatever the record says about them is reconstructed from memory or from a pipeline configuration that may since have changed. 2. **Recognition is not enumeration.** An inspection reports components it can identify. Code that was compiled in, vendored into the tree, or copied over from an earlier build stage looks like ordinary files, so it is silently absent from the listing — and absence reads exactly like "not present". 3. **Discarded inputs are invisible.** Much of what decides an artifact's contents — compilers, generators, caches, intermediate sources — never ships inside it. Only the build ever saw them. ## Keying the records to the digest Both records are statements *about a specific artifact*, so they must name that artifact the way it is actually identified: by **digest**. Keying them to a movable name means the record can end up describing bytes nobody is running, and the pairing quietly becomes false without anything changing in either file. A digest names one artifact for as long as it exists, so the pairing stays true. The practical shape is: build produces the image, computes its digest, emits both records naming that digest as their subject, and publishes them alongside the artifact so that whoever gets the image gets its records. ## Where this ends This leaf is about **who produces the records and when**, not about what they look like or what a consumer does with them. The formats those records are written in, the assurance ladder that grades how trustworthy a build platform's record is, the signing that binds a record to an identity, and the gate that refuses to run an artifact whose records do not check out are all separate subjects with their own owners. What belongs here is the simpler and prior point: **a record that is written after the fact by a party who did not see the inputs is a reconstruction**, and a reconstruction is only as good as somebody's recollection. ## The failure mode to expect The usual estate does not have zero records; it has records of uneven provenance. Some builds emit them, some have them synthesised at release time from a template, and nothing distinguishes the two downstream. The realistic floor is to require every release build to emit both, fail the build when it cannot, and treat a hand-written record as a defect rather than a fallback — because the whole value of the record is that it came from the only place that knew.
- What breaks if the inventory is keyed to the image's tag instead of its digest?A tag can be moved to different bytes, so the record can end up describing an image nobody is running, while both files still look intact. A digest names exactly one artifact, so the pairing stays true for as long as that artifact exists. Keying to the digest is what makes the record a statement about a specific image rather than about a name.
- Which parts of a build are genuinely invisible to an inspection of the finished image?Everything about the run: the source revision, the bytes the base resolved to, the parameters, the order of steps and the machine. Plus everything consumed and discarded — toolchains, generated intermediates and caches that never reached the final image, yet decided what it contains.
saying these in an interview costs you the question
- Says an image inspection recovers everything the build knew
- Treats the inventory and the how-it-was-made record as one thing
- Assumes compiled-in dependencies appear in a component listing
- Keys the records to a movable tag instead of the digest
- Thinks the records can be hand-written at release time