A build job uploads a jar and the deploy job ships whatever it downloads - how do you close that gap?
answer
- resolved by name, not by content
- who else can write that store?
- record the digest at build time
- the expected value needs its own channel
- pull by digest, not by tag or filename
basics
~20 sBind the handoff to content, not to a name. Have the build record the jar's digest, have the deploy recompute it and refuse on a mismatch, and carry the expected digest where the run's other jobs cannot rewrite it.
solid answer
~50 sThe deploy job resolves the artifact by **name** from a store that every job in the run can usually write. Anything that can write that name - a compromised parallel job, a test job running third-party fixtures - substitutes the bytes, and the deploy ships them, because a filename says nothing about content. Close it by binding the reference to content: the build computes the jar's digest and records it where the other jobs cannot touch, and the deploy recomputes the digest of what it downloaded and aborts on mismatch. Uploading a `digest.txt` beside the jar does not count - whoever overwrites one overwrites both. Better still, skip the run-local store: publish to a repository and pull **by digest**, so the reference *is* the content. Be clear what this buys - the bytes are the ones the build recorded, not that the build was trustworthy.
go deeper
Understand that downloading an artifact by name gives you whatever currently sits under that name. A digest identifies content; a filename only labels it.
Be able to describe the mechanism: a run-scoped artifact store is mutable and usually writable by every job, so the deploy consumes bytes no one re-checked against what the build produced.
Show the full fix and its weak point - digest recorded at build, recomputed at deploy, and an expected value that travels where sibling jobs cannot rewrite it - plus what a digest match does and does not prove.
Argue the platform-level move: publish to a content-addressed store and deploy by digest everywhere, so substitution becomes impossible by construction rather than something each team remembers to verify.
## The shape of the gap A typical pipeline splits build from deploy: job A compiles and packages a jar and uploads it under a name; job B downloads that name and pushes it to production. The reason to split is sound - the deploy job is the one holding production credentials, and you do not want compilation happening in the same job as those credentials. But look at what job B actually trusts. It asks the run's artifact store for `service.jar` and gets bytes. The store is a mutable, name-keyed area that, in most CI systems, **every job in the run can write**, and sometimes jobs of other runs in the same project as well. Nothing in the download tells job B which job produced those bytes or whether they changed after upload. The name is not a claim about content; it is a label. ## Who can exploit it The interesting attacker here is not an outsider. It is a *sibling job on the same run*: a matrix leg building a variant, a test job that executes third-party test fixtures or install hooks, a lint job running a plugin someone added last week. Any of those runs arbitrary code with the run's own artifact-store credentials. If one is compromised, it does not need to touch the build job or the source at all - it overwrites the uploaded jar between the upload and the download. The asset at risk is the deployed artifact and, through it, the availability and integrity of the running service. This is also why "the build job is hardened" is not an answer. The gap is the store between the jobs, not the job that filled it. ## Closing it **Step 1 - make the reference content-addressed.** The build job computes a cryptographic digest of the file it produced. The deploy job recomputes the digest of what it downloaded and compares. On mismatch it fails loudly; it never falls back to "deploy anyway". **Step 2 - protect the expected value.** This is the step people skip. If the digest travels as a second file in the same store, the attacker overwrites both and the comparison passes. The expected digest must arrive by a path the sibling jobs cannot rewrite: a job-output/metadata channel written only by the producing job, an append-only record, or a record signed by an identity the other jobs do not hold. Ask of any design: *who else can write the thing I am comparing against?* **Step 3 - prefer a store where the name is the content.** Publishing to a repository or an OCI registry and having the deploy pull by digest removes the class of problem instead of detecting it: a digest reference cannot resolve to different bytes, so substitution is not something you check for, it is something that cannot happen. The residual question becomes how the deploy job learns which digest to deploy - which is again a channel-integrity question, and a much smaller one. **Step 4 - shrink the writer set.** Fewer jobs with write access to the handoff store, distinct identities per job where the platform allows it, and no untrusted code in any job that holds those credentials. ## What this does and does not prove Get the direction right, because interviewers push here. A digest match proves **the bytes are the ones that were recorded** - an integrity property of the handoff. It does not prove the build was hermetic, that its inputs were trustworthy, or that the jar contains no vulnerable or malicious dependency. Those are different questions with different controls. Conversely, no amount of scanning the jar in the deploy job substitutes for digest binding: a scan tells you about content quality, not about whether the content is the one your build produced. ## Detection, if you did none of this The cheap forensic control is to log the digest at both ends. The build job prints the digest of what it uploaded; the deploy job prints the digest of what it shipped; both live in retained logs. Then a deployed digest that no build ever recorded is an alertable event, and after an incident you can answer "was this release the artifact we built?" without rebuilding. That is detective, not preventive, and it is worth having as a backstop even when digest verification is enforced - because it also catches the case where verification was accidentally disabled. ## A note on the tempting shortcut "Just merge build and deploy into one job" does close the handoff, and it is occasionally right for a small pipeline. But it puts production credentials in the same process as compilation, dependency resolution and any build-time code execution those bring - usually a worse trade than the one you were fixing. Splitting the jobs and binding the handoff by digest gives you both properties at once.
- The deploy job compares the jar against a digest file the build job uploaded alongside it. Why is that not enough?Both live in the same mutable store with the same writers. Anyone who can overwrite the jar can overwrite the digest file to match it, and the comparison then passes on attacker-chosen bytes. A verification is only as strong as the integrity of the expected value, so that value has to come through a channel the substituting job cannot write.
- Does verifying the digest tell you the artifact is safe to deploy?No. It tells you the bytes are the ones the build recorded - an integrity property of the handoff only. It says nothing about whether the build's inputs were trustworthy, whether the code is vulnerable, or whether the build itself was compromised. Treat it as closing one specific substitution gap, not as a general assurance about the artifact.
- How would you detect that a substitution already happened in a pipeline that never verified digests?Compare the digest recorded in each build job's logs against the digest of what was actually deployed and what sits in the registry today. A deployed digest that no build ever produced is the signal. This requires that both ends logged a digest and that logs are retained long enough, which is worth setting up even before enforcement lands.
Collecting a parcel by the name on the shelf is not the same as collecting one whose reference is a fingerprint of its contents. In the first case anyone who can reach the shelf decides what you take home.
saying these in an interview costs you the question
- Trusting an artifact because the filename matches
- Storing the expected digest in the same writable store
- Assuming only the build job can write run artifacts
- Claiming a digest match proves the build was trustworthy
- Scanning the artifact instead of binding its identity