skip to content

Your team rebuilds the renderer image for each environment from the same commit — what does a passing staging test then prove?

level: middleimportance: must knowfreq 60%

answer

  1. a test is evidence about the tested unit
  2. one commit, two builds, two artifacts
  3. dependencies and base contents resolve again
  4. rollback becomes a rebuild under incident pressure
  5. keep the bytes, move only the values

basics

~20 s

It proves that staging's build of that commit passed. It says nothing checkable about production's build, because the second build resolves its dependencies, base contents and any compiled-in values again, so production runs bytes no test ever touched.

solid answer

~50 s

A test is evidence about the unit that was tested. Rebuilding per environment produces a different unit for each one, so the staging result covers the staging build and the production build inherits nothing but the source revision. That gap is not theoretical: a second build resolves the dependency set and the base layers again, runs on whatever build tooling was current, and — the reason teams rebuild in the first place — compiles that environment's own values in, so the production artifact contains at least one thing no test ever exercised. The claim you want is "what was tested is what ships", and it only holds when the tested bytes move forward untouched. Build once, keep that artifact, have every environment's deployment reference that exact artifact, and change only the values supplied around it.

go deeper

for a junior

Hold on to the core sentence: a test tells you about the artifact it ran against. Two builds are two artifacts, even from one commit.

for a middle

Explain the mechanism, not the slogan — dependencies and base contents are resolved again, and the per-environment values compiled in were never exercised by any test.

for a senior

Bring the operational argument: rollback becomes a rebuild under incident pressure, and debugging loses the ability to rule the artifact out and go straight to the value set.

for a principal

Own the exception path. Decide where a per-target build is genuinely unavoidable, what compensating evidence it must carry, and how a pipeline is prevented from drifting into it by default.

## What a test is evidence about A passing test is a statement about **the unit you ran it against**. Not about the source it was built from, not about the intent behind it — about the artifact. Everything the practice of promoting one artifact buys you follows from that one sentence. So when a pipeline builds separately for each environment, the honest reading of a green staging run is: *the build that staging produced, from that commit, behaved correctly in staging.* Production is running a **different artifact**, and the only thing the two provably share is a source revision. ## Why the same commit does not mean the same bytes Teams find this counter-intuitive, because the source is pinned and the build script is the same. But a build is a resolution step, not a pure function of the source: - **Dependencies are resolved again.** What a version range, a floating label or a mirror returns can differ between two runs, sometimes by a patch level, sometimes by a rebuilt package. - **The base layers are fetched again.** If the base is referenced by a name that can move, the second build can start from a different filesystem than the first. - **Build tooling moves.** A different builder version, a different setting, a different cached layer reused or not reused. - **The environment's own values are compiled in.** This is usually the *reason* the team builds per environment, and it guarantees that at least one thing in the production artifact was never exercised by any test. - **Time and ordering leak in.** Timestamps, generated identifiers and non-deterministic ordering make byte-identical rebuilds a deliberate engineering achievement rather than something you get by default. None of these is exotic. Each on its own is enough to make the shipped artifact a thing your evidence does not cover. ## What the two models actually claim | | Rebuild per environment | Build once, move the artifact | |---|---|---| | What moves forward | the source revision | the artifact's bytes | | What the staging pass covers | staging's build only | the artifact that production will run | | Where per-environment values live | inside each build | supplied around one fixed artifact | | Rolling back means | rebuilding an old commit and hoping | redeploying an artifact that already exists | | Failure mode | "it worked in staging" with no shared unit to inspect | a value set difference, which is diffable | ## The rollback consequence, which is the one teams feel The evidential argument is the principled one; the argument that wins meetings is rollback. If production's artifact was built by production's build job, then going back a release means **rebuilding** an older commit — with the same resolution risks, at the worst possible moment, under an incident. When the artifact is the thing that moved forward, the previous artifact is still sitting there and rolling back is a deployment, not a build. ## What restores the claim 1. **Build once, in one job, and keep the output.** Everything downstream consumes that artifact rather than the source. 2. **Reference the exact artifact.** Every environment's deployment must name the identity that cannot have been repointed since the test — otherwise "the same image" is a convention rather than a fact. 3. **Move the per-environment values out.** Every value that was the reason for rebuilding becomes something supplied at start-up around the fixed artifact. 4. **Make the lower environments run the artifact, not a lookalike.** Testing a debug build and shipping a release build is the same defect in smaller clothes. ## What the honest limits are One artifact everywhere does not mean the lower environment's pass is a guarantee. It means the difference between the two runs has been reduced to the values and the surroundings — data, traffic, dependencies, scale — instead of including "and also some unknown delta in the bytes". That is a large reduction in the search space when something goes wrong in production, which is most of what the discipline buys day to day: when you are debugging, you can rule the artifact out and go straight to the value set and the environment. There are kinds of software where shipping one artifact to every target genuinely is not possible, and those cases deserve compensating evidence rather than a shrug — but they are the exception being argued for, not the default a pipeline should fall into by accident.

  • If the build is fully reproducible and both runs really do produce identical bytes, is rebuilding per environment fine?
    Then the evidential objection largely goes away, and what remains is operational. Reproducibility is a property you have to keep proving, it is only as good as the pinning underneath it, and it does nothing for rollback — recovering an old release still means a build rather than a redeploy of something that already exists. Most teams who claim it have not verified it lately.
  • What is the first thing to fix in a pipeline that builds per environment because each one needs different settings?
    Take the settings out of the build. Give the artifact safe defaults, have the deployment supply the per-stage values, and the reason for three build jobs disappears. Then collapse them into one job whose output every environment references. Doing it in the other order — merging the jobs while values are still compiled in — just moves the problem.

Crash-testing one car and then shipping a second car assembled from the same drawings a week later. The drawings passed; the car on the road never did.

saying these in an interview costs you the question

  • Believes one commit guarantees identical bytes from two builds
  • Says the staging pass covers production's separately built image
  • Plans rollback as rebuilding the previous commit
  • Tests a debug build and ships a separately built release build
  • Treats per-environment values as a reason to build per environment