After editing one source file, why did the image rebuild re-execute every build step that came after it?
answer
- reuse is a chain, not a set
- each step starts from the one below
- the parent layer is an input
- one miss ends every later match
- steps before the edit stay reused
basics
~20 sEach build step's reuse depends on the layer it starts from. The edit changed the inputs of the step that copies that file, so the step ran again and wrote a new layer - and every later step then started somewhere new and missed.
solid answer
~40 sA builder reuses a step's layer only when that step's inputs match the previous run's. The inputs are the identity of the layer beneath it, the step's own definition, and, for a step that copies files in, the content of those files. Editing the source file changed the third input for the copy step, so it executed and wrote a **new** layer with a new identity. That identity is the first input of the next step, so the next step missed as well, even though its own definition was identical - and so did everything after it. Reuse is a chain rather than a set: the first miss ends it. Steps *before* the miss are still served from the previous run.
go deeper
Be able to say that steps are reused in order, and that the first step whose inputs changed ends the reuse for every step after it.
Explain the inputs a step is matched on, and why a step whose own definition is identical still misses once the layer beneath it is new.
Tie the cascade to time in a real build: the bill is whatever expensive work sits after the first miss, so the fix is where volatile inputs are placed, not more hardware.
Decide how much of a team's discipline to spend on step ordering conventions versus simply accepting slower builds, and make the choice explicit rather than folklore.
## What a build cache actually holds An image build is an ordered sequence of **steps**. Each step that executes produces a **layer** - the set of filesystem changes that step made - and the image is those layers stacked in order. Alongside them the builder keeps bookkeeping from previous runs: a record that maps *what went into a step* to *the layer that came out of it*. When a step is about to run, the builder computes that step's inputs, looks for a matching record, and either reuses the stored layer or executes the step for real. Two things follow immediately. First, the cache is **local bookkeeping**, not part of the image: an image carries layers, not the record of how they were matched. Second, reuse is decided **per step, in order**, before the step runs - not by comparing two finished images afterwards. ## The inputs a step is matched on A step's reuse decision is computed from inputs the builder can see before executing it: - the **identity of the layer the step starts from** - its parent, i.e. the result of the step immediately before it; - the **step's own definition**, including any argument values substituted into it; - for a step that **copies files in**, the content of the files it copies (designs differ here - some builders compare content digests, others file metadata). What is *not* an input is just as important: what the step produced last time, how long it took, how big the layer was, or anything the step's command fetched from outside the build. ## Why one edit cascades Walk the sequence with a single edited source file: 1. The steps before the copy have unchanged inputs, so they are **reused**; their layers keep the same identities. 2. The step that copies the edited file now has different file content among its inputs. It **misses**, executes, and writes a brand-new layer with a brand-new identity. 3. The next step's *first* input is that parent identity - which is now new. So it misses too, **regardless of its own definition being byte-identical to last time**. 4. Step 3 repeats to the end of the build. That is the whole mechanism: one changed input invalidates everything after it, because every later step's starting point moved. | step | own definition changed? | parent layer changed? | reused? | |---|---|---|---| | before the copy | no | no | yes | | the copy of the edited file | no | no | **no - file content changed** | | install / compile after it | no | **yes** | **no** | | final packaging step | no | **yes** | **no** | The middle rows are the ones candidates get wrong. Nothing about the install step changed, and it is re-executed anyway. ## What it costs, and what it does not The cost of an edit is **whatever work sits after the first miss** - not the size of the edit and not the number of files touched. A one-character change and a thousand-line change cost exactly the same if they land in the same step. Equally, the edit does not invalidate the whole build: every step before the miss is still served from the previous run, which is why a build that re-runs *everything* usually means the miss is at or near the first step, or that the builder has no record of a previous run at all. That last case is worth recognising. On a builder that has never run this build, every lookup misses because there is nothing to match against. The build is not broken; it simply has no previous run to reuse. ## Keeping the damage local Because the miss point decides the bill, the lever is **where volatile inputs sit in the order**: - put slow, rarely-changing work early, so that the things that change many times a day sit after it; - keep each step's copied file set narrow, so that unrelated edits are not inputs to it; - expect a changed parent to be enough on its own - an unchanged step definition is no protection once something beneath it moved; - remember that a base image that was refetched and is now a different set of bytes changes the first parent identity, and therefore invalidates the entire chain. The interview answer in one line: reuse is a chain of parents, an edit breaks the chain at the step whose inputs it touched, and everything downstream of the break is executed again.
- Do the steps before the edited one re-run as well?No. Their inputs - the layer beneath them, their own definitions, and any files they copy - are unchanged, so the builder serves them from the previous run's layers. The cost of an edit is whatever work sits after the first miss, never the whole build.
- Why does the same unchanged build take full time on a builder that has not run it before?The matching record is bookkeeping left behind by previous runs on that builder. Where there is none, every step's lookup misses and every step executes. Nothing is wrong with the build; there is simply nothing to match against, which is why a first run is never a fair timing sample.
- Does adding a step near the end of the build invalidate the steps before it?No. Reuse flows forward only: a step's inputs include what came before it, never what comes after. Appending a step leaves every earlier step's inputs untouched, so they are all reused and only the new step executes.
Reusing build layers is like resuming from saved checkpoints in a game: you can pick up at any checkpoint, but only if every move before it happened exactly as before. Change an early move and every later save describes a game that no longer exists.
saying these in an interview costs you the question
- Thinks only the edited step re-runs and later ones stay cached
- Believes a step is matched on its own definition alone
- Says the builder compares the two finished images and skips unchanged work
- Claims a reused layer expires by age rather than by changed inputs
- Assumes an unchanged step definition guarantees a reused layer