When a build has a toolchain stage and a separate final stage, what actually crosses between them?
answer
- nothing is inherited
- two independent filesystems
- only named files cross
- packages, environment, users stay behind
- ownership bits travel, accounts do not
basics
~20 sOnly the files you explicitly copy, at the paths you name. The final stage begins from its own base with an independent filesystem, so installed packages, environment values, created users, the working directory and the entry command do not carry over.
solid answer
~50 sNothing is inherited. The final stage starts from whatever base you gave it, with that base's filesystem and nothing else; the toolchain stage is a separate filesystem that the build produced along the way. The only thing that crosses is file content you name explicitly — a path out of the earlier stage's filesystem, copied to a path in this one — together with the recorded ownership and permission bits on those files. Everything else stays behind: packages installed there, environment values set there, users and groups created there, the working directory, and the command the earlier stage would have run. That is the whole point of the boundary, and it is also where people get caught: an artifact that resolved its libraries from the toolchain base has to find them again in the final base.
go deeper
The one thing to hold on to: nothing carries over automatically. A later stage gets its own base and only the files someone explicitly copied into it, at the paths they chose.
Be able to list what does not cross — packages, environment values, accounts, working directory, entry command — and explain why: the final stage is built on its own base's layers, not on the earlier stage's.
Show the failure you have actually seen: an artifact that started fine in the build stage and exited immediately in the final one, or a permission error because the copied files carried an id that means nothing in the new base. Say how you confirmed which it was.
The lead-level angle is standardisation: a shared final-stage base whose contents are known, so teams stop rediscovering the missing certificate store and the missing timezone data one incident at a time.
## Two stages, two filesystems A staged build produces more than one filesystem. The **toolchain stage** builds one: a base with a compiler, plus everything the build installed and produced. The **final stage** builds another, starting from its own base — usually a small one — and it starts empty of everything the first stage did. They are not parent and child. The final stage does not sit on top of the toolchain stage's layers; it sits on top of its own base's layers. That is why the shipped image is small: the toolchain stage's layers are not ancestors of the final image, so they are not part of what is pushed to a registry or pulled by a host. And it is why the boundary is strict: if the final stage inherited anything, it would have to inherit layers, and the size win would be gone. ## What crosses, and how 1. You name a **source path inside the earlier stage's filesystem**. 2. You name a **destination path inside the final stage's filesystem**. 3. The builder copies the file content at that path, with its recorded ownership and permission bits, into a new layer of the final stage. That is the entire mechanism. It is a file copy between two filesystems, not an import of an environment. ## What does not cross | Produced in the toolchain stage | Present in the final stage? | |---|---| | A package installed with the package manager | No — unless you copy the files it placed | | An environment value set for the build | No — set it again in the final stage if it is needed | | A user or group created during the build | No — the id is recorded on copied files, the account is not | | The working directory the build used | No — the final stage has its own | | The command the stage would run | No — the final stage declares its own | | Files written outside the path you copy | No — they stay in a filesystem that is never shipped | The practical consequences follow directly: - **Missing run-time dependencies.** The artifact linked against whatever the toolchain base provided. If the final base does not ship those files, the artifact is copied in fine and fails the moment it starts. - **Ownership and permissions.** The copy preserves the numeric owner and the mode recorded on the file, but the account behind that number may not exist in the final base. If the workload runs as a different identity, it may not be able to read or write what you copied. Set ownership and mode in the toolchain stage, or as part of the copy, rather than assuming. - **Run-time data is easy to forget.** A certificate store, timezone data, a locale set, a default config file — all of these lived in the toolchain base without anyone noticing. Copy them in, or pick a final base that ships them. - **Only the copied path exists.** If the build wrote a report, a generated schema or a second helper file somewhere else, it is not in the shipped image unless you copied it too. ## A useful side effect Because the toolchain stage is not an ancestor of the final image, whatever happened there is not in the layers you push. A dependency fetched from a private source, a large intermediate archive, an unpacked source tree — none of them are in the shipped artefact. This is a consequence of the stage boundary, and it is worth stating precisely: it means those bytes are not in the final image's layers. It is not, by itself, a secret-management strategy, and how a credential should be supplied to a build in the first place, whether it can still surface through an exported build cache or the build's own recorded history, and how it gets rotated are a different subject with its own answers. ## How to think about it in an interview The sentence that demonstrates understanding is: *the final stage inherits nothing; it starts from its own base, and only named files cross*. Everything else in this material falls out of that — why the image is small, why a start-up failure appears only after the split, why file ownership suddenly matters, and why the config file and certificate store have to be accounted for explicitly. The stage boundary is a filesystem boundary, and it is deliberately narrow.
- A credential was used in the toolchain stage to fetch a private dependency — is it in the shipped image?Not in the final image's layers, provided it was never copied across: the toolchain stage is not an ancestor of what is pushed. That is a property of the stage boundary rather than a way of handling credentials — how the value should be supplied to the build, whether an exported build cache or recorded build history can still expose it, and when it is rotated belong to the build-time secrets subject.
- Copied files are owned by a numeric id that has no account in the final base — what happens?The ownership and mode recorded on the file are preserved, so the number survives while the account name does not. If the workload runs as a different identity, reads or writes against those files can fail with a permission error at start-up. Fix it by setting ownership and mode deliberately in the toolchain stage or as part of the copy.
saying these in an interview costs you the question
- Thinks the final stage inherits packages installed in the toolchain stage
- Expects environment values from the earlier stage to still be set
- Assumes the final image contains both stages' layers
- Forgets run-time data such as a certificate store or timezone files
- Believes copying a file also brings the account that owned it