skip to content

Image Build Model

What the builder is handed, what it reuses from last time, and what gets baked into a layer for good. These choices decide image size, build time, and whether a credential ships to every puller.

on this pageshow

questions

27

A service image moves from a full distribution base to a minimal base — what goes away with it?

level: juniorimportance: must knowfreq 72%

answer

  1. convenience against footprint
  2. the base is a userland
  3. everything your artifact did not pack
  4. shell, package manager, certificate store
  5. timezone data and system libraries too

basics

~20 s

Everything the distribution shipped that the process does not carry itself: the shell and its utilities, the package manager, usually the trusted-certificate store, timezone and locale data, and most system libraries. Only the artifact and whatever the build copied in remain.

solid answer

~50 s

A base image is a working userland, and most of it exists for humans and for convenience rather than for the process. A full distribution base ships a shell and its utilities, a package manager, a trusted-certificate store, timezone and locale data, user and group files, and a wide set of system libraries. A minimal base keeps a small fraction of that; an empty base keeps none of it, so the image contains only what the build copied in. What you gain is a smaller transfer on first pull, fewer packages for someone to patch, and less surface inside the image. What you pay is that anything the process quietly relied on — certificate verification, timezone conversion, a shared library, a shell to run a wrapper line — must now be supplied deliberately, and you cannot open a prompt inside the image to find out which one is missing.

go deeper

for a junior

Be able to name what a base image supplies: a shell and utilities, a package manager, certificates, timezone data, system libraries. Then say plainly that a minimal base keeps little of it and an empty base keeps none.

for a middle

Explain the invisible dependencies rather than just the visible ones — why verification, time conversion and a dynamically linked artifact all depend on files the base happened to ship, and why removing the package manager means every addition is a rebuild.

for a senior

Show that you decide from the run-time requirement set, validate the final image in the pipeline before publishing, and have already answered how the thing gets diagnosed once there is no prompt inside it.

for a principal

Frame it as an estate trade-off: a base choice sets the patch stream, the debugging story and the migration bill for every team that adopts it, so the interesting question is which small set of bases you are willing to maintain.

## What a base image actually is A **base image** is the filesystem a build starts from: read-only layers holding a *userland* — programs, libraries and data files — onto which the build copies your own artifact. It is not an operating system in the full sense. Every container on a host shares that host's kernel; the base supplies everything *above* the kernel that a process might reach for. That is the whole of what this choice is about. Three points on the spectrum come up in practice: - **A full distribution base** — a complete working userland: a shell and its command-line utilities, a package manager, a maintained set of root certificates, timezone and locale data, user and group files, and a broad library set. Hundreds of megabytes is normal. - **A minimal base** — a few megabytes. Typically a tiny library set and little else; some variants ship a certificate bundle and a stripped shell, many ship neither. Which one you picked matters, so read what it contains rather than assuming. - **An empty base** — nothing at all. The image is exactly what the build copied in, on top of an empty filesystem. ## What departs with the strip | What the full distribution shipped | Typically on a minimal base | On an empty base | What its absence costs you | |---|---|---|---| | Shell and command-line utilities | Sometimes, stripped | No | Start commands written as a shell line, wrapper scripts and health scripts stop working | | Package manager | No | No | You cannot add anything after the build; every addition is a rebuild | | Trusted root certificates | Sometimes | No | Outbound connections fail to verify the server's certificate | | Timezone and locale data | Rarely | No | Time conversion and formatting fall back or fail | | User and group files | Sometimes | No | A numeric identity with no name behind it; some processes object | | System libraries | A small set | No | A dynamically linked artifact will not start | ## The dependencies you did not know you had The surprises are never the things you would have listed. They are the ones the distribution supplied silently: - **Trusted root certificates.** A client verifies the chain a server presents against roots it reads from the local filesystem. No store, no verification, and every outbound call over an encrypted connection is refused. - **Timezone and locale data.** Converting an instant into a named zone reads data files. Strip them and the process either falls back to a single zone or refuses the conversion. - **Name-resolution configuration and helpers.** Some processes read these from files the base supplied. - **User and group files.** Running under a numeric identity that no name maps to is legal, and some software still refuses. - **System libraries.** An artifact linked dynamically names the libraries it needs and resolves them at start-up. If the image does not hold them, it dies before its own first line runs. - **A shell.** Not only for you: anything phrased as a shell line — a start command with a pipe in it, a wrapper, a periodic check — needs one to exist. ## What you gain, honestly - **Transfer cost.** Size is paid on the first pull by each host that does not already hold those layers, and again whenever the base itself changes. On an estate of several hundred hosts that is real, but it is a one-off per host, not a per-request cost. - **Fewer things to patch.** Every package present is something whose changes someone must track. Removing packages you never call removes that work — provided you track what you added in their place. - **Fewer moving parts.** An image with one artifact and three libraries has an obvious inventory; an image with a whole distribution does not. - **Less inside the boundary.** A smaller userland is less to reach for after something goes wrong inside the container. ## Choosing between the three 1. **Start from what the artifact needs at run time**, not from what the build needed. Those are different sets, and most of the second one has no business in the final image. 2. **If the workload needs a runtime or interpreter shipped alongside it**, an empty base is not a candidate; the runtime and its own library needs come with it. 3. **If the artifact is genuinely self-contained**, a near-empty base is straightforward and small. 4. **Decide who patches what you add**, before you strip. Hand-copied files have no publisher behind them. 5. **Decide how you will diagnose it** at three in the morning, when there is no prompt to be had inside the image. ## How teams get this wrong - Treating the choice as purely a size number, and discovering the missing certificate store in production rather than in the pipeline. - Stripping the base and copying half the distribution back in, one file at a time, until the image is the same size with none of the maintenance. - Assuming every minimal base is interchangeable with every other, when their library sets differ. - Planning to install the missing piece inside the running container, which is exactly what removing the package manager prevents.

  • The start command worked on the full distribution base and does nothing on the minimal one. Why?
    It was almost certainly written as a shell line — a command with a redirect, a pipe, a variable expansion or a chained second command — and something had to interpret it. A full distribution base supplied that interpreter; the minimal one does not. Express the start command as a direct executable with its arguments listed, so nothing needs interpreting.
  • Does a smaller base always mean less transferred across the estate?
    No. A host pulls only what it does not already hold. A slightly larger base that every image in the estate shares is pulled once per host and reused, while a unique tiny base per team is pulled once per team per host. Size matters most on first pull and on base changes, not on every deployment.

saying these in an interview costs you the question

  • Thinks a minimal base only changes the image's size
  • Assumes certificate verification works with no certificate store present
  • Believes a compiled artifact always carries its own system libraries
  • Treats an empty base as the right default for every workload
  • Expects to install a missing tool inside a running minimal image
open as a page

When a build is pointed at a directory, what is packaged and sent to the builder before the first instruction runs?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Everything under the directory the build was pointed at, recursively, minus whatever the ignore rules exclude. That packaged set is the build context, and it is transferred before step one. An instruction can only read files that are inside it.

open as a page

A build-time token is baked into an image layer — who can read it once the image is published?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Everyone who can pull the image, and everyone already holding a copy of it. The value sits in a read-only layer and in the build record that travels with it, so pull access is credential access.

open as a page

After editing one source file, why did the image rebuild re-execute every build step that came after it?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Each build step's reuse depends on the layer it starts from. The edit changed the inputs of the step that copies that file, so the step ran again and wrote a new layer - and every later step then started somewhere new and missed.

open as a page

Why is a nightly batch job's image over a gigabyte when the compiled artifact it runs is 40 MB?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A single-stage build ships everything the build needed, not only what the run needs: the compiler toolchain, development headers, the whole source tree, intermediate output and package-manager caches all stay in the image's layers beside the 40 MB artifact.

open as a page

A service on a near-empty base fails certificate verification on every outbound call — why?

level: middleimportance: must knowfreq 58%

basics

~20 s

The image carries no trusted-certificate store. A full distribution base ships a maintained set of root certificates; a stripped base does not, so the client has nothing to check the presented chain against and rejects every server it meets.

open as a page

A build context included a signing key and a broad copy instruction swept it into the image — what is the real fix?

level: middleimportance: must knowfreq 51%

basics

~20 s

Keep the file out of the set that is sent. A copy instruction can only take files that were packaged, so ignore rules at the context root close the hole for every instruction and every future one; narrowing that single copy only closes it once.

open as a page

Your image build must authenticate to a private dependency feed — what does a step-scoped secret mount give you that a build argument does not?

level: middleimportance: must knowfreq 57%

basics

~20 s

A step-scoped mount exposes the value only while one step runs and leaves it out of both that step's layer and the recorded build inputs. A build argument or environment value is part of the build and ships inside the image.

open as a page

An image build slowed from forty seconds to nine minutes with its dependency list unchanged - what step ordering explains that?

level: middleimportance: must knowfreq 66%

basics

~20 s

Almost always the build copies the whole project into the image before it installs dependencies. That makes every source file an input of the copy step, so each commit invalidates it and forces a full dependency fetch on every single build.

open as a page

On an empty base image, a binary that ran fine on the build host exits immediately — why?

level: middleimportance: should knowfreq 48%

basics

~20 s

The artifact is linked dynamically against system libraries the build host supplied and the empty base does not. Nothing resolves its dependencies at start-up, so it dies before the first line of its own code runs.

open as a page

Why does a build context take two minutes to package and transfer before ten seconds of actual build work?

level: middleimportance: should knowfreq 47%

basics

~20 s

Because the build was pointed at a working tree full of things no instruction needs — installed dependencies, earlier build output, test fixtures, version-control metadata — and the packaging step takes all of it. The lever is the set that is sent, not the machine.

open as a page

A build writes a feed credential to a file and removes it in a later step — why does the published image still carry it?

level: middleimportance: should knowfreq 50%

basics

~20 s

Because layers stack rather than overwrite. The later step records that the path is hidden; the earlier layer that wrote the credential is unchanged, still transferred on every pull, and still readable by unpacking the layers.

open as a page

Which inputs does a builder compare to decide whether a build step can reuse the layer from the previous run?

level: middleimportance: should knowfreq 48%

basics

~20 s

Three: the identity of the layer the step starts from, the step's own definition including any argument values in it, and, for a step that copies files in, the content of those files. What the step produced is not an input.

open as a page

A rebuild of the same source produces a different image digest — which build inputs typically differ between the two runs?

level: middleimportance: should knowfreq 48%

basics

~20 s

Non-deterministic inputs the build recorded: embedded build timestamps, file modification times and archive entry ordering, absolute build paths and host names, and a base image whose reference resolved to different bytes. Each changes layer bytes, so the digest changes.

open as a page

When a build has a toolchain stage and a separate final stage, what actually crosses between them?

level: middleimportance: should knowfreq 55%

basics

~20 s

Only the files you explicitly copy, at the paths you name. The final stage begins from its own base with an independent filesystem, so installed packages, environment values, created users, the working directory and the entry command do not carry over.

open as a page

After moving a service image to a near-empty base, who patches the system libraries it still carries?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The image's owner does. A full distribution base carries a maintained package set that a rebase refreshes; anything hand-copied onto a minimal or empty base has no upstream feed behind it, so nothing triggers a rebuild when it changes.

open as a page

Why can a build context still carry a file whose name appears in the build's ignore rules?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Because a rule that is written is not a rule that matched. It can be anchored differently than assumed, be a pattern that matches nothing at all, be undone by a later re-including rule, or belong to a context root that this particular invocation did not use.

open as a page

You find a published image ships a valid private-feed token — what has to happen beyond fixing the build, and why?

level: seniorimportance: should knowfreq 44%

basics

~10 s

Make the token stop being accepted at whatever issued it. Publication is one-way: every copy of that image stays pullable and readable, so no registry cleanup and no rebuild can take the value back.

open as a page

A dependency-fetch step reused its cached layer for weeks and shipped stale package versions - why did the builder not re-run it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Because nothing the builder compares had changed: the step's definition was fixed, the layer beneath it was the same, and it copied no files. Reuse is a promise about a step's inputs, never about the freshness of what it fetched.

open as a page

Why should the build itself emit a component inventory and a record of how the image was made, rather than deriving both later from the finished image?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Because the builder is the only place the inputs are visible. Finished bytes show what ended up in the image, not which source revision, base bytes, steps and parameters produced it, or what was used during the build and discarded.

open as a page

An auditor rebuilds a shipped ledger image from its source revision and gets a different digest — what does that prove?

level: seniorimportance: should knowfreq 42%

basics

~20 s

By itself, almost nothing. Unless that build was already known to produce the same bytes twice, a mismatch cannot separate tampering or an undeclared input from ordinary build non-determinism, so it only becomes evidence once reproducibility has been established first.

open as a page

A staged build's final image was given a package manager and a compiler because the service failed to start — what went wrong?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The artifact needed something the final base does not ship — a system library, a certificate store, locale or timezone data, or an interpreter it was never linked statically against. Installing a toolchain to supply it restores the size, the package manager and the attack surface the staged build removed.

open as a page

Your platform team wants one minimal base image mandated for every service in the estate — what do you weigh?

level: principalimportance: should knowfreq 33%

basics

~20 s

Weigh what the mandate buys — one patch stream, one review, one answer to "is everything rebuilt?" — against what it costs teams whose workloads need a shipped runtime, whose diagnosis assumes a shell, and who must fund migrating hundreds of existing images.

open as a page

How can a build run its test suite without the test tooling reaching the shipped image?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Give the tests their own stage built on the toolchain stage. The final stage copies nothing from it, so no test runner, fixture or report reaches the shipped layers, and the build can be asked to stop at that stage when you want the tests or their output.

open as a page

What changes about a build context when it is supplied as an archive stream or a remote location instead of a local directory?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

The set stops being a walk of your disk. Whoever produces the stream or whatever the builder fetches decides the contents, so local uncommitted files never participate and the ignore-rules file becomes just another file that may or may not be applied.

open as a page

When does merging several build steps into one earn the coarser cache invalidation that comes with it?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

When the merged work always changes together anyway, or when its intermediate output must not survive as a layer of its own. Splitting pays only where one part's inputs change far less often than the other's.

open as a page

For release builds, which do you fund first with one quarter of platform effort: byte-identical reproducibility, or inventory and provenance on every build?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Emitted records first, for most estates: they can be adopted one build at a time, help immediately, and answer the questions people actually ask. Byte-identical reproducibility costs far more and is worth buying for the few releases that must be independently re-derivable.

open as a page