Your platform team wants one minimal base image mandated for every service in the estate — what do you weigh?
answer
- one stream or many
- who pays the migration
- not every workload can go empty
- sanctioned set plus an exception path
- cheap rebase decides it
basics
~20 sWeigh what the mandate buys — one patch stream, one review, one answer to "is everything rebuilt?" — against what it costs teams whose workloads need a shipped runtime, whose diagnosis assumes a shell, and who must fund migrating hundreds of existing images.
solid answer
~50 sThe case for one base is concentration: a single userland that one group keeps current, one review of what is inside it, and a fleet-wide question — is every image rebuilt on the current base? — that finally has an answer. That is worth a lot at estate scale. The case against is that it fails as a blanket rule. Workloads that need an interpreter or a runtime shipped with them cannot go near-empty, teams that diagnose through a prompt inside the container lose that, and someone has to pay for migrating every existing image. What usually survives contact is a small sanctioned set — a minimal default, one variant carrying a runtime, one full distribution for the exceptions — with a named exception path and, above all, a rebuild path cheap enough that changing the base is routine. A standard nobody can rebuild against is just a document.
go deeper
Understand why an organisation would want one base at all: fewer things to keep current, and one answer to what is inside the images it runs. The tension is that not every workload fits the same base.
Be able to name the concrete blockers — a shipped runtime, a shell-style start command, a vendor image, a library family mismatch — rather than arguing about minimal images in general terms.
Show the operational half: a sanctioned set instead of a single mandate, a written exception with an owner and a review date, and adoption measured as images rebuilt on the current base.
Sequence it. The rebuild path comes before the mandate, because the standard's whole value is that one change reaches everything, and name who funds the migration before announcing the policy.
## What a single base actually buys The value of standardising is not the megabytes. It is that a large number of independent decisions collapse into one, and a set of questions that were previously unanswerable become answerable: - **One userland to keep current**, maintained by a group whose job that is, instead of a different one per team. - **One review** of what is inside the images the organisation runs, rather than one per repository. - **A fleet-wide question with an answer**: which images are not on the current base? That question is meaningless when every team picked differently. - **Predictable failure modes.** When every image behaves the same way with no shell and no package manager, the runbook is written once. - **Cheaper onboarding.** A new service inherits a decision that has already been argued through. ## What it costs, and to whom | Who | What the mandate gives them | What it takes from them | |---|---|---| | Platform team | One thing to maintain and measure | Ownership of every exception and every migration escalation | | Service teams | A default they need not justify | Build changes, revalidation, and a workload that may not fit | | On-call | Uniform images and one runbook | The prompt inside the container they used to reach for | | Security reviewers | One userland to assess | A new, distributed set of hand-copied additions to watch | The second row is where mandates die. A migration is not free, and if the platform team declares the standard while service teams fund the work, adoption stalls at exactly the teams with the least slack — which are usually the ones running the oldest images. ## Where the mandate genuinely breaks - **Workloads that need a runtime or interpreter shipped alongside them.** The runtime and its own library needs come with it, so a near-empty base is not a candidate at all. - **Workloads that start helper processes through a shell**, or whose start command, periodic check or wrapper is written as a shell line. - **Vendor artifacts you do not build**, which arrive as an image someone else chose the base for. - **Workloads whose diagnosis genuinely needs tools in the image**, where the platform has not yet provided an alternative way to get a view inside a running container. - **Artifacts built against a library implementation family the mandated base does not ship**, which fail at start-up rather than politely. ## A shape that survives contact 1. **Sanction a small set rather than exactly one.** A minimal default, one variant carrying a runtime, one full distribution for the genuine exceptions. Three maintained bases beat one mandated base and forty unofficial ones. 2. **Make the default the path of least resistance**, so teams land on it by doing nothing rather than by complying. 3. **Write the exception path down**: a named owner, the reason, the base used instead, a review date, and the same patch obligations. An exception with an expiry is a decision; one without is a fork. 4. **Fund the migration centrally** or accept that it will not finish. Name who does the work for a team that has none to spare. 5. **Invest in the rebuild path first.** If rebuilding and redeploying every image on a newer base is a week of coordinated effort, the standard will age no matter how good it is. 6. **Measure rebuilt images, not signed-up teams.** Adoption is the count of images actually running the current base, and the lag from a base change to the estate being on it. ## The question to ask before mandating anything *Can we rebuild and redeploy every image in the estate on a new base this week?* If the answer is no, the base standard is not the first problem, because the standard's entire value is that a change to one userland reaches everything. Fix the rebuild path, then mandate; done the other way round, the mandate produces a document, a compliance spreadsheet and a set of images pinned to whatever the base was on the day each team adopted it. ## How to answer this in a room Do not argue for or against minimal images in the abstract — that argument has no end. State what the standard is *for* (one patch stream that provably reaches everything), name the two groups who pay (teams with awkward workloads, and whoever funds migration), propose the small sanctioned set with an explicit exception path, and put the rebuild path ahead of the mandate in the sequence. Then say how you would know it worked, in a number, a quarter from now. That shape — purpose, cost, mechanism, measurement — is what distinguishes a judgment answer from a preference.
- A quarter in, how do you know the standard is working?Count images actually running on the current base, not teams who agreed, and measure the lag from a base change to the estate being rebuilt on it. Track the number of live exceptions and whether any are past their review date. If adoption is high but lag is weeks, you standardised the base and not the rebuild path.
- One team genuinely cannot use the minimal base. What does a good exception look like?A named owner, the specific reason, the base used instead, a review date, and the same patch and rebuild obligations written down explicitly. The point is that the exception stays visible and expires. Without an owner and a date it is not an exception — it is a second standard nobody maintains.
saying these in an interview costs you the question
- Mandates one base with no exception path at all
- Assumes every workload can run without a shell or runtime
- Counts adoption by policy sign-off rather than rebuilt images
- Ignores who funds migrating hundreds of existing images
- Treats image size as the goal instead of patch throughput