skip to content

A platform team proposes rebuilding every Docker image in the estate weekly against refreshed bases — what does that cost, and how do you make it safe?

level: principalimportance: should knowfreq 40%

answer

  1. New bytes, not just a new base
  2. Only one input should float
  3. The registry is not production
  4. Failures arrive as a queue
  5. Fewer bases beats faster rebuilds

basics

~20 s

A rebuild produces new bytes, so unless application inputs are pinned the base refresh smuggles in unreviewed upgrades. Make the base the only floating input, gate on each service's tests, stagger the rollout, and remember rebuilding is not deploying.

solid answer

~50 s

The policy is right in direction and dangerous in detail. A weekly rebuild means every service's shipped bytes change weekly, so the first requirement is that the **only** thing allowed to move is the base: application dependencies pinned by lockfile, package installs version-pinned, no unpinned downloads in the Dockerfile. Second, a rebuild is a release — it needs the service's own test suite and a staged rollout, because refreshed shared libraries under a natively compiled binary are a real behaviour risk. Third, rebuilding without deploying buys nothing; the registry looks healthy while hosts run last quarter's image. Fourth, expect cost: CI time, registry storage and fleet bandwidth for multi-gigabyte images, and a stream of rebuild failures from services whose builds were never reproducible. Reduce the surface first — fewer, smaller, centrally owned bases — so there are fewer refresh paths to run.

go deeper

for a junior

Recall the underlying rule the policy is built on: images get more vulnerable simply by ageing, and the fix is to rebuild them on a newer base rather than to patch inside them.

for a middle

Be able to explain what a rebuild actually changes — every layer above the base is produced again — and why unpinned dependencies make a scheduled rebuild unpredictable rather than routine.

for a senior

Argue the operational side: rebuilding is releasing, so it needs tests, staged rollout and a rollback path, and it needs a mechanism that actually reaches running hosts rather than stopping at the registry.

for a principal

Own the tradeoff between freshness and change risk across the whole estate: set cadence and exemption rules, standardise on a small set of centrally owned bases, name who owns the rebuild-failure queue, and choose metrics about base age and time-to-live rather than raw finding counts.

## What the policy is actually proposing The premise is sound: for most images, the number of known vulnerabilities is largely a function of how old the base is, and the only supported way to reduce it is to rebuild on a newer base. Turning that into "rebuild everything weekly" converts a security property into an operational programme, and the interesting part is everything that programme touches. ## Cost 1: a rebuild is a change of bytes, not a no-op If the Dockerfile contains anything unpinned — a package install without a version, a `curl | sh` installer, a dependency range resolved at build time — then the weekly rebuild silently upgrades those too. A base refresh has become an unreviewed dependency bump across the estate, and when something breaks nobody can say which of the two moved. The precondition for the policy is therefore a discipline that has nothing to do with bases: **pin everything except the base**, so the weekly diff has exactly one degree of freedom. Even with that discipline, the base itself is a real change. A payments reconciliation batch shipped as a C++ daemon links against the runtime shared libraries in the image; a refreshed base can bring a newer TLS stack or C++ runtime under a binary compiled against the previous one. Behaviour changes here are rare, and they are exactly the kind that surface in production at month-end rather than in a smoke test. ## Cost 2: rebuilding is not deploying The registry filling with fresh images is a metric that improves without anyone becoming safer. What matters is what is running, and getting there means recreating containers across the fleet on a schedule — with the usual costs of any rollout: capacity to run both versions briefly, drain and restart behaviour, and a rollback path. A weekly cadence multiplies those costs by 52. Deciding to rebuild weekly and deploy monthly is a legitimate answer, but it must be a decision, not an accident. ## Cost 3: the mechanical bill Every service's build runs, tests run, images are pushed, hosts pull. On images in the multi-gigabyte range — a runtime layer of roughly 2.3 GB is unremarkable for a fat native service — a fleet-wide weekly refresh is real bandwidth and real disk, and hosts that never reclaim untagged images will fill up on a predictable schedule. ## Cost 4: rebuild failures are a queue The first estate-wide run finds the services whose builds no longer work: a vanished upstream artefact, a pinned toolchain that has been yanked, a Dockerfile that only ever built on one engineer's machine. This is genuinely valuable — a scheduled rebuild is the best build-health canary an organisation has — but it arrives as a support queue, and it must have an owner or the policy dies in its third week. ## Making it safe - **Pin app inputs hard, float the base deliberately.** One moving part per rebuild. - **Gate on the service's own tests.** Automation opens the change (a bumped base digest or a `--pull` rebuild), CI builds and tests it, and it merges automatically only when green. Humans see failures, not routine successes. - **Stagger.** Do not rebuild and roll out everything on the same night. Canary a handful of services, and spread the estate across the week so one bad upstream base cannot break every deploy at once. - **Reduce the surface first.** A handful of centrally owned, slim golden bases refreshed by the platform, with services inheriting from them, means fewer distinct refresh paths, fewer scanner findings per image, and one place to fix a problem. Standardisation is worth more here than cadence. - **Prioritise by exposure, not by count.** An internet-facing edge service and an internal batch job do not deserve the same urgency, and a policy that treats them identically will be ignored where it matters most. - **Measure the right thing.** Track the age of the base each running service was built from, and the time from an upstream base rebuild to that base being live. Those numbers describe the pipeline you are trying to build; a raw count of findings does not. ## The judgement to state out loud Weekly is a starting hypothesis, not a law. The right cadence is the one the organisation can absorb without the rebuild queue becoming shelfware — for many estates that is fortnightly for the bulk, with an out-of-band path for a serious advisory. What is not negotiable is the shape: one floating input, tests as the gate, automation doing the routine work, a real path to running hosts, and someone who owns the failures.

  • How do you keep a scheduled rebuild from becoming an unreviewed dependency upgrade?
    Make the base the only unpinned input. Application dependencies come from a lockfile, any OS package installed by the Dockerfile carries an exact version, and nothing is downloaded from a moving URL during the build. Then a rebuild's diff has one explanation, and when a rebuilt image fails its tests you know immediately that the base is what moved.
  • Six services fail on the first estate-wide rebuild. Is the policy wrong?
    No — it just surfaced pre-existing debt. Those builds were not reproducible, which means those services had no working patch path at all and nobody knew. Treat the failures as the first deliverable: fix or retire each one, and keep the scheduled rebuild afterwards as a permanent canary that stops build rot from accumulating silently.
  • Would you rebuild every image weekly, or only those with findings?
    Rebuilding on a schedule is simpler and more honest: it keeps the path exercised, avoids depending on scanner coverage to decide what gets attention, and prevents an image from being untouched for a year. Findings-driven rebuilds still matter for urgency — a serious advisory should trigger an out-of-band rebuild — but they are a supplement to the cadence, not a replacement.
  • What do you say to a team that wants an exemption because their rebuilt image failed twice?
    An exemption from rebuilding is an exemption from ever being patched, so it should be time-boxed and owned, not granted. Usually the right answer is to fix the fragility the rebuild exposed — an unpinned input, a flaky test, an image that is far larger than it needs to be — and to help them do it, because the same fragility will block the emergency rebuild you will need from them one day.

saying these in an interview costs you the question

  • Treats a base-only rebuild as a zero-risk change
  • Measures success by images built rather than deployed
  • Lets application dependencies float during the rebuild
  • Rolls the whole estate out on one night
  • Ignores the queue of services that no longer build
  • Adds cadence without reducing the number of bases

context