skip to content

Your Docker images must now run on both arm64 and amd64 hosts. How do you decide which ones ship both, and keep the two from diverging?

level: principalimportance: nice to knowfreq 27%

answer

  1. Start from where each image is scheduled
  2. Three ways to produce the second architecture
  3. Building is not the same as behaving
  4. An emulated test proves start-up only
  5. Write down the deliberate exceptions

basics

~20 s

Decide per image from where it actually runs, not by default: dual-architecture for shared base images and anything developers run locally, single-architecture where a workload is pinned to one host pool. Then test each architecture on native hardware, because build success is not behavioural parity.

solid answer

~50 s

Treat it as an inventory and policy problem before a build problem. Start from where each image is actually scheduled: shared base images and anything engineers run on their laptops should be dual-architecture, while a batch job pinned to one host pool can stay single-architecture and say so explicitly. Then price the two ways of producing the second architecture — emulated builds on existing runners, which are cheap to set up and can turn a 3-minute build into a 40-minute one, versus a small pool of native builders that costs money but keeps the pipeline honest. The harder half is divergence: dependencies with no arm64 artefact, base images that lag on one architecture, and behaviour differences that only appear under load. Insist that each architecture is smoke-tested on native hardware, gate releases on a check that the published tag lists every architecture you deploy to, and keep an explicit list of the images that are deliberately single-architecture.

go deeper

for a junior

Know that one image tag can serve several architectures and that a host pulls the variant matching its CPU; the strategy questions above that are not expected of you yet.

for a middle

Be able to say what it costs to add a second architecture to a pipeline and why a build that succeeds does not prove the service behaves the same on both.

for a senior

Show the operational controls: native smoke tests per architecture, a release gate asserting the published tag covers every deployment target, and a canary when a workload first moves to new hardware.

for a principal

Own the policy and its economics — which images must be dual-architecture, how the second architecture is produced, what the divergence risks are, and when the honest answer is to stay on one architecture and record why.

## Frame it as inventory, cost, and parity "Make everything multi-platform" is the wrong default in both directions: it wastes build capacity on images that will only ever run in one place, and it gives false confidence that anything which builds for arm64 works on arm64. A defensible position has three parts. ### 1. Inventory: which images actually need it Start from where each image runs, not from a blanket rule. - **Always dual-architecture.** Shared base images and toolchain images — if the platform team's base image is amd64-only, every downstream team inherits the constraint and nobody can move. Anything an engineer runs on a laptop, since developer hardware is the most mixed estate you have. Anything scheduled onto a heterogeneous host pool where you do not control placement. - **Defensibly single-architecture.** A batch job pinned to one node group, an image wrapping a vendor binary published only for amd64, a legacy service with a scheduled decommission date. The key discipline is that this is a *recorded* decision with an owner and a reason, not an accident nobody noticed until a container failed to start. The artefact of this step is a short list, reviewed periodically, not a wiki page written once. ### 2. Cost: how the second architecture gets built Three production strategies, with genuinely different economics. - **Emulated builds** on existing runners. No new infrastructure, works today, and catastrophically slow for anything CPU-bound: a compile-heavy image whose native build takes about three and a half minutes can take three quarters of an hour emulated, and some toolchains are simply unreliable under emulation. Acceptable for images built rarely. - **Native builder nodes** — a small pool of machines of each architecture that the builder farms work out to. Fastest and most faithful, but it is a fleet to own: capacity, patching, credentials, and idle cost between builds. - **Cross-compilation** in the Dockerfile, where the toolchain stage runs natively and produces output for the target architecture. Nearly free at runtime and the cheapest of the three, but it is per-language work and not every stack cooperates; for JVM services it is almost trivial, for something with heavy native dependencies it can be a project. Most organisations end up with a blend, and the leadership call is which blend, not which single answer. The number to hold in the argument is total pipeline wall-clock and its effect on deploy frequency, not the cost of the runners in isolation. ### 3. Parity: the part that actually causes incidents A successful build proves the image exists. It does not prove the service behaves. The recurring sources of divergence: - **Missing artefacts.** A dependency with no prebuilt arm64 package falls back to building from source at install time, or fails; the failure is loud, the silent fallback to a slower or older path is worse. - **Base image lag.** A distribution or runtime image can publish one architecture days ahead of the other, so a "same tag" deploy is not the same content everywhere. - **Runtime behaviour.** Memory page size, atomics and floating-point details differ between architectures; JIT-compiled runtimes make different choices; a race that never surfaced on one CPU surfaces on the other. These appear under load, not in a start-up test. - **Performance shape.** Per-core throughput and price/performance are simply different, so capacity assumptions and autoscaling thresholds calibrated on one architecture do not transfer. The controls that address these are unglamorous: a smoke test that runs on *native* hardware for every architecture you ship, because an emulated test proves the image starts and nothing else; a release gate that inspects the published tag and fails unless every architecture you deploy to is present; and, when you first cross over, a canary on a small number of hosts of the new architecture with the same metrics you would watch for a code change. ## How to phase it A rollout that works: make the shared base images dual-architecture first, so teams are unblocked; add the release-time architecture assertion so nothing regresses; convert the highest-frequency builds to cross-compilation where the language makes it easy and leave the rest emulated; stand up a small native builder pool only when the emulated wall-clock demonstrably hurts deploy frequency. Then move workloads architecture by architecture behind a canary, with a clear rollback that puts the host pool back rather than rebuilding images under pressure. ## What to say about stopping It is equally principal to argue *against* going dual-architecture — for a small fleet on uniform hardware with no developer-laptop pressure, the build cost and the doubled test surface buy nothing. The judgement being assessed is whether you priced both sides, and whether the decision is written down where the next person will find it.

  • How do you keep CI honest that the arm64 image really works, not just that it built?
    Run the smoke test on native arm64 hardware, not under emulation. Emulation exercises the code path but distorts timing and CPU-specific behaviour, which is where cross-architecture bugs live. If native runners are not available for every pipeline, at minimum run the full suite natively on a nightly schedule and gate releases on that.
  • When would you deliberately keep an image single-architecture?
    When it wraps a vendor binary published for one architecture only, when it is pinned to a node group you control and no developer runs it locally, or when it is scheduled for decommission. The requirement is that it is an owned, recorded decision with the reason attached, so the constraint surfaces during planning rather than as a container that will not start.
  • What stops a team from silently shipping an amd64-only tag again?
    A pipeline step after the push that inspects the published tag and fails the job unless every architecture in the deployment target is listed. It is a few seconds of pipeline time and it converts a production start-up failure into a red build, which is the trade any platform team should take.

saying these in an interview costs you the question

  • Treating arm64 support as just another rebuild
  • Accepting emulated tests as proof of behavioural parity
  • Maintaining a separate tag per architecture indefinitely
  • Ignoring dependencies with no arm64 artefact
  • Assuming equal per-core performance across architectures
  • Leaving single-architecture images undocumented and unowned

context