Design the image supply strategy for an air-gapped environment: build agents and runtime hosts on a network with no route to the public internet still need upstream container images, kept reasonably current. What is your approach and what are the tradeoffs?
answer
- internal registry = only reference anyone uses
- git-tracked allowlist pinned by digest
- skopeo/crane copy --all, not docker save
- carry referrers: signatures, SBOMs
- scheduled refresh + retention + GC
basics
~20 sRun an internal registry as the single source of truth. On a connected staging host, copy an explicit, digest-pinned allowlist of images with a tool like skopeo or crane, transfer them across the boundary (registry-to-registry sync or tarball), and re-tag them into the internal registry. Rewrite all references internally; schedule refreshes.
solid answer
~50 sMake the **internal registry the only registry anyone references**. Nothing internal names docker.io. The pipeline: maintain a declarative **allowlist** of images pinned by digest, checked into git. A connected staging host copies each one with `skopeo copy --all` (or `crane cp`) — `--all` matters so the whole multi-arch index and its referrers come across, not just one platform. Move the payload across the boundary either as an OCI layout directory / tarball on approved media, or via a one-way replication path between two registries where policy allows. On the inside, push into the internal registry under a stable internal name. Tradeoffs to state explicitly: **freshness versus control** (a scheduled refresh with a vetting/scan gate, not ad hoc requests); **storage growth** (retention and garbage collection are a first-class problem); **provenance** (record source digests, and carry signatures/SBOM referrers across so verification still works inside); and **process cost** — the boundary crossing must be routine and automated or teams will smuggle images.
code
bash · 8 linesskopeo copy --all \
docker://docker.io/library/alpine@sha256:beefbeefbeefbeefbeefbeefbeefbeefbeefbeefbeefbeefbeefbeefbeefbeef \
oci:/mnt/transfer/alpine:3.20
# inside the boundary
skopeo copy --all \
oci:/mnt/transfer/alpine:3.20 \
docker://registry.internal/mirror/library/alpine:3.20go deeper
Know that air-gapped environments pull from an internal registry and that images must be copied in deliberately.
Describe the copy mechanics — digest-pinned allowlist, skopeo/crane, multi-arch --all, re-tag into the internal registry.
Own the pipeline: automation, scanning gate, referrers/signature transfer, retention and GC, and lint enforcement of internal references.
Frame it as governance: single registry of record, refresh cadence and CVE policy, boundary process cost versus bypass risk, and how verification works without reachable public infrastructure.
## The governing decision In an air-gapped estate the only sane architecture is: **one internal registry is the source of truth, and every image reference inside the boundary points at it.** Everything else — how bits cross the boundary, how often — is mechanics. Getting this wrong shows up as `FROM alpine:3.20` in some team's Dockerfile that works on a laptop and fails in the secure build. ## Curate an explicit allowlist Maintain a declarative list in version control: image name plus **digest**, plus who owns it and why it is needed. Digest pinning gives you a reproducible sync (you copy exactly what you reviewed) and an auditable record. Tools like `skopeo sync` consume such a YAML list directly. The list is also the natural gate for scanning and approval — an image enters the boundary because someone signed off on that digest. ## Copying: use registry-native tools, not save/load `docker save` / `docker load` works but is the weakest option: it goes through the daemon, historically flattens to a single platform, and drops registry-side artifacts. Prefer **`skopeo copy --all docker://upstream/img@sha256:… oci:/mnt/media/img:tag`** or `crane`/`oras`, which speak the registry API and preserve: - the **multi-arch index** (`--all`), so arm64 nodes are not stranded; - **referrers** — signatures, SBOMs and attestations attached to the image, which are separate objects and are *not* implied by copying the image; - the original **digests**, so what you verified upstream is what lands inside. Two transport shapes exist. **Staged media**: copy into an OCI layout directory, checksum it, move it on approved storage, import on the inside. **Guarded replication**: a data diode or approved one-way path between an outside registry and the inside one, which is far more automatable where policy allows it. ## Naming and rewriting references Decide the internal namespace convention up front — e.g. `registry.internal/mirror/library/alpine:3.20` preserving upstream path structure so rewrites are mechanical. Enforce it: a CI lint that rejects Dockerfiles and manifests referencing external registries is cheap and prevents the slow leak of exceptions. Where possible, also mirror the *convenience* of upstream (same tags) so developer muscle memory transfers. ## Freshness Air gaps push people toward 'copy it once and never touch it', which is how you end up running a base image with two years of unpatched CVEs. Schedule the sync: a recurring job re-resolves each allowlisted tag upstream, diffs against the pinned digest, scans the candidate, and raises a change for approval. The cadence is a policy decision — weekly for bases, immediate for critical CVEs — but it must exist and be owned. ## Storage and lifecycle Every synced image is stored forever unless you decide otherwise. Set retention (keep N versions of each base, plus anything referenced by a running deployment), and actually run registry **garbage collection**, remembering that untagged-but-referenced objects such as signatures can be destroyed by naive 'delete untagged' rules. Size the volume for the multi-arch reality: copying `--all` multiplies footprint by the number of platforms you keep. ## Verification inside the boundary If you verify signatures before deployment, the verification material must also cross: the signature/attestation objects and the trust roots. Sigstore's public transparency log is unreachable inside, so you either bundle verification material for offline verification or operate internal signing. Whatever the choice, decide it during design — discovering it during rollout is what turns a signing programme into a permanent exception list. ## Tradeoffs to name out loud Control and reproducibility improve dramatically; agility drops. The boundary crossing becomes a queue, and queues create pressure to bypass. The single strongest predictor of success is that the sync is **automated, scheduled and self-service enough** that following the process is easier than circumventing it.
- Why prefer `skopeo copy --all` over `docker save` and `docker load`?`docker save` goes through the daemon and historically exports a single platform, so arm64 or other architectures silently disappear and the multi-arch index is lost. It also ignores registry-side artifacts such as cosign signatures and SBOM referrers, which are separate objects attached to the image. `skopeo` speaks the registry API directly, preserves digests and the whole index, and needs no daemon.
- How do you keep the internal registry from growing without bound?Pair a retention policy with actual garbage collection: keep a bounded number of versions per base plus anything a live deployment references, then run the registry's GC to reclaim unreferenced blobs. Be careful with blanket 'delete untagged manifests' rules — signatures and attestations can be stored as separate objects and are easy to destroy accidentally, which then breaks verification.
It is customs for images: an approved manifest of goods, a single port of entry, a stamped record of what came in — and a scheduled shipment, because closing the port entirely just breeds smuggling.
saying these in an interview costs you the question
- Letting internal Dockerfiles keep upstream `FROM` references 'because the mirror will handle it'.
- Using `docker save`/`load` and losing multi-arch indexes or attached signatures.
- Copying by mutable tag rather than digest, so the audit trail cannot say what actually crossed.
- Treating the sync as a one-time event with no refresh cadence or CVE path.
- Ignoring storage growth and garbage collection until the registry volume fills.