skip to content

Your Dockerfile installs OS packages unpinned and curls a tarball. What does that cost you?

level: middleimportance: should knowfreq 44%

answer

  1. The Dockerfile is fixed, the answers are not
  2. Every fetch is resolved at build time
  3. Tested artefact differs from promoted artefact
  4. Version the URL, verify the checksum
  5. Pin the base image by digest

basics

~20 s

Every fetch resolves at build time against something you do not control, so two builds of the same commit produce different filesystems. The image you promote is not the image you tested, and the build breaks when upstream moves.

solid answer

~50 s

It makes the build non-hermetic: `apt-get install` takes whatever the mirror serves today, and a `latest` tarball URL is answered by the web server, not by your repository. Unlike a timestamp difference, this changes real content - the image you promote can contain a different JRE patch or a different agent than the one you tested, and rebuilding an old commit no longer reconstructs the artefact you shipped. It is also an availability risk: when a mirror retires a version or a URL moves, a Dockerfile that changed nothing stops building. The fixes, in order of value: pin `FROM` by sha256 digest, install application dependencies from a lockfile, fetch versioned URLs and verify them with `sha256sum -c` in the same `RUN`, pin OS package versions where the archive still serves them, and prefer copying a vetted artefact from the build context over fetching it.

code

dockerfile · 8 lines
dockerfile
# syntax=docker/dockerfile:1
FROM eclipse-temurin:21-jre-noble
RUN apt-get update \
 && apt-get install -y --no-install-recommends curl \
 && rm -rf /var/lib/apt/lists/*
RUN curl -fsSL https://example.com/agent/latest.tar.gz | tar -xz -C /opt
COPY target/fraud-scoring.jar /app/app.jar
CMD ["java", "-jar", "/app/app.jar"]

go deeper

for a junior

Be able to point at the lines in a Dockerfile that reach out to the network and say why their result is not fixed by the repository - package installs and downloaded archives are the usual two.

for a middle

Explain the difference between metadata drift and content drift, and name the concrete remedies: digest-pinned base, lockfiles, versioned URLs, checksum verification in the same RUN step.

for a senior

Show the operational consequence - a rollback that does not reconstruct the shipped artefact, a build that breaks when a mirror retires a version - and how you would stage the fixes on a live service.

for a principal

Own the policy: which classes of image must be hermetic, where the mirrors and artefact stores live, and who is accountable for moving pins so reproducibility does not turn into an unpatched estate.

### The shape of the problem Every network fetch inside a build is a resolution performed at build time against a system you do not control. `apt-get install curl` asks a mirror what version `curl` is today. `curl -fsSL https://example.com/agent/latest.tar.gz` asks a web server what "latest" means today. A dependency install without a lockfile asks a package index the same question. The Dockerfile is fixed; the answers are not. Two builds of the same commit, a week apart, legitimately produce different filesystems - and unlike a timestamp difference, this one changes what the image actually runs. That is the real cost, and it is worth stating in the interview before the reproducibility framing: an unpinned build means the artefact you tested in the pipeline on Monday is not the artefact you promoted on Friday. You get bugs that reproduce only in the newer image, rollbacks that do not roll anything back because rebuilding the old tag re-resolves to new packages, and a supply-chain surface where anything the upstream index serves lands in your image unreviewed. ### Why it also breaks the build itself Non-hermetic builds fail in a second way: they stop working. A mirror retires an old point release, a download host changes a URL, a rate limit trips, a corporate proxy blocks an endpoint - and a Dockerfile that built fine for a year fails on a commit that changed nothing. Every fetch is an availability dependency of your build, not just a determinism one. ### The remedies, in order of value **Pin the base by digest.** `FROM eclipse-temurin:21-jre-noble` is a mutable pointer; `FROM eclipse-temurin:21-jre-noble@sha256:...` is not. This is one line and removes the largest block of unpinned content in a typical image. It creates an obligation: something has to move that pin when a security update lands, whether that is a human on a schedule or an update bot. **Lock the application dependencies.** The build should install from a lockfile the repository owns, not resolve ranges at build time. For a JVM service that means resolving against a fixed dependency set rather than dynamic versions; the same principle applies to every ecosystem. **Pin OS packages, with eyes open.** `apt-get install -y curl="${CURL_VERSION}"` pins the version, and it is worth doing, but be honest about the failure mode: a distribution's main archive generally serves only current versions, so a pinned version disappears when the mirror moves on and the build breaks. Teams that need this to hold either point at a date-pinned snapshot of the distribution archive or mirror the exact package files into an internal artefact store. **Verify what you download.** If a build must fetch a file, fetch a versioned URL, never a `latest` alias, and check it against a recorded checksum in the same `RUN` step. `sha256sum -c` turns "whatever the host served" into "exactly this file, or the build fails". The strongest form is to stop fetching in the build at all: pull the artefact into your own store, then `COPY` it from the build context, which makes it visible to code review. **Move toward a hermetic build.** The end state, for the images where it is worth the effort, is a build whose only inputs are the build context and pinned, digest-addressed base layers - no unmediated network access at all. ### A worked example A Spring Boot fraud-scoring service builds on a JRE base, installs two OS packages, downloads an observability agent from a `latest` URL, and copies its fat jar in - roughly 2.3 GB of dependency layer. Rebuilding the commit that was released six weeks earlier, in order to reproduce a scoring discrepancy, produces an image with a newer JRE patch, newer OS packages and a different agent build. The discrepancy no longer reproduces, and nobody can tell whether that is because the bug was in the agent, in the JRE, or in the code. Pinning the base by digest and copying a checksum-verified, versioned agent tarball out of the internal store makes that rebuild an actual reconstruction rather than a fresh guess. ### What to say about how far to go Not every image needs a hermetic build. The pragmatic line most teams can hold is: base pinned by digest, application dependencies locked, downloaded artefacts versioned and checksum-verified, OS packages pinned where the archive supports it - and an explicit owner for moving the pins so that "reproducible" does not quietly become "unpatched".

  • You pin an OS package version and six weeks later the build fails because that version is gone. What now?
    That is the known cost of pinning against an archive that serves only current versions. The options are to point the build at a date-pinned snapshot of the distribution archive, to mirror the exact package files into an internal artefact store the build reads from, or to accept a floating version and move the reproducibility guarantee up to the base image, which you rebuild and pin by digest on a schedule you control.
  • Pinning the base by digest freezes its security patches. How do you avoid shipping a stale base forever?
    Make moving the pin a routine, owned change rather than an accident of rebuilding. An automated update proposal that bumps the digest, plus a scheduled rebuild of your own base images, keeps the pin fresh while keeping every individual build reproducible. The property you want is that the base only changes when someone decided it should, not that it never changes.
  • Is fetching over the network during a build always wrong?
    No - it is a tradeoff, and for a low-stakes internal tool the convenience usually wins. It becomes wrong when the image is released, rolled back, or audited, because then you need to reconstruct exactly what shipped. The middle ground most teams land on is: fetches allowed, but only from versioned URLs, only checksum-verified, and only through a mirror you control.

saying these in an interview costs you the question

  • Says apt-get install is fine because the tag is fixed
  • Treats a latest download URL as a stable input
  • Thinks rebuilding an old commit reconstructs the shipped image
  • Pins the base by tag and calls it pinned
  • Downloads an artefact without verifying any checksum
  • Ignores that every fetch is also a build availability risk

context