Why does rebuilding a container image from the same Git commit produce a different digest?
answer
- Same source is not same bytes
- Git does not record file mtimes
- The config carries a created timestamp
- Unpinned installs resolve at build time
- Cold cache exposes what local builds hide
basics
~20 sAn image records file bytes and metadata, not your commit. A fresh checkout gives every file a new mtime, the image config gets a new created timestamp, unpinned installs pull newer packages, and changed build args alter the config.
solid answer
~50 sA digest hashes the image's layers and its config JSON, so any changed byte in either changes it. Four causes dominate. **Timestamps**: the config carries a `created` field taken from the clock, and `COPY` preserves each file's mtime - Git stores no mtimes, so a fresh checkout stamps everything with the checkout time and the layer tar differs even though content is identical. **Injected metadata**: `LABEL`s such as a build date and `--build-arg` values baked into `ENV` sit in the config. **Unpinned inputs**: a moving `FROM` tag, `apt-get install`, or a `curl` of a latest tarball resolve to whatever upstream serves today. **Non-deterministic steps**: generated files with embedded dates, archives with variable entry order. Locally you often miss all of this because the build cache reuses the previous layers; on a cold CI runner the digest moves.
code
bash · 3 linesdocker build --no-cache -q -t fraud-scoring:a .
docker build --no-cache -q -t fraud-scoring:b .
docker image inspect --format '{{.Id}} {{.Created}}' fraud-scoring:a fraud-scoring:bgo deeper
Be ready to say that an image is bytes plus metadata, not a pointer at your commit, and to name one concrete cause such as a build timestamp or an unpinned package install.
Explain the mechanics: what the config's created field is, why COPY carries file mtimes out of the build context, and why a fresh Git checkout changes them without changing content.
Show how you would prove which layer moved and decide which causes are worth fixing on a real service, including why a warm local cache hides the problem that a cold CI runner exposes.
Own the position: which images in the estate need reproducible rebuilds at all, what pinning discipline you mandate, and how you keep a moving base-image pin from becoming an unpatched one.
### What a digest actually covers A container image's digest is a sha256 over its manifest, and the manifest names the config blob and every layer blob by their own sha256. Nothing in that structure mentions your source control system. Two builds are "the same image" only if the bytes match: every entry of every layer tar, with its path, mode, owner and modification time, plus every field of the config JSON. Your Git commit is an *input* to the build, not a *component* of the image, so "same commit" and "same bytes" are different claims. Four families of variance sit between them. ### Family 1 - timestamps The image config carries a `created` field that the builder fills from the wall clock unless it is told otherwise, and each entry in the config's `history` array carries one as well. That alone guarantees a new config blob, and therefore a new digest, on every build. Underneath, layers are tar streams, and tar records an mtime per entry. `COPY` preserves the modification time of the file it read from the build context. Git does not store mtimes, so a fresh clone or a CI checkout stamps every working-tree file with the moment it was written to disk. The file *contents* are byte-identical; the tar *entries* are not; the layer digest differs, so the manifest digest differs. ### Family 2 - build metadata you inject Values passed with `--build-arg` that the Dockerfile bakes into `ENV`, and `LABEL`s such as a build date, a build number or a VCS ref, all live in the config JSON. A pipeline that stamps `org.opencontainers.image.created` with the output of `date -u +%FT%TZ` has deliberately made its images unreproducible - the filesystem may be identical and the digest still changes. ### Family 3 - unpinned inputs `FROM eclipse-temurin:21-jre-noble` resolves a mutable tag at build time; the publisher moves that tag whenever it rebuilds for a CVE. `apt-get install` installs whatever the mirror is serving today. `curl` of a "latest" tarball, a dependency resolution without a lockfile - each is a network fetch whose result is a function of the calendar, not of your repository. This family is the one that changes real *content* rather than metadata, and it is the one that actually hurts: it means the image you shipped is not the image you tested. ### Family 4 - non-deterministic build steps Code generators that embed a timestamp or a hostname, archives whose entry order follows a directory walk, tools that consult `$RANDOM` or mint a UUID, compilers that record absolute build paths. The determinism of an application toolchain is that toolchain's own problem, but it surfaces here as a layer whose bytes differ every time. ### Why you rarely notice it locally BuildKit's cache key for `COPY` is a content checksum, not an mtime, so a second build on the same machine usually hits the build cache and reuses the layer it produced before - identical digest, and a comforting but false sense of determinism. On an ephemeral CI runner with a cold cache, or after `--no-cache`, the same source produces different bytes. ### A worked example A Spring Boot fraud-scoring service builds `FROM eclipse-temurin:21-jre-noble`, installs two OS packages, then copies a fat jar into a 2.3 GB dependency layer. Two CI runs of commit `a91c4f7`, 40 minutes apart, produce image IDs that differ. Peeling it apart: the config `created` timestamps differ by 2412 seconds, which alone changes the config digest. Beyond that, the layer that copies the application also differs, because the second runner checked the repository out fresh and every file carries a later mtime. The base image layers are shared, so the tag had not moved in that window - but a run the next morning picks up a rebuilt base and changes several more layers. ### Making it stable Work the families in order of value. Pin the base image by digest so `FROM` stops being a moving target. Pin application dependencies with a lockfile and OS packages with explicit versions, or fetch pre-resolved artefacts from an internal store instead of the public internet. Verify anything downloaded against a recorded checksum. Drop build-time labels, or accept that they cost you bit-identity. Set `SOURCE_DATE_EPOCH` to the commit timestamp so BuildKit writes that value into the config and into file timestamps instead of reading the clock. Then prove it: build twice on a cold cache and compare image IDs, rather than assuming. ### What "reproducible" is worth Perfect bit-identity is a spectrum, not a switch, and the last few percent - a toolchain that embeds paths, a package archive that no longer serves the version you pinned - can cost more than it returns. The cheap 80% (digest-pinned base, lockfiles, normalised timestamps) is worth doing on every image, because it is what makes "the digest changed, why?" an answerable question instead of background noise.
- Two builds of the same commit differ. How do you narrow it down to one instruction?Run `docker history` on both images and compare step sizes and commands, then compare the two layer lists - the first position where they diverge names the guilty instruction. Export that image with `docker save` and list the layer tar with `tar -tvf` including timestamps: if paths and sizes match but mtimes differ, it is a timestamp problem; if file sizes or file sets differ, an input changed.
- Why doesn't using the same FROM tag guarantee the same base layers?A tag is a mutable pointer in the registry. Publishers move `21-jre-noble` onto a freshly rebuilt image whenever they patch the OS packages inside it, so the same tag resolves to different layers on different days. Pinning `FROM name@sha256:...` is what makes the base an immutable input; you then need a process to move that pin deliberately.
- Does stamping the Git commit into a LABEL make the build reproducible?No. It records provenance, which is useful, but it does not lock any input - the same label can sit on two images built from different upstream packages. Worse for this property, a label that also carries a build date guarantees the config blob changes on every build, so it actively prevents bit-identical rebuilds.
Photocopying the same page twice gives you the same text but two different sheets of paper: the digest is hashing the paper, not the manuscript.
saying these in an interview costs you the question
- Thinks the same Dockerfile always yields the same digest
- Believes the registry assigns a new digest on each push
- Assumes an unchanged tag means unchanged base layers
- Thinks identical file contents guarantee identical layer bytes
- Concludes builds are deterministic from a warm local cache
- Adds a build-date label and still expects bit-identity