Why tag a container image with the short git commit SHA rather than an incrementing build number?
answer
- Three properties a build tag must buy
- One of them a SHA does not give
- A counter needs someone to allocate it
- Re-running the same commit is the trap
- Most teams end up composing both
basics
~20 sA commit SHA ties the image to the source that produced it, needs no shared counter, and stays unique across branches and parallel builds. Its weakness is that SHAs carry no order, so many teams combine a build number with the SHA.
solid answer
~50 sA commit-derived tag such as `indexer:7e1c94d` answers the question you actually ask during an incident: which source produced the bytes running right now. It is generated by the checkout itself, so it needs no counter shared between pipelines, it never collides across branches or parallel builds, and anyone can go straight from the tag to the diff. A build number such as `indexer:4183` answers a different question — which run produced it — and has one property the SHA lacks: it sorts, so you can see at a glance that production is behind staging. Neither is safe alone if a commit can be rebuilt: the same SHA tag pushed twice from different base images means one name has meant two different sets of bytes. That is why a combined tag like `1.4.2-4183-7e1c94d` is common — ordered, traceable and unique per run.
go deeper
Know that the deploy tag is generated by the pipeline from the commit or the run, never typed by a person, and that its whole point is to name one specific build permanently.
Be able to compare the two directly: traceability and no coordination on one side, sortability and a link to the run's evidence on the other, and explain why a combined tag is so common.
Bring up the rebuild hazard unprompted — the same commit built twice under the same tag — and say how you would stop a pipeline from silently overwriting a build tag.
Frame it as an identity scheme other systems will depend on. Decide what must be derivable from a tag alone, and accept that once deploy records and dashboards parse it, the format is expensive to change.
### What the immutable tag has to buy you Every deploy tag scheme is trying to buy three things: **traceability** (from a running container back to source), **uniqueness** (no two builds ever share a name), and **order** (which of these two is newer). A short commit SHA and an incrementing build number each buy some of that, and the interesting part of this question is which they miss. ### Where the commit SHA wins *Traceability is direct.* `indexer:7e1c94d` is not a lookup key into a build system — it is the commit. During an incident you paste it into `git show` and read the diff. If your build system is down, unavailable to you, or has rotated its history, the tag still means something. A build number is a foreign key into a database you must be able to reach. *It needs no coordination.* The SHA exists the moment the code is checked out. Nothing has to allocate it, so two pipelines, a re-run of an old branch and a developer building locally all produce non-colliding tags with no shared counter. Build numbers are allocated per pipeline definition, so the same integer can be produced by two different pipelines, and moving or re-creating a pipeline often resets the sequence. *It is stable across retries.* Re-running a failed pipeline on the same commit gives you the same tag — usually what you want, since the intent is "produce the image for this commit". ### Where the commit SHA loses *There is no order.* This is the real cost, and it shows up constantly in operations. Given `indexer:7e1c94d` in production and `indexer:a4f0b12` in staging, nothing in those two strings tells you which is ahead. A registry tag listing sorted alphabetically is noise. Answering "is production behind?" requires going back to git, or to a deployment record, every single time. *Re-running is a double-edged sword.* Because a re-run produces the same tag, the second push **re-points that tag to different bytes** if anything outside the commit has changed — a floating base image tag, a package index that moved, a changed build argument. The name is now ambiguous, and it is ambiguous in the worst possible way, because everyone assumes a SHA tag is immutable. Two defences exist: make the pipeline refuse to overwrite an existing build tag (registries can be configured to reject a tag that already exists, though that is a registry-side control), or make the tag unique per run by including the run identifier. *Short SHAs can collide.* A seven-character abbreviation is fine for a small repository and gets less fine as history grows; git itself lengthens abbreviations over time for this reason. It is a small risk, but it is a reason not to shorten aggressively. ### Where the build number wins It sorts. `4183` is obviously later than `4102`. It is also short, easy to say out loud, and maps one-to-one onto a pipeline run with its logs, its test results and its approvals — useful when the question is "where is the evidence for this build" rather than "what source is this". ### Why teams end up combining them The common resolution is not to choose. A tag such as `indexer:1.4.2-4183-7e1c94d` carries a version for humans, a build number that gives order and unique-per-run identity, and a commit SHA that gives traceability. It is long, and nobody types it by hand — but nobody should be typing deploy tags by hand anyway; the pipeline emits it and the deployment record stores it. Whatever the composition, the properties that actually matter are the same: - **Written exactly once.** No later build may ever reuse that string. This is the property the whole rollback story rests on. - **Derivable from the build, not invented at deploy time.** A tag that a human types during a release is a tag that will eventually be typed wrong. - **Reversible to a commit.** You must be able to get from a container back to source without guessing. ### A concrete check A good test of any scheme: page yourself at 03:00 with only the output of `docker ps` on one host. Can you tell, from the image reference alone, which commit it is and whether it is the newest build? A bare SHA tag answers the first and not the second; a bare build number answers the second and not the first. The scheme you pick should answer both without asking you to look anything up first.
- What breaks if the same commit is built twice and pushed under the same SHA tag?The second push re-points that tag, so one name has now meant two different sets of bytes — and since everybody treats a SHA tag as immutable, nobody looks for that. It happens whenever something outside the commit changed: a floating base image, a package index, a build argument. Defences are making the pipeline refuse to overwrite an existing build tag, or making the tag unique per run by including the run identifier.
- Why not just tag with the release version, like `indexer:1.4.2`?A release version is a promise about behaviour, not an identity for a build. Not every build is a release, so most builds would have no name at all, and a patched rebuild of the same version is a strong temptation to re-point the tag. Versions are worth carrying — usually as part of a longer tag or as a separate alias — but the tag that deployments and rollbacks reference should be one the pipeline generates for every single build.
saying these in an interview costs you the question
- Says a SHA tag tells you which build is newer
- Assumes build numbers are globally unique across pipelines
- Thinks re-running a build cannot change the image
- Relies on a human to type the deploy tag
- Tags only releases, leaving most builds unnamed
- Believes any short SHA is collision-proof forever