Two data centres run the same service from the same image tag yet behave differently — how is that possible?
answer
- same string, two moments
- a pointer that moved in between
- resolution happens once, at start
- a host keeps what it already holds
- compare resolved digests, not labels
basics
~20 sA tag is a movable label, so the two sites resolved it at different moments and the publisher repointed it in between — identical configuration, two different images, and nothing in the deployment record disagrees.
solid answer
~40 sBoth sites hold the same string in configuration, but a tag is a pointer the publisher can move, and each site turned that pointer into actual content at the moment it started its workload. If the tag was repointed between those two moments, the two sites run different bytes under one name. A host that already holds what it resolved earlier can keep running it long after the tag has moved, so the divergence outlives a restart. Every artefact of the deployment agrees — configuration, change ticket, dashboards — because all of them record the label and none of them records the content. Read the resolved content digest back from each site and compare; that is the only comparison that settles it.
code
pseudocode · 15 lines# each site turns the reference into content once, when it starts the workload
function startWorkload(reference, site):
digest = registryLookup(reference.repository, reference.tag) # a tag: whatever it points at NOW
if not site.hasLocally(digest):
site.fetch(digest)
site.record(requestedReference = reference, resolvedDigest = digest, at = now())
site.run(digest)
# timeline
09:00 startWorkload("ledger:1.4", siteA) -> digest = hash-of-content:a1b2c3 # build 812
10:20 publisher repoints "ledger:1.4" -> the label now names build 907
11:45 startWorkload("ledger:1.4", siteB) -> digest = hash-of-content:d4e5f6 # build 907
# both sites' configuration still reads exactly "ledger:1.4"
# only site.record() holds anything that distinguishes themgo deeper
Know that a tag is a label someone can move, so two machines that read the same label at different times need not be running the same thing.
Explain when resolution happens — once, as the workload starts — and why a host that already holds content it resolved earlier need not go and look again.
Show how you would confirm it: read the content identity back from each site rather than comparing configuration, which is identical by construction here.
Address the standard that prevents a recurrence — what your delivery path records at deploy time, and whether an unpinned reference should reach production at all.
## The failure, stated precisely Two data centres run the same payments ledger. Both configurations name the same image by the same tag. The two sites nonetheless behave differently — one rejects a class of request the other accepts, or one shows a latency change the other does not. Nothing in configuration management is out of sync, and that is exactly what makes the incident long. A tag is a **movable pointer**: a publisher chooses what content it points at and may repoint it whenever they like. **Resolution** — turning that pointer into actual content — happens per host, at the moment the workload starts. Two resolutions separated in time are two independent questions asked of a value that may have changed in between. ## The timeline that produces it 1. Site A starts its workload at 09:00 and resolves the release tag to the content published as build 812. 2. At 10:20 the publisher republishes under that same tag — a hotfix, a rebuilt base, or a pipeline that republishes on every merge to a release branch. 3. Site B starts at 11:45, resolves the same string, and gets build 907. 4. Both sites' configuration still reads the same reference, and will keep reading it through every subsequent review. Nothing in that sequence is a fault. Every component did what it is specified to do. The label meant one thing at 09:00 and something else at 11:45, and only the label was ever written down. ## Why it outlives a restart A host that already holds the content it resolved has no reason to go and ask again. Platforms differ here, and the difference matters: some re-resolve a reference on every start, some only when the content is absent locally, and some only under an explicit instruction. Whichever design you are on, the practical result is the same — the divergence can persist for weeks and survive restarts and reboots. The nastiest variant is a site that becomes internally inconsistent. The original hosts keep running what they resolved long ago; a host added later during a scale-up resolves afresh and joins the same site carrying different content. Requests then succeed or fail depending on which replica answers, which reads as an intermittent bug rather than a deployment problem. ## What does and does not reveal it | where you look | what it says | reveals the divergence | |---|---|---| | deployment configuration | the reference, which is a label | no — identical on both sites by construction | | change ticket | the release name people agreed on | no | | dashboards and log labels | the version string the label implies | no | | content digest read from a running host | the exact bytes in use | yes | | deploy record holding the resolved digest | what each site resolved, and when | yes, and after the replicas are gone | The first three rows are the trap. They are accurate, they are consistent with each other, and they are describing two different pieces of software. ## How to confirm and close it - read the content digest back from a replica at each site and compare the two — that comparison *is* the diagnosis - capture it before restarting or replacing anything, because a replacement may resolve again and destroy the evidence - if the two identities match, the divergence is not in the image and you should be looking at configuration, data or dependencies instead - converge both sites onto one explicitly chosen content identity rather than redeploying from the tag and hoping the timing works out - record the resolved identity at deploy time so the next occurrence is visible in minutes rather than days - treat an unpinned reference in a production deployment as the defect, and the behavioural difference as merely its symptom ## Where a moving tag is still the right choice None of this makes tags wrong. A moving tag is the mechanism by which a publisher ships a fix to every consumer without any of them editing configuration, and in a development environment, a scratch cluster or an internal tool that is precisely what you want: everyone picks up the newest build for free. The property that broke the ledger — a reference whose meaning depends on when it was read — is the same property, seen from the other side. The judgment is about where you can afford it, not about whether the mechanism is sound.
- Why do the configuration files, the change ticket and the dashboards all fail to reveal the divergence?Because each of them records the reference, and the reference is the label. Nothing in that chain stores what the label resolved to, so every artefact is accurate and mutually consistent while describing two different builds. The divergence becomes visible only where a content identity is recorded at deploy time, or read back from a running host.
- Does redeploying both sites from the same tag fix it?It converges them onto whatever the tag points at at that moment, which is not the same as fixing the class of problem: the next repoint reopens it, and the redeploy destroys the evidence of what each site was running. Capture and compare both content identities first, then deploy both sites from one explicit identity.
saying these in an interview costs you the question
- Says the two sites must be running different configuration somewhere.
- Blames the registry for serving inconsistent content to different sites.
- Assumes restarting a workload always re-checks what the tag points at.
- Thinks a version-shaped release tag cannot have been repointed.
- Tries to confirm it by comparing references, which are identical by definition.