An autoscaled service deployed from a reviewed digest runs different bytes on nodes added hours later - how do you diagnose it?
answer
- ask where the reference resolves last
- read the persisted spec, not the pipeline
- group running digests by start time
- node cache explains the partial spread
- tag history dates the remap
basics
~20 sFind the last place the image reference is resolved. If the persisted workload spec still holds a tag, every later pull re-resolves it, so scale-out and node replacement drift the fleet onto whatever that tag names now.
solid answer
~50 sFirst capture evidence before anything is rolled: record the image digest each running instance reports, grouped by start time. A clean split by age - old instances on the reviewed digest, newer ones on another - points straight at pull-time re-resolution rather than a bad deployment. Then read the persisted workload spec: if it contains a tag, the digest you reviewed was resolved somewhere upstream and thrown away, and every instance created afterwards asks the registry what the tag means now. Confirm with the registry's tag history to see when the mapping moved and which identity pushed. Local image caches explain why only part of the fleet changed: nodes that already held the bytes never re-pulled. The fix is that the digest must live in the spec that scale-out reads, backed by a deploy-time gate that refuses any image reference without one.
go deeper
Understand that new instances are created from a stored specification, and if that specification names a tag then each new pull asks the registry what the tag means at that moment.
Be able to trace a reference through every step that could resolve it and say which step persists its result, and explain why caching makes the drift partial rather than fleet-wide.
Demonstrate the investigation: preserve evidence before rolling, group running digests by start time, correlate with registry tag history, then scope what the unreviewed workload could reach and fix the persisted spec.
Own the standard that a pinning claim is only meaningful at the last resolution point, and require continuous comparison of running digests against approved ones rather than trusting pipeline-time evidence.
## The principle behind the symptom A pin is only as strong as the **last place the reference is resolved**. A reference passes through many hands - a developer's file, a build job, a deploy tool, a cluster controller, the node agent that actually pulls, and the local cache on that node. If any step downstream of your pin still holds a mutable name, the pin was decorative: the name gets resolved again, later, under conditions nobody reviewed. That is what a multi-tenant autoscaling service exposes so cruelly. The reviewed digest was true at deploy time. Then a scale-out event, a node replacement, a spot reclaim or a rolling eviction creates new instances from the *stored* spec, hours later, and each of those pulls resolves the name against the registry as it is now. ## Diagnose in this order **1. Freeze and collect, do not roll.** The instinct to redeploy destroys the only evidence you have. For every running instance record the resolved image digest the runtime reports and the instance start time. Mixed digests across one workload is the finding. **2. Read the persisted spec, not the pipeline.** Ask what the platform has actually stored for this workload. If it says `svc:release-2026.04`, you have your answer: the digest existed only in the pipeline's logs or in a pre-flight check, and was discarded before anything durable was written. **3. Correlate against the tag history.** Registries record when a tag was remapped and by which identity. Match the timestamp against the moment the first divergent instance started. That distinguishes a deliberate push from a stale mirror or a restored backup. **4. Explain the partial spread.** People misread "only half the fleet" as randomness or a registry fault. It is neither. Instances started before the remap keep running the bytes they already have; nodes holding a cached copy may not re-pull at all depending on pull policy; only fresh pulls after the remap fetch new content. Age of the instance, not configuration, decides who is affected. If the split does **not** follow start time, look elsewhere - a mirror serving stale content, or two clusters pointed at different registries. **5. Establish what actually ran.** For a multi-tenant API the security question is not only "is it malicious" but "what did unreviewed code have access to": the service account, the tenant data reachable through it, and any credential mounted into it. Assume the unreviewed bytes had everything the workload has, and scope the response from there. ## The fix, in order of durability - **Put the digest in the persisted spec.** Whatever writes the deployment - the pipeline, a controller reconciling from git, a template - must emit `svc@sha256:...`. If a tool resolves a tag to a digest for verification and then applies the tag anyway, the resolution is thrown away at exactly the moment it matters. - **Gate the deployment.** A check that refuses any workload whose image reference is not digest-pinned makes the failure mode loud instead of silent, and turns "we agreed to pin" into something enforced rather than remembered. - **Close the registry side too.** Forbidding remaps of existing tags means the substitution is rejected at the source; restricting who holds push rights shrinks the set of identities that could have done it. Neither replaces the digest in the spec, because a mirror, another registry or a restore sits outside those rules. - **Detect drift continuously.** Compare the digests actually running against the digests approved for that release, on a schedule. This is the control that would have found the problem in minutes instead of at the next incident. ## The trap to name explicitly "We pin in CI" is the single most common false comfort here. Resolving a digest during the build or the deploy proves what *that step* saw; it says nothing about what the platform will resolve when it creates instance number forty at three in the morning. The question to ask of any pinning story is always: **what string is stored in the thing that scale-out reads?** If the answer is a tag, there is no pin.
- Why does only part of the fleet change rather than all of it?Instances started before the remap keep running the bytes already on their nodes, and a node holding a cached copy may skip the pull entirely. Only fresh pulls after the remap fetch the new content, so the split follows instance age. A split that does not follow age suggests a stale mirror instead.
- Your deploy tool resolves the tag to a digest before applying - is that enough?Only if the resolved digest is what gets persisted. If the tool resolves it for logging or a pre-flight check and then applies the tag, the resolution is discarded and everything created later re-resolves. Always inspect what the platform stored, not what the pipeline printed.
- Which single control would have prevented this drift?Refusing to admit any workload whose image reference lacks a digest, so a re-resolvable name cannot reach the persisted spec at all. Registry-side immutability is a valuable second layer, but it is scoped to one registry and does not travel with the artifact.
- What do you do first once you confirm unreviewed bytes ran?Preserve the evidence - the divergent digest, the affected instances and the registry push record - then treat every credential and dataset reachable from that workload as exposed. Replace the instances with the approved digest, and only then work backwards to which identity pushed and how it held that right.
saying these in an interview costs you the question
- Rolls the fleet before capturing running digests
- Blames image caching alone and clears caches
- Assumes a digest pinned in CI reaches the running spec
- Treats mixed digests as a registry bug
- Says a redeploy of the same tag proves what runs