skip to content

How do you prove an emergency dependency patch actually reached production everywhere?

level: seniorimportance: nice to knowfreq 37%

answer

  1. intent is not state
  2. scan what runs, not what is committed
  3. digests identify content, tags name it
  4. enumerate every environment explicitly
  5. record when the last vulnerable instance stopped

basics

~20 s

A merged pull request and a closed ticket prove intent, not state. Proof is a chain: the fixed version inside a built artifact, that artifact identified by digest, that digest running in every environment, and a re-scan of the running inventory returning no matches.

solid answer

~50 s

Close the loop on evidence about what is *running*, not about what was *decided*. The chain has four links: the patched version appears in the resolved dependency set of the built artifact; the artifact is identified by its content digest rather than a mutable tag; every deployment in every environment references that digest; and a re-scan of the running inventory, not the source repositories, returns zero matches for the advisory. Where the interim control was a mitigation rather than an upgrade, add the effective running configuration as a fifth link — the flag as the process sees it, not as the config file declares it. The gaps that bite are predictable: replicas that never restarted, an autoscaling group still launching an old image, a forgotten environment, a service that was down during the rollout and came back on the old version, and anything you already published to customers. That last one is not a deployment problem at all; it is a disclosure obligation.

go deeper

for a junior

Know that merging the version bump is not the end: someone has to confirm the new build is actually running in production before the advisory can be considered handled.

for a middle

Explain the evidence chain — fixed version in the resolved build, artifact identified by digest, that digest deployed, running inventory re-scanned — and why a digest is stronger evidence than a tag.

for a senior

Name the places verification breaks: unrestarted replicas, autoscaling images, scaled-to-zero services, vendored copies and caches, and define an explicit exit condition before the rollout starts.

for a principal

Own the audit-truth angle: the exposure window and the environment list are what a post-incident review and any customer notification rest on, so capturing timestamps live is a deliberate organisational requirement.

## Intent versus state The most common way an emergency response fails is not that nobody patched. It is that everybody patched *something* and nobody established what is running. A merged pull request, a green pipeline and a closed ticket are records of intent. The question after an incident — asked by a post-incident review, by a customer, and eventually by an auditor — is a question about state: on which systems, between which times, was the vulnerable component running? Evidence has to be collected in the direction that answers that question, and it has to be collected while the incident is live, because reconstructing it a week later is guesswork. ## The chain of evidence **Link 1: the fixed version is in the artifact.** Not in the manifest, in the *resolved* set — the lockfile or the inventory generated from the actual build. A manifest declares an intention about versions; the resolved set records what was really assembled. If your build emits a component inventory, this link is free. **Link 2: the artifact is identified by content, not by name.** A digest is a cryptographic hash of the bytes: it identifies content. A tag is a label that can be repointed at different content tomorrow. "We deployed 2.3.1" is an unfalsifiable claim; "we deployed sha256:…" is a checkable one. If your deployments reference tags, you cannot close this chain, and fixing that is a durable improvement worth taking out of the incident. **Link 3: that digest is what is running, everywhere.** Everywhere is the load-bearing word. Enumerate environments rather than assuming the list: production, every region, staging and pre-production, disaster-recovery standbys, the long-lived customer demo nobody owns, the batch cluster that only runs at 02:00, and any edge or on-premise footprint. **Link 4: re-scan the running inventory.** Scan what is deployed, not what is committed. A repository scan tells you the source is fixed; it says nothing about the instance still serving traffic on last month's build. **Link 5 (if you mitigated): the effective configuration.** A flag is only applied if the running process sees it. Read it back from the process or the platform, not from the file in the repository. This is the same distinction as links 1 and 2, applied to configuration. ## Where the chain breaks in practice - **Long-lived processes.** A rolling update that stalled at 90%, or replicas that were never restarted because the deployment object changed but nothing triggered a roll. - **Autoscaling and node images.** The deployment is patched; the scaling group still launches from an old image, so the vulnerable version reappears at the next scale-out. This one is nasty because it verifies clean at the moment you check. - **Anything that was down.** A service scaled to zero, a paused job, a standby that comes back on the pre-incident version after the rollout finished. - **Copies that are not dependencies.** A vendored or bundled copy, a file checked into the repository, a component baked into a base layer several builds ago. Version-based checking finds none of these. - **Caches.** A build cache or a mirrored artifact that serves the old resolution to the next build. - **Published artifacts.** Anything you have already shipped to consumers is outside your deployment entirely. You cannot patch it for them; you can only tell them, which makes it a disclosure task rather than a rollout task. ## What "done" should mean Define it before you start, so the incident has an exit condition rather than a feeling. A reasonable bar: no running workload in any enumerated environment resolves to an affected version; every deployment references a digest containing the fix; interim mitigations are either removed or explicitly retained with an owner; and the evidence is captured somewhere durable with timestamps. The timestamps matter more than they look. The two questions that decide whether you owe anyone a notification are *were we exposed* and *for how long*, and the answer to the second is the interval between the vulnerable version first running and the last instance of it stopping. If you did not record when the last instance stopped, you cannot answer it, and "we believe it was fixed promptly" is not an answer that survives scrutiny. ## Why this is the link that gets skipped By the time the fix is merged the adrenaline is gone, the channel has quieted, and verification is tedious work with no visible reward. That is exactly why interviewers ask about it: it separates people who have run a response to its actual end from people who have watched one. The strongest answers name specific evidence — a digest, a re-scan of running workloads, an enumerated environment list — rather than saying they would "confirm with the teams".

  • You verified clean, then the vulnerable version reappears the next morning. What is the usual cause?
    Something outside the deployment path regenerated it: an autoscaling group or node pool launching from an old image, a cached build resolution, or a floating version range that resolved differently on the next build. Verification checked the running set at one instant, which is why the artifact reference should be an immutable digest and why the fixed version should be pinned until the rollout is complete.
  • Why does re-scanning your repositories not close an emergency patch?
    It shows the source is fixed, which is intent. It cannot see a workload still serving traffic on an older build, a replica that never restarted, or a vendored copy that no manifest declares. The scan has to run against the deployed inventory, and the artifacts have to be identified by digest, or you are asserting rather than proving.
  • What do you record while the incident is live so the post-incident review can be answered honestly?
    The time the advisory landed, the time each mitigation went live and on which workloads, the digest that contains the fix, the time the last vulnerable instance stopped running, and the enumerated environment list you checked against. The exposure window is the interval between first vulnerable run and last, and it is not reconstructible a week later from memory.

A merged pull request is a receipt saying you asked the bank to transfer money. The digest running in production is the statement showing it arrived. Incidents are closed on statements, not receipts.

saying these in an interview costs you the question

  • Treats a merged pull request and a closed ticket as proof
  • Re-scans source repositories instead of the running inventory
  • Verifies by mutable tag rather than by content digest
  • Forgets standbys, batch clusters and scaled-to-zero services
  • Assumes patching deployments also protects already-published artifacts

context