How do you determine which systems actually ran a malicious dependency version?
answer
- current lockfile is intent, not history
- ask what is deployed, not what is committed
- artifact manifests over repository files
- pull records from the caching proxy
- a logging gap is unknown, not clean
basics
~20 sNot from the current lockfile, which describes today's intent. Reconstruct it from lockfile history across the exposure window, from the deployed inventory of running artifacts and their manifests, and from registry or caching-proxy pull records.
solid answer
~50 sThree sources, and you need all three. First, lockfile history rather than the lockfile on the default branch today — a dependency range may have briefly resolved to the bad version and been re-resolved away since, so you diff across the exposure window and include every branch and fork that built. Second, the deployed inventory, which is usually the surprising one: what runs in production is artifacts built at many different times, and services still on older images are described by no current lockfile at all, so you enumerate what is deployed and read each artifact's own manifest or SBOM. Third, pull records from your registry or internal caching proxy: if the bad version was live for forty minutes, those logs are the only evidence of who actually fetched it, and they routinely contradict "nobody installed it". A gap in those logs is unknown, not negative.
go deeper
Know that the lockfile in the repository describes what the next build will install, not what is running in production today, and that those two answers routinely differ.
Be ready to name the three evidence sources and what each one misses: lockfile history, deployed artifact manifests, and registry or proxy pull records. Explaining why you need all three is the answer.
Demonstrate scoping under uncertainty: bound the exposure window correctly, mark logging gaps as unknown rather than clean, and drive the list by environment rather than by repository.
Own the capability gap this exposes. If nobody can answer 'what is deployed and what is inside it' within an hour, that is an inventory investment to fund, and the incident is the moment to make that case with evidence.
## The question behind the question "Where is this actually running?" is the hardest hour of a malicious-dependency response, and interviewers ask it because the obvious answer — grep the lockfiles — is confidently wrong in three different directions. ## Source 1: lockfile history, not the current lockfile A lockfile is a snapshot of a resolution decision. Reading the one on the default branch today tells you what the next build intends to install, which is exactly the thing you have possibly already changed. What you need is the set of resolutions that were *live during the exposure window* — from the moment the malicious version was published to the moment it was yanked or your controls blocked it. That means diffing lockfile history across the window, and it means every branch, release branch and long-lived fork that ran a build, not just the trunk. It also means remembering the asymmetry: a repository whose lockfile never named the bad version can still have installed it if a build ran with resolution unlocked, in a CI mode that ignored the lockfile, or in a container image whose base layer resolved it independently. ## Source 2: the deployed inventory Source control describes intent. Production runs artifacts. Those two drift apart constantly and the gap is where the response goes wrong. A realistic picture: your platform has forty services, the current lockfiles are clean, and nine services are still running images built months ago from lockfile states nobody has looked at since. No current repository file describes those nine. The only authoritative statement about what is inside a running artifact is the artifact itself — its recorded manifest, its embedded SBOM if it has one, or a fresh inventory taken from the image. So you enumerate what is *deployed* — image digests running in each environment, including the ones nobody redeploys — and match each against the affected package and version range, using its build timestamp to bound the question: an artifact built before the malicious version existed cannot contain it, and that is the cheapest way to shrink the list. ## Source 3: pull records Most organisations put an internal caching proxy or a pull-through mirror between their builds and public registries. That component is the closest thing you have to a flight recorder. Its logs answer the question no repository can: **who fetched this version, and when.** This is where "nobody installed it" goes to die. Someone's exploratory branch, a laptop running a build outside CI, a nightly job — the proxy saw it even if no commit records it. The reverse reasoning does not hold, and saying so is what separates a good answer from a confident one. Absence of a pull record proves nothing unless you can show that (a) every consumer is forced through the proxy, (b) the proxy logs every fetch including cache hits, and (c) no copy existed already — in a warm cache, a vendored directory, a pre-baked builder image, or an offline mirror. A cache hit that is not logged looks identical to a package that was never fetched. Treat every gap as unknown and resolve it another way. ## Putting it together The practical output is a single list with a status per environment, not per repository: | Evidence | Answers | Blind spot | | --- | --- | --- | | Lockfile history over the window | what builds intended to resolve | unlocked builds, forks, base images | | Deployed artifact manifests | what is running right now | artifacts with no recorded inventory | | Proxy and registry pull records | who actually fetched it | unlogged cache hits, direct fetches | Each source covers another's blind spot, which is why the answer is all three and not the best one. ## Bounding the window The exposure window starts at publication and ends when the version stopped being resolvable *for you* — which may be later than the yank, because your own cache may have kept serving it. Anything that installed inside that window is in scope; anything demonstrably outside it is not. Getting the window's end wrong in the optimistic direction is the most common scoping error, and the caching proxy is usually the reason. ## What a strong answer sounds like "I'd scope by environment, not by repository. Lockfile history over the window tells me intent, the deployed inventory tells me what is actually running — including the services nobody has rebuilt in months — and proxy pull logs tell me who really fetched it. Where the proxy has no record I mark it unknown rather than clean, because a cache hit may never have been logged."
- Nine services still run images built months ago. Are they in scope?In scope until proven otherwise. Those are exactly the artifacts no current lockfile describes. Bound them by build timestamp first — anything built before the malicious version was published is out — and for the rest read the image's own recorded inventory rather than the repository it came from.
- The caching proxy has no record of that version. Can we call it clean?Only if every consumer must go through the proxy, cache hits are logged as well as misses, and no pre-existing copy sat in a warm cache, a vendored directory or a builder image. Otherwise a silent cache hit and a never-fetched package look the same. Mark it unknown.
- How do you decide when the exposure window ends?Not at the upstream yank, but when the version stopped being resolvable for you. If your own mirror or proxy kept serving a cached copy after the yank, your window is longer than the public one, and that difference is usually where a missed environment hides.
saying these in an interview costs you the question
- Greps only the current default-branch lockfiles
- Assumes a yank means nothing was ever pulled
- Treats repository inventory as deployed inventory
- Reads missing proxy logs as proof of absence
- Ignores forks, release branches and base images