skip to content

Malicious Dependency Response

It is already installed somewhere. The first hour is blast radius from lockfiles, assuming install-time code ran and rotating what it could reach, then freeze, pin or roll forward.

on this pageshow

questions

4

Why is deleting a malicious dependency and reinstalling not a complete response?

level: juniorimportance: must knowfreq 72%

answer

  1. the install already happened
  2. removal only fixes future resolutions
  3. resolving a package runs its code
  4. treat it as credential exposure
  5. rotate what those environments could read

basics

~20 s

Because the package already ran code on every machine that installed it. Removing it only changes what the next install resolves. Treat those environments as compromised: rotate every credential they could read, and rebuild what they built.

solid answer

~50 s

Removing the package changes what the next install resolves; it does nothing to the machines where the bad version already landed. Installing a dependency runs publisher-controlled code on the installing host, so every laptop, CI runner and image build that resolved that version has to be treated as having executed attacker code with that environment's privileges. The response therefore opens with scoping and rotation, not a lockfile edit: list the environments that installed it during the exposure window, treat every credential readable from them as stolen — deploy roles, registry and publish tokens, the cloud role a runner assumes, developer tokens on laptops — and rotate them. Then rebuild any artifact produced on those machines in that window, because fixing the source tree does not fix images that were already built and pushed. The clean lockfile is the last step, not the response.

go deeper

for a junior

Be ready to say plainly that installing a package can execute its code, so removal fixes the next install and not the machines already affected. Naming credential rotation as the first move is what the screener is listening for.

for a middle

Expect to explain what a lockfile actually governs — future resolution — and to list concretely what is not covered by it: injected environment secrets, credential files, already-built images, and actions taken with a stolen token.

for a senior

Show ordering under time pressure. Cut re-entry paths first, then rotate by reach, then rebuild artifacts produced on affected hosts, and only then land the dependency fix. Be able to say why the dependency fix is last.

for a principal

Own the framing that this is a code-execution incident on your build estate, not a dependency bug, and make sure the organisation's playbook reflects that ordering before it is needed — including who is allowed to declare 'assume compromised' without waiting for proof.

## What this question is really testing When a malicious version of a package reaches your dependency graph, the instinct is to treat it as a bad dependency: remove it, resolve something clean, commit the lockfile, move on. That instinct is wrong in one specific way, and this question exists to find out whether you know which way. ## Installing is executing Installing a dependency is not the same as copying a file into a folder. Across mainstream ecosystems — npm, PyPI, RubyGems, and build-plugin ecosystems on the JVM — resolving and installing a package can execute publisher-supplied code on the host doing the install, and even where it does not, the unpacked artifact is immediately compiled, loaded or executed by the next build step. The working assumption for a response is therefore blunt: **if a machine resolved and installed the malicious version, attacker-controlled code ran there with the privileges of whoever ran the install.** That is not a dependency defect. That is code execution on your build hosts and your engineers' laptops. This is also why the usual triage vocabulary does not apply. Reachability — whether your application actually calls the vulnerable function — is the right question for a vulnerable dependency, because a flaw that is never called cannot be exercised. It is the wrong question for a malicious dependency, because the code did not wait to be called. "We never import that package" is not a defence. ## What removal actually changes A lockfile is a statement about what a future resolution should produce. Reverting it makes the *next* install clean. It does not touch: - **Credentials** that were readable from the affected environments: environment variables injected into the job, files in the workspace, credential files in a developer's home directory, and any ambient role or token the runner could assume. - **Artifacts** already produced on those machines. An image built on a hostile runner is suspect even if the malicious package does not appear anywhere in that image's own dependency list — the build environment, not the dependency list, was the problem. - **Actions already taken** with anything that was stolen: artifacts published, keys or users created, webhooks or automation added. ## The order that shows you understand it 1. **Scope the environments.** Which laptops, runners and build hosts resolved the bad version inside the exposure window. 2. **Cut re-entry first.** Revoke or rotate anything that grants deploy, publish or signing power before anything else — otherwise you clean a system an attacker can immediately re-poison. 3. **Rotate the rest**, ranked by what the credential reaches. 4. **Rebuild** artifacts produced on affected machines during the window, on a host you trust. 5. **Then** fix the dependency — remove, pin or roll forward — and only then redeploy. Notice that the dependency fix, which is where the naive answer starts, is step five. ## Wrong answers worth recognising - *"The registry yanked the version."* A yank stops new installs. It un-runs nothing, and it does not tell you who already pulled it. - *"We rolled back the deploy."* Rollback restores the running version of your service. It does not rotate a token that left the building an hour ago. - *"Only CI installed it, and CI is ephemeral."* Ephemeral runners limit persistence on the host; they do not limit what a job could read while it was alive. A short-lived token still had full authority during its life. - *"The payload only read environment variables."* Analysis of the published payload is useful for *prioritising* rotation, not for skipping it — and what you fetched for analysis is not necessarily what was served to every consumer. ## A worked framing A nightly ETL job pulled a malicious version of one of its dependencies. The job is unglamorous — it moves rows — but it holds a credential with write access to the production ledger. The blast radius here is not customer personal data; it is money and audit truth. The response is the same shape: assume the credential is gone, rotate it, and then go read the ledger's own change history for the window, because rotating the credential does not undo a write it already made. ## What a strong answer sounds like "Assume every environment that installed it during the window ran attacker code. Rotate everything readable from those environments, starting with anything that can deploy or publish. Rebuild artifacts they produced. Fix the dependency last, and verify by rebuilding from a clean host — not by looking at the lockfile."

  • Which machines count as having installed it — just the CI runners?
    Every host that resolved the version in the window: CI runners, developer laptops, image and base-image builds, and any pre-warmed dependency cache that fetched it. Laptops are frequently the worst of them, because a personal token there usually has broader scope than a runner's job-scoped role.
  • Our application never imports that package. Does that narrow the scope?
    No. Whether your code calls the package is a reachability argument, and reachability matters for a vulnerable dependency, not a malicious one. The code ran at install and build time regardless of whether anything imported it. It may narrow runtime impact on the deployed service; it does not narrow host or credential exposure.
  • The registry pulled the version within the hour. Are we finished?
    No. A yank prevents new installs and nothing else. It does not tell you who fetched it during the window, does not clear caches or mirrors that already have it, and does not touch credentials or artifacts. It shortens the window, which is useful for scoping, and that is all.

Firing a contractor stops them coming back tomorrow. It does not get back the keys they already copied, or undo the work they did while they had them.

saying these in an interview costs you the question

  • Says reverting the lockfile undoes the compromise
  • Assumes nothing ran because the app never imported it
  • Rotates only the token named in the failing pipeline
  • Treats a registry yank as remediation
  • Waits for a clean scan before rotating anything

context

open as a page

After a malicious dependency version lands, do you freeze the lockfile, pin around it, or roll forward?

level: seniorimportance: must knowfreq 54%

basics

~20 s

Usually all three in order: freeze to stop further drift, pin or override the bad transitive package as the interim fix, then roll forward to a version you can justify trusting. Order by what is fastest to undo.

open as a page

How do you determine which systems actually ran a malicious dependency version?

level: middleimportance: should knowfreq 58%

basics

~20 s

Not from the current lockfile, which describes today's intent. Reconstruct it from lockfile history across the exposure window, from the deployed inventory of running artifacts and their manifests, and from registry or caching-proxy pull records.

open as a page

Which credentials do you rotate first after a malicious build plugin ran on your CI runners?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Anything that lets an attacker back in or ship code: deploy roles, publish and registry tokens, signing keys. Then data-plane credentials ranked by what they reach. Rotation alone is not remediation — also audit what the stolen credentials already did.

open as a page