A company publishes an internal package called `acme-auth-utils` on a private registry, but a developer's laptop still has the default public registry configured. One day a build silently pulls in a different `acme-auth-utils` package that was never written by anyone at the company. What attack is this, and how does the resolution logic actually let it happen?
answer
- Birsan 2021 bug bounty
- flat namespace, highest version wins
- scope→registry pinning closes it
- reserve the name publicly too
- install scripts run before app code
basics
~20 sAn attacker publishes a public package with the exact same name as a company's private internal package. If a build tool checks the public registry and grabs whichever version looks newest/highest, it installs the attacker's fake package instead of the real internal one — no phishing or malware download needed, just a naming collision.
solid answer
~50 sThis is dependency confusion (popularized by Alex Birsan's 2021 research). Many package managers historically treated all configured registries as one flat namespace and resolved a bare package name to whichever source had the highest matching version, without regard to which registry 'owns' that name. An attacker who learns or guesses an internal package name — often from a leaked package.json, a public repo, or job postings — publishes a same-named package publicly with an artificially high version number. Any machine (dev laptop, CI runner) that isn't strictly pinned to pull that name only from the internal registry will resolve to the attacker's package and execute its install scripts. The fix is registry scoping: explicitly map internal package names/scopes to the private registry so the public registry is never consulted for them, plus reserving the name publicly as a placeholder.
go deeper
Can explain in plain terms that a same-named public package can get installed instead of the private one, and that this is dangerous because install scripts run automatically.
Knows the concrete mitigation — registry/scope pinning in .npmrc/pip config — and can explain why a lockfile alone doesn't prevent the first poisoned resolution.
Designs and enforces the control organization-wide: registry proxy as the single install path, reserved placeholder packages, and CI checks that fail a build if a dependency resolves from an unexpected source.
Weighs this against the broader supply-chain program: prioritizes which of dozens of ecosystems (npm, pip, Maven, NuGet, Go modules) actually carry sensitive internal names, sequences rollout of registry proxying without breaking existing pipelines, and defines the exception process.
## The assumption it exploits Dependency confusion (sometimes called a **namespace** or **substitution attack**) exploits an assumption baked into many package managers' default resolution logic: that a bare package name lives in one, unambiguous namespace, and that when multiple registries are configured, the 'right' version to install is simply the highest version number that satisfies the requested range, regardless of which registry it came from. Large organizations commonly publish internal packages — shared auth helpers, logging wrappers, build tooling — to a private, company-run registry (npm's private registries, an internal PyPI mirror, an Artifactory/Nexus instance) while also relying on the public registry (npmjs.com, PyPI, Maven Central) for everything else. If a developer's or CI machine's package-manager configuration doesn't explicitly say 'this name/scope always comes from the private registry, full stop,' the tool will happily check both, and **the highest version wins**. ## How the attack runs An attacker doesn't need to compromise anything — they just need to learn or guess the name of an internal package and publish a same-named package publicly with a very high version number, such as `999.0.0`. Where the name comes from: - leaked in a public GitHub repo - a job posting mentioning internal tooling - an error message - or simple enumeration Once installed, the attacker's package runs arbitrary code immediately via install-time hooks (npm's `postinstall`, Python's `setup.py`) — long before any of the application's own code executes — typically to exfiltrate environment variables, CI secrets, or SSH keys, or to establish a foothold for further compromise. ## The trade-off underneath This class of attack exists because package ecosystems were built for openness and low friction: anyone can publish a package under any unclaimed name, and resolution logic was optimized for convenience (grab the newest thing that satisfies the version range) rather than provenance (only trust names from where you registered them). That design choice is exactly the trade-off: an open, flat namespace lets ecosystems grow fast and lets small teams publish without gatekeeping, but it also means 'internal' and 'private' are conventions, not enforced boundaries, unless someone configures the tooling to enforce them. The cost of closing the hole is operational: every package manager, every CI image, and every developer machine in the org has to be configured with explicit registry-to-scope mappings (for example, npm's `.npmrc` supporting `@acme:registry=https://npm.internal.acme.com`), and that configuration has to be enforced consistently — one misconfigured laptop or one unpinned CI job is enough to open the door. ## The failure mode in production The failure mode in production is almost always **silent and fast**: a routine `npm install` or `pip install -r requirements.txt` in CI pulls the malicious package, its install script runs with whatever privileges the CI job has (often including registry publish tokens, cloud credentials, or access to the deploy pipeline), and the compromise can be complete before anyone notices a diff in the lockfile — if the lockfile was even regenerated rather than reused. Detection typically comes from anomaly monitoring (unexpected outbound network calls during a build), a security researcher's disclosure, or, worse, an actual incident. ## The canonical case The canonical real-world case is Alex Birsan's 2021 dependency confusion research, in which he registered public npm/PyPI/RubyGems packages matching internal package names he found referenced in leaked `package.json` files, build logs, and JavaScript source of major companies. His proof-of-concept packages 'phoned home' rather than doing anything harmful, but they successfully executed inside internal build systems at Apple, Microsoft, PayPal, Netflix, Uber, Shopify, and roughly 30 other companies, each of whom paid bug bounties (Birsan collected over $130,000 total). The research made clear this wasn't a theoretical risk but a systemic gap across essentially every major package ecosystem. ## Mitigation is layered Mitigation is layered rather than a single fix. 1. **First, explicit registry scoping/pinning** so internal names are never resolved from the public registry — npm scopes (`@acme/...`) mapped to the private registry, pip's `--index-url/--extra-index-url` locked down rather than left to fall through, and equivalent controls for Maven/Gradle repository ordering. 2. **Second, defensively reserving the internal package names** on the public registries too, even as empty placeholder packages, so an attacker can't claim them. 3. **Third, using a registry proxy/mirror** (Artifactory, Nexus, npm's private registry with upstream proxying) as the single source of truth so individual machines never talk to the public registry directly and policy is enforced centrally rather than per-laptop. 4. **Fourth, lockfile integrity checks** and CI-side verification that installed package hashes match an approved manifest, which catches drift even if resolution briefly goes wrong. None of these is sufficient alone — the Birsan disclosures hit companies with mature security programs precisely because registry scoping wasn't uniformly enforced everywhere dependency resolution happened.
- Why doesn't a lockfile (package-lock.json, poetry.lock) fully protect you from this on a fresh clone?A lockfile only pins what was resolved the first time it was generated — if that initial resolution already picked up the attacker's public package (because registry scoping wasn't configured when the lockfile was created), the lockfile will faithfully keep re-installing the malicious version on every subsequent install. Lockfiles protect against future drift, not against a poisoned baseline.
- What's the single most effective organizational control against dependency confusion?Routing all package installs through a private registry proxy (Artifactory, Nexus, or the ecosystem's own private-registry feature) configured with an explicit, enforced mapping of internal scopes/names to the internal source, so individual developer and CI configuration can't accidentally fall through to the public registry. Centralizing the policy removes the 'one misconfigured machine' failure mode.
- Is reserving your internal package names on the public registry enough by itself?No — it blocks name-squatting but doesn't fix the underlying resolution ambiguity or protect ecosystems/tools that don't require reservation, and it's easy to miss a name (e.g., a new internal package) before an attacker claims it. It's a cheap, useful supplement to explicit registry scoping, not a replacement for it.
It's like a mailroom that delivers to whoever answers to a name first, without checking which company that name is registered to — if an impostor shouts 'I'm Acme Corp' louder (a higher version number), the mailroom hands them Acme's package instead of the real Acme.
saying these in an interview costs you the question
- thinks lockfiles alone prevent this on first install
- doesn't know registry scoping/private registry config exists
- conflates with typosquatting
- assumes private repo name is automatically protected