A nightly build that always resolved an internal Python package from the company index installed a public copy of the same name, with no manifest change — how do you work out what let it in?
answer
- the build went green — that is the problem
- what changed was outside the manifest
- which host served each artifact
- not-found, outage, or cold cache
- assume the runner's secrets are gone
basics
~20 sSomething stopped the internal source answering for that name — a not-found on a new or transitive requirement, an outage, or a cold cache on a fresh runner — while a public source stayed reachable. Start from the build's resolution log.
solid answer
~50 sNothing in the manifest changed, so what changed is either the name's availability internally or its availability publicly. Pull the resolution and download log for that run and establish, per artifact, which host served it, at which version, and when. Then check three things: whether the installed version exists internally at all, whether the internal index logged errors, timeouts or not-founds in that window, and when the public copy of the name was first registered — a name freshly published by a stranger is an attack, not an accident. Also check whether the requirement is newly introduced or transitive, since a name the internal index never held falls straight through. In parallel treat the runner as compromised: rotate every credential the build environment could reach, because exfiltration happens the moment the package is installed and the build goes green either way.
go deeper
Know that a build can fetch a package from a different source than expected without anyone editing the manifest, and that the first thing to look at is the build log showing where each package came from.
Explain the fall-through paths — the internal source answering not-found for a new or transitive name, an outage causing failover, a fresh runner with no cache — and why a newly registered public name changes the outcome with no change on your side.
Demonstrate a real investigation order: per-artifact source evidence, internal index errors in the window, public first-publish facts, requirement provenance, then immediate credential rotation for everything the runner could reach.
Own the detection gap rather than the single incident: decide that every build must be able to answer which host served each artifact, and that an internal name resolved externally is an alert, then fund the work to make that true across the estate.
## Why this is hard: the build succeeded The defining property of this failure is that it does not look like a failure. The pipeline is green, the tests passed, the artifact published. The only signal is that one dependency came from somewhere else, and unless the build records where each artifact came from, there is nothing to notice. Treat 'which source served each artifact' as the evidence you must be able to produce; if you cannot produce it, that is the first finding. ## Three ways a public copy gets in without a manifest change **Not-found fall-through.** The internal index is asked for a name it does not hold and answers not-found, so the client moves on to the next configured source, which does hold it. The name may be new (a requirement added inside another internal package rather than in this manifest), renamed, yanked, or simply never published internally for the platform or interpreter version this runner used. This is the most common path, and it is the one that makes a *transitive* requirement so dangerous — nobody in the change review ever saw the name. **Availability failover.** The internal index times out, returns a server error, or briefly rejects the build's credentials. Clients configured with more than one source are frequently tolerant of a source being down, and tolerance means continuing with the sources that answered. A short internal outage therefore becomes a resolution decision. **Cold cache.** Developer machines and long-lived runners had the internal artifact cached, so the question of where it comes from never arose. A fresh ephemeral runner resolves from scratch, and that first honest resolution is the one that takes the public copy. This is the usual explanation for 'it worked for months'. Overlaying all three is the trigger that changes nothing on your side: **the name was registered publicly for the first time**. Until that moment the public source had nothing to offer, and every fall-through resolved harmlessly or failed loudly. The day someone claims the name, the same configuration starts producing a different outcome. ## The insider variant worth naming Consider a retail bank whose data-platform builds are configured with a public index alongside the internal one. A contractor holds read-only access to the internal Python index — legitimate, unremarkable access. They browse the private wheel names, register three of them on the public index at higher versions, and wait. The build agent installs them and the package reads the environment it landed in: cloud credentials, index tokens, source checkouts. Nothing the contractor did inside the bank looked like an attack, and the write happened somewhere the bank has no authority at all. This is why 'they only had read access to the index' is not a mitigating fact in this attack class — names are the whole input. ## What to collect, in order 1. **Resolution evidence for the run**: for the offending package, the host that served it, the version, the digest and the timestamp. If the client's default verbosity does not record the host, that is a gap to close before anything else. 2. **Internal availability**: does that version exist internally? Did the internal index log 404s, 5xx, auth failures or latency in that window? Was the package recently renamed or removed? 3. **Public registration facts**: when the public copy was first published, by whom, and whether several of your internal names were registered around the same time — clustering across your namespace is close to conclusive. 4. **Requirement provenance**: is the name declared directly, or pulled in transitively by an internal package whose own requirement string is an open range? 5. **Content**: does the package do anything on install or import beyond what a stub would, and does it resemble your library at all? An empty package with collection behaviour is the signature of an opportunistic registration. ## Assume the environment is burned Treat every credential reachable from that build environment as disclosed: cloud role credentials, registry and index tokens, signing material, source access, anything in environment variables or on disk. Rotation is not a post-incident tidy-up item — the exfiltration, if any, already happened at install time and the window is measured in seconds. Then look outward: egress from the runner in that window, and whether the resulting artifact was published or deployed anywhere, because the blast radius follows the artifact. ## Telling attack from accident A public package can innocently share a name. Publisher identity, first-publish date, whether the content is a plausible unrelated project, and whether the version is absurdly high all separate the cases. Absent an innocent explanation, treat it as an attack, and remember that the durable answer is not to fix this one name but to close the fall-through path so an internal name can never resolve externally — a build-configuration decision, and a separate conversation from this diagnosis.
- How would you tell this apart from an unrelated public package that happens to share the name?Look at publisher identity and first-publish date, whether the version is implausibly high, whether the content is a real project or a stub with collection behaviour, and whether several of your internal names were registered around the same time. Clustering across your namespace, by a new account, at high versions, is effectively conclusive.
- The person who published only had read access to the internal index. Does that lower the severity?No. Read access to the namespace is the entire prerequisite, because the malicious write happens on a public index you do not control. Severity should be judged by what the build environment held — credentials, tokens, source — not by the privilege level the attacker had inside your systems.
- What would you want logged so this takes minutes to answer next time?Per-artifact resolution records for every build: name, version, digest, serving host and timestamp, retained and queryable. On top of that, an alert whenever a name in an internal namespace is resolved from an external host — that single rule turns a silent success into a page.
saying these in an interview costs you the question
- The build succeeded, so nothing needs investigating
- The manifest did not change, so it cannot be a dependency issue
- Read-only access to the internal index is harmless
- Credential rotation can wait for the postmortem
- The internal source is listed first, so it always wins
- Delete the public package and the incident is closed