Your ingest gate blocks any dependency name within edit distance two of one you already use, and developers hit false positives weekly. How should that policy work?
answer
- similarity is a signal, not a verdict
- near-names are legitimate and common
- route to quarantine, decide once
- the allow-list is the durable artifact
- watch for developers routing around it
basics
~20 sName similarity is a routing signal, not a verdict. Route flagged and never-before-seen names into a quarantine lane where a person decides once, record every decision on an allow-list, and watch for teams bypassing the gate.
solid answer
~50 sA hard block on edit distance two will misfire forever, because healthy ecosystems are full of legitimate near-names — a package and its companion, a language-prefixed port, a scoped fork. Rather than tuning the threshold until it is useless, split the decision: similarity is a signal that *routes* a name into quarantine, and a person makes the call once. Reserve a hard stop for the near-certain case, an identical confusable-folded skeleton against a name you already use. The allow-list is the durable artifact: once a name is approved, no later build pays anything, so steady-state cost tracks new dependency adoptions rather than build volume. Add a separate first-time-seen rule, the only thing covering an invented name that resembles nothing you use. Then measure whether people are vendoring code to avoid the gate — a bypassed gate is worse than none, because it also removes your visibility.
go deeper
Be ready to explain what an allow-list and a quarantine lane are for, and why blocking a name that merely looks similar to one you use will sometimes be wrong.
Explain why similarity has an irreducible false-positive rate in real ecosystems, and how normalising names before measuring distance makes the signal meaningful rather than noisy.
Design the flow end to end: which observation hard-blocks, which routes to quarantine, who decides, and how the recorded decision makes the check cost nothing on later builds.
Own the trade between coverage and latency, state the steady-state cost in terms leadership can weigh, and treat shadow ingest — not denial counts — as the measure of whether the control is real.
## Why the strict rule is unfixable as stated Edit distance measures string proximity, and string proximity is not the same as impersonation. Real ecosystems produce near-names constantly: a core package and its `-utils` companion, a language-prefixed binding of the same library, a scoped fork, two unrelated projects that both wanted a common word. Any threshold loose enough to catch a determined near-miss will flag those, and any threshold tight enough to stop flagging them will miss a single-character substitution. There is no setting that resolves this, because the metric cannot distinguish intent. Treating it as a verdict guarantees a steady stream of blocked-but-legitimate packages, and the predictable result is not more security but a workforce that learns to route around the gate. ## Reframe: signal, lane, decision, record **Signal.** Compute similarity *after* normalising — lowercase, fold separator runs, map confusables to a canonical skeleton — and compare against the set that actually matters: the names your estate already depends on, plus the ecosystem's most-fetched names. Grade the outcome rather than thresholding it: | Observation | Response | |---|---| | Identical skeleton to a name you already use | hard stop; this is impersonation until proven otherwise | | Non-ASCII characters in a package name | quarantine; rare enough to be cheap | | Edit distance one or two from a name you use | quarantine for a human look | | Same name, different publisher or namespace | quarantine; string distance scores this as zero | | Never seen before anywhere in the estate | quarantine, regardless of similarity | **Lane.** Quarantine means the name is fetched into an inspection area and is not resolvable by builds until released. It is a lane, not a wall: the point is that the request keeps moving toward a decision instead of returning an error to a developer with no next step. **Decision.** A named rota, a stated response time measured in hours rather than days, and a documented escape path for urgent cases that produces an audit record instead of a bypass. Who is on that rota is an organisational choice with real cost, and pretending the check is free is how these programmes die. **Record.** The allow-list is the artifact that makes the whole thing affordable. Every decision — approved, rejected, and why — is stored against the name, so a name is reviewed once and never again. That turns the steady-state cost into a function of how many *new* dependencies the organisation adopts per week, which is a much smaller number than most people expect and is the figure to put in front of anyone questioning the programme. ## The gap the similarity check cannot cover Similarity detection assumes the attacker is impersonating something. An invented name — the kind a coding assistant produces, which an attacker then registers — impersonates nothing in your inventory, so it sits at a large distance from every known name and sails through. The rule that catches it is orthogonal: **first time this organisation has ever seen this name**. Any serious ingest design needs both, and being able to say which control covers which attack is the answer that separates a designed gate from a copied one. Neither rule says anything about whether a correctly-named package is safe; both are about *identity*, and they sit in front of, not instead of, whatever review you do on the code itself. ## What to measure - **Median and tail time in quarantine.** The tail is what people remember and what drives bypass. - **Share of requests auto-approved from the allow-list.** This should climb steadily; if it does not, either your inventory is churning or the list is not being written back. - **False-positive rate on flagged names**, tracked per rule, so you can retire a rule that never catches anything. - **Shadow ingest.** Vendored copies appearing in repositories, direct source-URL dependencies, installs performed on laptops. This is the metric that actually decides whether the gate works, and it is the one most programmes never look at. A bypassed gate is worse than no gate, because leadership believes dependencies are screened while visibility has quietly moved outside the system. ## The judgment to state out loud Security controls that produce frequent, unexplained denials get engineered around, and the engineering-around is invisible. So the design goal is not maximum blocking; it is maximum *coverage of first-time names* at a latency engineers will tolerate. Buy that latency with automation on the boring cases and spend human attention on the small number of genuinely novel names — that is the trade, and it is one a lead should be able to defend in numbers.
- Which observation would you allow to hard-block a build outright?An identical name after confusable folding and separator normalisation, matching something the estate already depends on. That is impersonation of a package you actually use, and there is no benign explanation worth the latency of a review. Everything softer — edit distance one or two, a publisher change, an unfamiliar name — routes to quarantine instead, because those have legitimate explanations often enough that blocking teaches people to avoid the gate.
- Which control covers an invented name that resembles nothing you already use?Only the first-time-seen rule. Similarity metrics assume impersonation of a known name, and an invented name is far from every name in your inventory, so no threshold fires. Routing never-before-seen names to a human decision is the control that covers it, and because the decision is recorded on an allow-list, it costs nothing on any subsequent build.
- How do you justify the programme's cost to engineering leadership?Show the steady-state number, not the launch number. Because decisions are recorded per name, the recurring cost is proportional to genuinely new dependency adoptions per week rather than to build volume, and that figure is usually small. Pair it with median and tail quarantine latency, and with the shadow-ingest measurement, so the conversation is about a measured tax and its coverage rather than about a feeling of friction.
- What single metric tells you the gate is failing?The rate of dependencies entering the estate outside it — vendored copies committed into repositories, direct source-URL dependencies, packages installed straight onto laptops. Denials and queue length describe friction; bypass describes failure. A gate people avoid is worse than no gate, because the organisation believes screening is happening while the actual ingest path has moved somewhere nobody is watching.
An airport screening lane: the detector's job is to route a traveller to a person, not to make the final call. A machine that simply refuses entry on a beep would empty the terminal, not secure it.
saying these in an interview costs you the question
- Tunes the edit-distance threshold instead of splitting the decision
- Treats every similarity hit as a hard build failure
- Forgets that invented names trip no similarity check
- Reviews the same package name on every build
- Measures denials but never measures bypass