Why can a code reviewer miss a homoglyph dependency name in a manifest diff?
answer
- eyes read meaning, not glyphs
- confusable characters and letter pairs
- capital I, Cyrillic vowel, rn for m
- normalise to a skeleton before comparing
- the lockfile then freezes it
basics
~20 sReviewers read names for meaning, not glyph by glyph, and confusable characters render nearly identically: a Cyrillic vowel, a capital I standing in for lowercase l, the pair rn for m. Catching these needs mechanical normalisation, not attention.
solid answer
~50 sHuman review is the wrong instrument for this. A reviewer scanning a manifest recognises a dependency as a familiar token — "the analytics SDK" — rather than decoding it character by character, and the substitutions attackers use are chosen to survive exactly that: non-Latin lookalikes such as a Cyrillic vowel, Latin ambiguities like capital I for lowercase l or zero for O, and multi-character pairs like `rn` reading as `m` in a proportional font. The mechanical answer is to normalise before comparing: case-fold, fold separators, map every character to a canonical skeleton so confusables collapse together, and then compare the candidate against the inventory of names you already depend on. A skeleton collision with a name you use is a strong signal; a small edit distance is a weaker one. Ingest can also simply refuse non-ASCII package names, which legitimate packages rarely need.
code
json · 7 lines{
"dependencies": {
"streamline-utils": "^2.3.1",
"streamIine-utils": "^2.3.1",
"...": "..."
}
}go deeper
Be ready to name a couple of concrete lookalikes — a capital I for lowercase l, rn reading as m, a Cyrillic vowel — and to say that eyeballing a manifest is not a control.
Explain the mechanics: normalise case and separators, fold confusables to a skeleton, then compare against names you already use, and say why a skeleton collision is a stronger signal than a small edit distance.
Show where the check belongs in the flow and what happens after approval — the lockfile pins the wrong name and every later build reproduces it faithfully, so detection after merge is far more expensive than screening at ingest.
Weigh a hard block against a routed review: decide which signals justify stopping a build outright, who resolves the rest, and how you keep false positives from teaching engineers to bypass the check.
## The classes of near-miss name Attackers are not limited to dropped letters. Roughly, in increasing order of how badly they defeat the eye: 1. **Letter-level typos** — a dropped, doubled or transposed character. Visible on careful reading. 2. **Latin confusables** — capital `I` for lowercase `l`, digit `1` for either, digit `0` for `O`. Invisible in most sans-serif proportional fonts, which is what code-review web UIs and pull-request descriptions use. 3. **Multi-character confusables** — `rn` read as `m`, `cl` read as `d`, `vv` read as `w`. These survive even a monospace font at small sizes. 4. **Unicode homoglyphs** — Cyrillic and Greek vowels, fullwidth forms, and other codepoints that render identically or nearly so to their Latin counterparts. Nothing about the rendered string is wrong; the bytes are different. 5. **Separator and case variants** — the same words joined with `-`, `_`, `.` or nothing, or a different case pattern. 6. **Padding** — a suffix or prefix the ecosystem uses constantly, so a name gains or loses `-js`, `-core`, `-utils` or a language prefix and still reads as the same thing. 7. **Namespace shift** — the identical package name published under a different owner or scope, where the eye reads the familiar half and skips the owner. ## Why review does not catch classes 2 through 7 Reading is pattern matching against expectation. A reviewer looking at a dependency bump has already decided what the diff is *about*, and the package name is a label confirming that decision rather than data to be verified. Three further things work against them: - **The rendering is honest.** With a homoglyph there is no visual defect to notice, only different bytes behind an identical picture. - **The lockfile is consistent.** After resolution, the wrong name appears many times with matching integrity hashes and a coherent transitive tree. Internal consistency reads as normal, and a large lock diff is skimmed by almost everyone. - **Approval happens once, the effect is permanent.** The lockfile now pins that name to a digest, so every subsequent build faithfully and reproducibly fetches the attacker's artifact. Reproducibility guarantees you keep getting the same wrong thing; it does not tell you the thing was right. Consider the shape of the worst case: a mobile release build resolves a homoglyph-named analytics SDK, the artifact is signed and shipped through a store, and it lands on a million handsets. The asset at risk is customer data on end-user devices, and the remediation is a store update — you cannot recall what has already been installed. This is why the control has to fire at ingest rather than at incident response. ## The mechanical check that works Compare a *normalised* form of the candidate name against a *known set*: - **Normalise:** lowercase; fold `-`, `_` and `.` runs to a single separator or remove them; then map each character through a confusable table to one canonical representative per lookalike group (the "skeleton"). Some indexes already normalise case and separator runs themselves, which closes class 5 at the ecosystem level but does nothing for homoglyphs. - **Compare against two sets:** the names you already depend on — which is the set that matters, because the attacker is impersonating something *you* use — and a list of the most-downloaded names in the ecosystem. - **Grade the signals:** an identical skeleton to a name you already use is close to conclusive and deserves a hard stop. An edit distance of one or two is suggestive and deserves a human look, not a block, because legitimate near-names are common. A non-ASCII codepoint anywhere in a package name is rare enough in practice to be worth quarantining on its own. - **Check the owner too**, not just the string. Same name, different namespace or publisher is a distinct case that string distance scores as zero. ## What this is not This is a detective control on the *name*. It says nothing about whether the package that carries the right name is safe, and it cannot see an invented name that resembles nothing you already use. Pair it with a rule that routes never-before-seen names to a human, and keep the decision so it is made once.
- What would you compare a candidate package name against, and why that set?Primarily the names your own estate already depends on, because a near-miss attack only pays off if it impersonates something you actually pull; secondarily the ecosystem's most-downloaded names, to catch impersonation of a package you are about to adopt. Comparing against the whole index is noise — millions of legitimate names sit within a small edit distance of each other, so the false-positive rate makes the check unusable.
- Is an edit-distance threshold enough on its own?No. It scores zero for a namespace or owner change, misses multi-character confusables unless you normalise first, and produces constant false positives because ecosystems are full of legitimate near-names such as a package and its `-core` or `-utils` companion. Use skeleton collision as the strong signal, edit distance as a routing hint to a human, and check publisher identity separately from the string.
- Why does pinning to a digest not help here?A digest pins content — it guarantees you get the same bytes every time. It says nothing about whether those bytes were the right package to begin with. Pinning the wrong name simply makes the compromise reproducible, and the resulting builds look impeccably deterministic while shipping the attacker's code on every run.
Proofreading your own writing: you see the word you intended, not the letters actually on the page. The attacker is writing for that reader, not for a parser.
saying these in an interview costs you the question
- Says careful code review is enough to catch it
- Confuses pinning to a digest with verifying the name
- Applies edit distance without normalising confusables first
- Compares against the whole index instead of your own dependencies
- Ignores publisher or namespace changes because the string matches