skip to content

Why does an exact-match term screen miss a blocked word when a user swaps in a look-alike codepoint?

level: juniorimportance: must knowfreq 64%

answer

  1. two relations, not one
  2. the screen never looks at the glyph
  3. renders alike, compares unequal
  4. codepoints are characters, not shapes
  5. U+0430 is not U+0061

basics

~20 s

The screen compares codepoints, not appearance. A Cyrillic letter drawn like a Latin one is a different codepoint, so the word no longer equals any stored term and the comparison returns no match, while the glyphs on screen are unchanged.

solid answer

~40 s

A term screen holds a list of stored strings and compares the incoming message against them. String equality is equality over codepoints: two strings match when they are the same sequence of codepoints, and how they will be drawn never enters into the comparison. Unicode assigns codepoints to characters, not to shapes, and many distinct codepoints are drawn with the same glyph — U+0430 CYRILLIC SMALL LETTER A is drawn like U+0061 LATIN SMALL LETTER A, and Greek omicron U+03BF like Latin `o`. Replace one letter with its look-alike and the rendering is identical while the comparison is unequal. Two relations people treat as one come apart: renders-alike is about glyphs, compares-equal is about codepoints. Nothing is hidden here — the character is fully visible, it is simply a different character.

go deeper

for a junior

Be ready to say that string comparison runs over codepoints and that two different codepoints can be drawn as the same glyph. Name one concrete pair from different scripts.

for a middle

Explain what the screen compares and at which stage it runs, and state precisely what a no-match verdict does and does not prove about the message that produced it.

for a senior

Expect to be pushed on evidence: which stage records exist for one message, what each can attribute, and how you separate the screen staying silent from the model not refusing.

for a principal

Own the framing that a surface-string comparison and a meaning-reading model answer different questions about the same input, and say what an assurance claim resting on the first can honestly cover.

## What the screen actually compares In a consumer chat product with no tools — say a study-help assistant where the only thing that can leave the system is text on the user's screen — there is often very little to defend, so what sits in front of the model is a term screen: a list of stored strings and a comparison. Before the message reaches the model, the product asks whether the message contains any stored term. In the simplest and most widely deployed form that comparison is exact: is this run of characters *equal* to that stored string? String equality is equality over codepoints. Two strings are equal when they are the same sequence of codepoints and unequal otherwise. How the sequence will be drawn on a display is not part of the question the comparison answers. ## What Unicode does and does not guarantee Unicode assigns codepoints to *characters*, not to *shapes*. Nothing in the standard says that two characters drawn the same way are the same character — in fact a great many distinct codepoints are drawn identically or near-identically in ordinary fonts, because different scripts inherited the same letterforms: - U+0430 CYRILLIC SMALL LETTER A is drawn like U+0061 LATIN SMALL LETTER A - U+0435 CYRILLIC SMALL LETTER IE is drawn like U+0065 LATIN SMALL LETTER E - U+03BF GREEK SMALL LETTER OMICRON is drawn like U+006F LATIN SMALL LETTER O So the word a person types, with one Latin letter replaced by its Cyrillic counterpart, renders exactly as the original renders and compares unequal to it. This is the whole property. Two relations that everyday intuition fuses are in fact separate: **renders-alike** is a property of glyphs and fonts; **compares-equal** is a property of codepoint sequences. Any control that matches against stored strings is answering the second question. The substitution changes the second while leaving the first untouched. ## What the construction costs the person doing it Almost nothing, which is why it is the first thing tried against a term screen. In this shape of product the attacker and the user are the same person and the channel is their own message field: no document to plant, no retrieval window to wait for, no third party to involve, no reviewer to get past. One substituted codepoint in a message they were already going to send. Note what it is *not*. This is not concealment. The glyph is on screen and legible. A character that renders as nothing at all is a different construction with different properties; here the point is that the character renders as the expected shape. ## What a miss proves, and what it does not This is where triage usually goes wrong. A no-match verdict from the screen proves exactly one thing: **under that comparison, the message equalled no stored term.** It does not prove the message is harmless, it does not prove anything about what the model will do with it, and it does not prove that no other stage fired. If the model then produces content it would ordinarily decline, two independent events have occurred: the screen did not match, and the model's own trained refusal did not fire. Those are different mechanisms with different failure modes, and reporting them as one event — "the filter was bypassed" — hides which of the two actually let the answer through. ## Where the class stops working - **At the model.** The substituted word is a different byte sequence and therefore a different token sequence. The model may or may not recover the intended word from context; the construction has a success rate, not a guarantee. - **At any stage whose comparison is not raw equality over the stored strings.** The moment the comparison stops asking the identity question, the property being bought stops paying. - **At the payoff ceiling.** In a product with no tools, no store to write to and no second reader, the entire prize is text on screen that would otherwise have been declined. There is nothing else in reach. The short version an interviewer wants: the screen compares identity, the substitution preserves appearance and changes identity, and those are different relations over the same string.

  • Does the substituted character have to be invisible for this to work?
    No, and it is not invisible. The codepoint renders a normal visible glyph — it just renders the *expected* glyph. That is a different property from a character that renders as nothing at all, which is a separate construction with its own survival problem through copy and transport. Here the string still reads naturally to anyone who sees it, and that costs the person typing it nothing.
  • The screen returned no match and the model answered. What does that pair of facts prove?
    Only that the message equalled no stored term under that comparison, and that the model's own refusal did not fire on this turn. It does not prove the message was harmless, that the screen is broken, or that the outcome is repeatable — those are three separate claims, and a report that runs them together cannot be triaged to an owner.
  • Why would a screen be built on exact comparison at all if this is so cheap to get past?
    Because exact comparison is fast, has no false positives against ordinary text, and is trivially auditable — you can say precisely which term fired. Those are real properties. The point is not that the comparison is badly built but that it decides identity, and identity is the one relation a look-alike substitution is free to change.

Two banknotes can be printed from the same plate and still carry different serial numbers. A cashier checking the picture and a cashier checking the serial are asking different questions about the same note.

saying these in an interview costs you the question

  • Says the substituted character is invisible or hidden
  • Assumes strings that look identical must compare equal
  • Thinks a no-match verdict means the message was harmless
  • Confuses the screen staying silent with the model refusing
  • Believes the model sees the same bytes the screen compared

context