skip to content

A look-alike substitution report says "a person read it and it looked normal" — what did it get wrong?

level: seniorimportance: should knowfreq 40%

answer

  1. enabling property, not defeated control
  2. the eye was never asked the question
  3. identity changed, appearance did not
  4. two mechanisms, two different questions
  5. a limit files differently from a defect

basics

~20 s

It named the wrong cause. Looking normal is the property the substitution buys, not the control it defeated. The stage that was defeated compares identity, and nobody was ever asked to compare that string against a list.

solid answer

~40 s

Rendering normally is not the failure — it is the enabling property, deliberately preserved so the message still reads as language. The control actually got past is the term screen's comparison, which decides whether a string equals a stored one; the substitution changes exactly that relation and nothing else. The human in the story is not a control at all: nobody asked them to check the string against a list, they read for sense and got sense. Written as a root cause, the finding should say that two relations over one string — renders-alike and compares-equal — were treated as one, so a stage deciding identity and a model deciding meaning disagreed about the same input. That framing also gives the verdict: a property of the arrangement, not an oversight by a reader.

go deeper

for a junior

Learn the separation first: appearance is what the substitution keeps, identity is what it changes, and only the second one has anything to do with the stage that stayed quiet.

for a middle

Be able to write the root cause in one sentence that names both mechanisms, says which question each answers, and blames neither of them individually.

for a senior

Show triage judgement: route the finding away from the reader, state impact at the ceiling the product actually allows, and report the observed rate rather than a single success.

for a principal

Own the bug-versus-limit verdict, and be able to say what a deployment relying on surface-string comparison in front of a meaning reader can be claimed to prevent at all.

## Why the wrong cause is so attractive "A person read it and it looked normal" feels like an explanation because it describes something true that happened. It is still the wrong cause, and the mistake is worth being precise about, because it determines who the finding goes to and what a verdict on it can honestly say. The substitution does two things to one string: 1. It **preserves appearance** — every substitute was chosen to render the same glyph. 2. It **changes identity** — the codepoint sequence is now different from any stored term. Only the second one gets past anything. The first is the property being *bought*: it is what keeps the message readable as ordinary language so that the model behind the screen still has something to understand. Naming it as the cause reverses the roles of the enabling condition and the defeated control. ## The human is not the control In a consumer chat product with no tools, nobody is standing between the message and the model comparing strings against a term list. If a person saw the message at all, they read it the way people read: for sense. They got sense, because the substitution preserves it. There is no lapse here to attribute — a reader who "noticed nothing odd" performed exactly the task reading is, and no amount of care at that task answers a question about codepoint identity, which is not a question the eye can be asked. So a finding that blames the reader has named an actor who was never in the decision path, which means it cannot be routed to an owner and cannot be closed by any change. ## What the root cause line should actually say Something close to: *a stage deciding string identity sits in front of a reader deciding meaning, and a substitution that changes identity while preserving meaning makes the two disagree about one input.* That single sentence carries everything a triager needs: - It names the stage that stayed silent (the comparison over stored terms) without claiming that stage malfunctioned. It did not; it answered its own question correctly. - It names the reader that acted (the model, which recovered the word from context) without claiming the model malfunctioned either. - It locates the defect in the *relationship* between the two rather than in either one. ## Bug or design limit — writing the verdict This is the call somebody has to make, and the honest version is uncomfortable: the divergence is not an oversight, it is what these two mechanisms are. A comparison over surface strings decides membership in a finite list of exact sequences. A model decides what an input plausibly means. There is no threshold, no list length and no amount of care that makes those two decisions coincide for all inputs, because they are not answers to the same question. That does not make the report worthless — quite the opposite. It makes it a report about a **limit** rather than a **defect**, and a limit is filed differently: it constrains what the deployment can be claimed to prevent, rather than describing something that was supposed to work and did not. Where a specific stored term matters enough, the follow-up is a question about which relation anything in front of the model is comparing — a question this finding raises and does not itself answer. ## The claims to keep straight when writing it up - The screen staying silent proves the input equalled no stored term under that comparison. It proves nothing about harm and nothing about the model. - The model answering proves the model produced that text on that turn. It does not prove the model would do so again — generation is sampled, and a construction that leans on the model recovering meaning has a rate. - The reader seeing nothing odd proves the substitution preserved appearance. That is a statement about the construction working as designed, not about the reader. - Getting content that would otherwise be declined is the entire payoff in this shape of product. There is no tool to reach, no store to write, no second reader downstream. Saying so plainly keeps the finding's impact honest rather than inflated by association with agentic outcomes it cannot produce. ## What an interviewer is scoring Not whether you can name the technique — everyone can. Whether you can separate the enabling property from the defeated control, decline to blame a human who was not in the loop, and then state the verdict in a form somebody can act on: which relation each mechanism compares, and what follows about the claims the deployment can support.

  • The reporter insists a more careful reviewer would have caught it. How do you answer?
    By pointing out what care would have to consist of: comparing codepoint identity by eye, which is not something reading does at any level of attention. The substituted glyph is the expected glyph. Attributing the outcome to attention sets an owner an impossible task and hides the mechanism that actually decided the message went through.
  • Does calling it a design limit mean the report should be closed with no action?
    No. A limit constrains what the deployment can be claimed to prevent, which is itself an output somebody needs. It changes the shape of the follow-up from "restore intended behaviour" to "decide what this stage is relied on for", and it stops the finding being routed to a person whose only available fix is to read harder.
  • What impact claim can this finding honestly carry in a product with no tools?
    Content the model would otherwise have declined, delivered as text on one user's screen, at some observed rate. There is no tool call, no data belonging to another user, no write that outlives the session. Borrowing severity language from agentic findings would be the fastest way to lose the reader who has to act on it.

Blaming the reader here is like closing a mismatched-serial-number report by noting that the cashier thought the note looked fine. Looking fine was never the check that failed.

saying these in an interview costs you the question

  • Blames a human reviewer who was never comparing strings
  • Treats the identical rendering as the control that failed
  • Says the term screen malfunctioned when it answered correctly
  • Reports one successful turn as a reliable result
  • Inflates impact with outcomes the product cannot produce

context