Why does text written in non-rendering codepoints still reach a model when no screen shows it?
answer
- a screen is not the string
- glyphs are a view; codepoints are the input
- zero width means zero glyphs, not zero tokens
- the tokenizer never sees the rendering
basics
~20 sA screen is a lossy view of stored text. Codepoints such as zero-width format characters and Unicode tag characters produce no glyph, but they are still stored, still copied, and still turned into tokens, so the model reads them.
solid answer
~50 sUnicode defines codepoints whose specified behaviour is to draw nothing: the zero-width format characters (U+200B, U+200D and neighbours) and the Tags block U+E0000-U+E007F, which mirrors ASCII one codepoint for one and is default-ignorable. A renderer that shows nothing for them is behaving correctly, so the glyph stream a person reads is a *derived*, lossy view of the codepoint stream behind it. The model gets no such view. An internal data-analysis assistant whose query template pastes a free-text warehouse column into the prompt hands the model the stored string, and the tokenizer assigns tokens to those codepoints like any others. So an attacker who wrote a directive into a product form months earlier gets a value that the admin table renders as an ordinary account name and the model reads as an ordinary account name plus a span nobody has ever seen.
go deeper
Be ready to say, in one sentence, that a screen shows glyphs while a model reads codepoints, and to name one class of character that produces no glyph.
An interviewer expects the mechanics: which codepoints are defined as non-rendering, that a tokenizer assigns them tokens anyway, and that a text column and JSON transport preserve them untouched.
Show that you reason about which surface produced a claim. A row reviewed in an admin table, printed in a notebook or shown in a trace viewer has been read through three renderers and zero byte-level checks.
Own the consequence for assurance: if every human-facing surface is disqualified by construction, then any statement about a corpus has to rest on a machine check over stored values, and you should say so before somebody promises otherwise.
## The two strings Start with the setting, because the asymmetry only makes sense once you can see both readers. An internal data-analysis assistant answers questions over a product-analytics warehouse. It has no retrieval index and no public chat surface; a query template simply pastes free-text column values - account names, ticket subjects, feature-flag descriptions - into the prompt alongside the analyst's question. One of those values was written months ago through an ordinary product form by somebody who wanted the assistant to be told something. In every table anyone has opened it in, it reads as an ordinary account name. The bytes in the column are not that. ## Renderers drop things by specification, not by accident Unicode contains codepoints whose defined behaviour is to produce no glyph at all. - The zero-width **format characters** - zero-width space `U+200B`, zero-width joiner `U+200D`, zero-width non-joiner `U+200C` and their neighbours - carry General_Category `Cf`. They exist to influence joining, shaping and line breaking, not to be seen. - The **Tags block**, `U+E0000` to `U+E007F`, mirrors *printable* ASCII: its assigned tag characters `U+E0020`-`U+E007E` correspond one for one to `0x20`-`0x7E`, with `U+E007F` as CANCEL TAG, and the block is defined as default-ignorable: a conforming renderer that does not implement the (long-deprecated) language-tag mechanism is expected to display nothing. - Bidirectional controls such as `U+202E` are visible in their *effect* on ordering but draw no glyph of their own. So the glyph stream a person reads is a lossy function of the codepoint stream that produced it. That is not a rendering bug. It is what the specification asks for, and it is the reason no amount of care in one viewer changes the outcome. ## What the model receives The model does not receive a picture. It receives the string, and its tokenizer maps that string - all of it - to tokens. Format characters do not disappear at tokenization; they are codepoints like any others. Whether a given model then interprets a tag-block sequence as its ASCII counterpart is model-dependent and varies, but that question sits *downstream* of the fact that matters at this level: the characters are inside the context window, and the analyst's eyes never had access to them. A short map of who can see what: | Surface | Reads | Sees the span? | |---|---|---| | Admin table where the row was reviewed | rendered glyphs | no | | Notebook cell printing the row | rendered glyphs | no | | Prompt template | stored codepoints | passes them on | | Tokenizer and model | stored codepoints | yes | ## Getting the direction of the claim right This is where junior answers go wrong, and it is worth stating flatly. **Seeing nothing proves that a surface drew nothing for the codepoints it renders. It proves nothing about the bytes.** The review that happened when the row was created was a review of a rendering. The same is true of the analyst's notebook output, and - the part people resist - of the platform team's trace viewer, which is also a renderer. The converse error is just as common: assuming something must have removed the characters along the way. Nothing removes them by default. A text column stores them, JSON transport carries them, a copy-paste preserves them, and a form that accepts free text has no reason to reject a format character. Removal is a step somebody has to have deliberately added; absence of that step is the ordinary case. ## Why this is a different problem from a look-alike character A look-alike codepoint renders as *something* - a glyph that resembles another one - so a reader has, in principle, something to notice. A non-rendering codepoint renders as *nothing*, so there is nothing to notice by looking, at any level of attention. That difference is why 'check it carefully' is not a strategy against this class: the checking faculty being appealed to is the one the carrier was selected to defeat. ## What a good answer sounds like Name the mechanism (codepoints defined to draw nothing), name the asymmetry (renderer sees glyphs, tokenizer sees codepoints), and name one concrete surface that has already misled somebody - the table where the row was reviewed. Then say the honest consequence: the only thing that can answer 'was there a hidden span in this value' is a check that looks at the stored codepoints, because every surface built for humans is disqualified by construction.
- How can a span made only of invisible codepoints carry anything readable?Two properties do it. The Tags block mirrors ASCII one codepoint for one, so a sequence there has a straightforward reading while drawing nothing on screen. Separately, zero-width format characters can be interleaved between ordinary letters, which leaves the words legible to a model reading tokens while breaking a byte-level string comparison. Neither needs a clever encoding; the carrier is doing the work, not the message.
- If a model ignores the hidden codepoints, is the problem gone?Not really, and it is the wrong thing to lean on. Behaviour varies by model, tokenizer and version, so an observation on one deployment is not a property of the class. More to the point, the characters are still in the stored value, still copied into every export and every answer that quotes the field, and still present the next time a different component reads that row.
- Does the value passing a form's validation tell you anything?Only that it satisfied whatever the form checked, which for a free-text field is usually a length limit and maybe a control-character rule. Format characters are legitimate text; there is no default reason for a product form to reject them, and no default reason for the column to store anything other than what it received.
A web page shows one space where the file has forty. The page is not lying, it is summarising - and the model reads the file, not the page.
saying these in an interview costs you the question
- Assumes the model can only read what a person can see
- Believes invisible characters get stripped somewhere by default
- Treats the admin table review as evidence about the stored bytes
- Confuses characters that draw nothing with look-alike glyphs
- Thinks a form accepting the value implies the value is printable