skip to content

Multimodal & Encoding-Based Injection

You will learn how instructions ride in on non-text channels and invisible characters that humans cannot see but the model reads, from Unicode tag smuggling to text embedded in images. Interviewers raise it because multimodal and copy-paste inputs are a fast-growing injection surface many defenders overlook.

on this pageshow

explore

questions

27

Why does text written in non-rendering codepoints still reach a model when no screen shows it?

level: juniorimportance: must knowfreq 68%

answer

  1. a screen is not the string
  2. glyphs are a view; codepoints are the input
  3. zero width means zero glyphs, not zero tokens
  4. the tokenizer never sees the rendering

basics

~20 s

A screen is a lossy view of stored text. Codepoints such as zero-width format characters and Unicode tag characters produce no glyph, but they are still stored, still copied, and still turned into tokens, so the model reads them.

solid answer

~50 s

Unicode defines codepoints whose specified behaviour is to draw nothing: the zero-width format characters (U+200B, U+200D and neighbours) and the Tags block U+E0000-U+E007F, which mirrors ASCII one codepoint for one and is default-ignorable. A renderer that shows nothing for them is behaving correctly, so the glyph stream a person reads is a *derived*, lossy view of the codepoint stream behind it. The model gets no such view. An internal data-analysis assistant whose query template pastes a free-text warehouse column into the prompt hands the model the stored string, and the tokenizer assigns tokens to those codepoints like any others. So an attacker who wrote a directive into a product form months earlier gets a value that the admin table renders as an ordinary account name and the model reads as an ordinary account name plus a span nobody has ever seen.

go deeper

for a junior

Be ready to say, in one sentence, that a screen shows glyphs while a model reads codepoints, and to name one class of character that produces no glyph.

for a middle

An interviewer expects the mechanics: which codepoints are defined as non-rendering, that a tokenizer assigns them tokens anyway, and that a text column and JSON transport preserve them untouched.

for a senior

Show that you reason about which surface produced a claim. A row reviewed in an admin table, printed in a notebook or shown in a trace viewer has been read through three renderers and zero byte-level checks.

for a principal

Own the consequence for assurance: if every human-facing surface is disqualified by construction, then any statement about a corpus has to rest on a machine check over stored values, and you should say so before somebody promises otherwise.

## The two strings Start with the setting, because the asymmetry only makes sense once you can see both readers. An internal data-analysis assistant answers questions over a product-analytics warehouse. It has no retrieval index and no public chat surface; a query template simply pastes free-text column values - account names, ticket subjects, feature-flag descriptions - into the prompt alongside the analyst's question. One of those values was written months ago through an ordinary product form by somebody who wanted the assistant to be told something. In every table anyone has opened it in, it reads as an ordinary account name. The bytes in the column are not that. ## Renderers drop things by specification, not by accident Unicode contains codepoints whose defined behaviour is to produce no glyph at all. - The zero-width **format characters** - zero-width space `U+200B`, zero-width joiner `U+200D`, zero-width non-joiner `U+200C` and their neighbours - carry General_Category `Cf`. They exist to influence joining, shaping and line breaking, not to be seen. - The **Tags block**, `U+E0000` to `U+E007F`, mirrors *printable* ASCII: its assigned tag characters `U+E0020`-`U+E007E` correspond one for one to `0x20`-`0x7E`, with `U+E007F` as CANCEL TAG, and the block is defined as default-ignorable: a conforming renderer that does not implement the (long-deprecated) language-tag mechanism is expected to display nothing. - Bidirectional controls such as `U+202E` are visible in their *effect* on ordering but draw no glyph of their own. So the glyph stream a person reads is a lossy function of the codepoint stream that produced it. That is not a rendering bug. It is what the specification asks for, and it is the reason no amount of care in one viewer changes the outcome. ## What the model receives The model does not receive a picture. It receives the string, and its tokenizer maps that string - all of it - to tokens. Format characters do not disappear at tokenization; they are codepoints like any others. Whether a given model then interprets a tag-block sequence as its ASCII counterpart is model-dependent and varies, but that question sits *downstream* of the fact that matters at this level: the characters are inside the context window, and the analyst's eyes never had access to them. A short map of who can see what: | Surface | Reads | Sees the span? | |---|---|---| | Admin table where the row was reviewed | rendered glyphs | no | | Notebook cell printing the row | rendered glyphs | no | | Prompt template | stored codepoints | passes them on | | Tokenizer and model | stored codepoints | yes | ## Getting the direction of the claim right This is where junior answers go wrong, and it is worth stating flatly. **Seeing nothing proves that a surface drew nothing for the codepoints it renders. It proves nothing about the bytes.** The review that happened when the row was created was a review of a rendering. The same is true of the analyst's notebook output, and - the part people resist - of the platform team's trace viewer, which is also a renderer. The converse error is just as common: assuming something must have removed the characters along the way. Nothing removes them by default. A text column stores them, JSON transport carries them, a copy-paste preserves them, and a form that accepts free text has no reason to reject a format character. Removal is a step somebody has to have deliberately added; absence of that step is the ordinary case. ## Why this is a different problem from a look-alike character A look-alike codepoint renders as *something* - a glyph that resembles another one - so a reader has, in principle, something to notice. A non-rendering codepoint renders as *nothing*, so there is nothing to notice by looking, at any level of attention. That difference is why 'check it carefully' is not a strategy against this class: the checking faculty being appealed to is the one the carrier was selected to defeat. ## What a good answer sounds like Name the mechanism (codepoints defined to draw nothing), name the asymmetry (renderer sees glyphs, tokenizer sees codepoints), and name one concrete surface that has already misled somebody - the table where the row was reviewed. Then say the honest consequence: the only thing that can answer 'was there a hidden span in this value' is a check that looks at the stored codepoints, because every surface built for humans is disqualified by construction.

  • How can a span made only of invisible codepoints carry anything readable?
    Two properties do it. The Tags block mirrors ASCII one codepoint for one, so a sequence there has a straightforward reading while drawing nothing on screen. Separately, zero-width format characters can be interleaved between ordinary letters, which leaves the words legible to a model reading tokens while breaking a byte-level string comparison. Neither needs a clever encoding; the carrier is doing the work, not the message.
  • If a model ignores the hidden codepoints, is the problem gone?
    Not really, and it is the wrong thing to lean on. Behaviour varies by model, tokenizer and version, so an observation on one deployment is not a property of the class. More to the point, the characters are still in the stored value, still copied into every export and every answer that quotes the field, and still present the next time a different component reads that row.
  • Does the value passing a form's validation tell you anything?
    Only that it satisfied whatever the form checked, which for a free-text field is usually a length limit and maybe a control-character rule. Format characters are legitimate text; there is no default reason for a product form to reject them, and no default reason for the column to store anything other than what it received.

A web page shows one space where the file has forty. The page is not lying, it is summarising - and the model reads the file, not the page.

saying these in an interview costs you the question

  • Assumes the model can only read what a person can see
  • Believes invisible characters get stripped somewhere by default
  • Treats the admin table review as evidence about the stored bytes
  • Confuses characters that draw nothing with look-alike glyphs
  • Thinks a form accepting the value implies the value is printable

context

open as a page

Why does an exact-match term screen miss a blocked word when a user swaps in a look-alike codepoint?

level: juniorimportance: must knowfreq 64%

basics

~20 s

The screen compares codepoints, not appearance. A Cyrillic letter drawn like a Latin one is a different codepoint, so the word no longer equals any stored term and the comparison returns no match, while the glyphs on screen are unchanged.

open as a page

A vendor invoice PDF passed human review — how can extracted text still carry a directive the reviewer never saw?

level: juniorimportance: must knowfreq 68%

basics

~20 s

The reviewer read the document as rendered; extraction deliberately recovers what rendering omits — form-field values, annotation contents, off-canvas or notice-sized runs. Those are two different texts, and only the extracted one reaches the model.

open as a page

Why doesn't a scan of an uploaded call recording catch an injection that appears in its transcript?

level: juniorimportance: must knowfreq 68%

basics

~20 s

The scan inspects audio bytes, and at that moment the instruction text does not exist. Transcription manufactures it afterwards. A clean verdict on the stored file says nothing about the transcript the model later reads.

open as a page

In an assistant that reads pasted screenshots, why is text inside the image untrusted input?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Pasting an image vouches for why the user wanted it, not for who wrote the words in it. A multimodal model reads every legible sentence in the frame, so a planted directive reaches the context as content the model may follow.

open as a page

Why is 'we normalise input, so hidden characters are gone' an incomplete claim against a carrier-choosing attacker?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A normalising pass removes a specific, enumerable set of carriers, not everything a span could hide in. 'We normalise' names that a pass runs; it does not say which carriers it strips and which it keeps. The attacker just picks a carrier from what survives.

open as a page

Why does an input screening model score an encoded span as benign when the generator obeys it?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Both components get the same characters, but only one resolves them. The screening model scores the encoded surface; a capable generator reconstructs the familiar encoding unasked and acts on the reconstructed meaning. No pipeline stage decoded anything in between.

open as a page

A trace viewer shows a clean prompt - why is that not evidence no invisible span reached the model?

level: seniorimportance: must knowfreq 55%

basics

~20 s

A trace viewer is a renderer too: it drops exactly the codepoints the carrier was chosen for. A clean display proves the viewer drew nothing, not that the prompt was clean. Only a codepoint-level check over stored values answers it.

open as a page

What must a non-rendering codepoint survive to work as a hidden carrier in stored free text?

level: middleimportance: should knowfreq 44%

basics

~10 s

A usable carrier clears two bars at once: it must round-trip through the form, the column and the prompt template unchanged, and draw nothing on every human-facing surface, not just the one someone checked.

open as a page

Why can a model still read the intended word when a look-alike codepoint has changed its tokens?

level: middleimportance: should knowfreq 48%

basics

~20 s

A model reads tokens, not characters, and predicts from context rather than looking strings up. The substitution splits the word into rarer fragments, but the surrounding words constrain what it must mean, so the model usually reads through it.

open as a page

In a PDF extraction pipeline, what property makes a container field eligible to carry an unseen directive span?

level: middleimportance: should knowfreq 46%

basics

~20 s

Eligibility is the asymmetry between two deliberate behaviours: extraction recovers everything machine-readable, rendering displays only what the layout puts under an eye. Any field emitted by the first and not shown by the second is a carrier.

open as a page

In a transcription pipeline, how does the recognition step's own error behaviour become part of the attack surface?

level: middleimportance: should knowfreq 41%

basics

~20 s

A transcript is a recogniser's guess, not a copy. Its mistakes can add text nobody spoke, and its normalisation rewrites what was said, so the span the model reads is shaped by the stage as much as the speaker.

open as a page

Why does directive text placed in a pasted screenshot survive the pasting user's glance?

level: middleimportance: should knowfreq 52%

basics

~20 s

Because salience and legibility are separate thresholds. The user looks at the image for the one thing they pasted it for, and their attention stops there, while the model reads everything in the frame that is resolvable at all.

open as a page

How does an attacker choose a carrier against a chain of lossy passes between a fetched page and the model?

level: middleimportance: should knowfreq 55%

basics

~20 s

Each pass strips some carriers and keeps others, so the attacker picks a carrier that survives the intersection of every pass on its path, not any single one. A carrier a screenshot flattens is useless even if the normaliser would have kept it. The choice is made against the whole chain.

open as a page

An LLM pipeline decodes inputs before screening them — why does an encoded span still get through?

level: middleimportance: should knowfreq 55%

basics

~20 s

Decoding covers the representations somebody listed, at the stages where the step runs. A generator reconstructs forms nobody enumerated, including malformed ones a strict decoder rejects, so the step narrows the unexpanded surface rather than closing the class.

open as a page

A look-alike substitution report says "a person read it and it looked normal" — what did it get wrong?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It named the wrong cause. Looking normal is the property the substitution buys, not the control it defeated. The stage that was defeated compares identity, and nobody was ever asked to compare that string against a list.

open as a page

A process owner proposes telling the extraction model to ignore hidden text — what does that actually close?

level: seniorimportance: should knowfreq 54%

basics

~10 s

Almost nothing deterministic. Extracted text arrives with no record of what the page displayed, so hidden is a fact about a viewer, not something the model observes.

open as a page

A directive found in one call transcript won't reappear on re-transcription — is that a finding?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Yes, if the exposure is decidable from what was produced. A failed replay produces a new transcript; it does not overwrite the one that actually reached the model. Separate whether it happened from how often it recurs.

open as a page

A pasted screenshot's instruction is obeyed, but the assistant's draft still needs a confirm click — what does that constrain about which argument values survive?

level: seniorimportance: should knowfreq 41%

basics

~20 s

The click reviews a rendered surface, not the call. Values the draft shows have to look like what the user asked for, so a value the user never chose survives only where the rendering rolls it up, truncates it or omits it.

open as a page

An input screen and an output screen both logged allow for a turn in which the assistant obeyed an encoded span — what does that record prove?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Two allow verdicts prove that two surface forms scored below threshold on the labels those screens carry. They say nothing about what the generator reconstructed, which call it made, or which argument values it chose.

open as a page

A carrier finding reproduced for the reporter, but their bug tracker stripped the carrier from the ticket: is it a finding?

level: principalimportance: should knowfreq 30%

basics

~20 s

Potentially yes: the bug tracker is itself another lossy pass, so a missing carrier in the ticket is expected, not disproof. Triage has to recover the carrier out-of-band and reproduce end-to-end. The destroyed evidence is a property of the reporting channel, not a verdict on the claim.

open as a page

An edit-distance matcher sits behind an exact term screen — what does that cost a look-alike substitution?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

It puts a floor on how many characters must change: one substitution is distance one from the stored term, inside any fuzzy window. Clearing the threshold means substituting more, and that price is paid entirely at the model.

open as a page

An approved invoice record's extracted fields differ from the PDF page — what does the approval log prove?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

It proves a decision was recorded against a record by an actor at a time. It does not prove which artefact was on screen; in this flow the approver usually reads the extracted summary, not the page.

open as a page

A carrier must survive weekly re-fetches and internal re-copies, not one run: how does that change the attacker's pick?

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

When the prize is repeat delivery, the attacker optimises for a carrier that survives every re-fetch, re-scrape and downstream copy of the republished output, not one pipeline run. That favours carriers riding in stable, re-read fields and rules out any that a single manual touch or export step in the recurring path would flatten.

open as a page

An obfuscation only the answering model resolves, not the screen: what is that method worth next quarter?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

The class outlives any single form. Generators keep resolving more representations unasked while screening models stay smaller, so the asymmetry widens; a decode step retires one enumerated form, and only a change in the generator's own behaviour touches the class.

open as a page

Asked which of 40M warehouse rows ever carried an invisible-codepoint span, what can you honestly claim?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Claim only what an instrument supports: a codepoint census shows which rows carry such a span now, across the whole population. It cannot show what overwritten values held, or where copies already travelled. Say where the boundary is.

open as a page

An owner funds a stronger upload scanner after a transcript-borne injection — what do you tell them?

level: principalimportance: nice to knowfreq 27%

basics

~10 s

The spend buys real things and buys nothing against this class, because the scanner inspects an artefact that never contained the span. The harder part is saying so before it is booked as remediation.

open as a page