skip to content

A trace viewer shows a clean prompt - why is that not evidence no invisible span reached the model?

level: seniorimportance: must knowfreq 55%

answer

  1. the log viewer is a renderer too
  2. both worlds produce the same picture
  3. an instrument that cannot separate hypotheses
  4. displayed glyphs against stored codepoints
  5. read the file, not the pretty print

basics

~20 s

A trace viewer is a renderer too: it drops exactly the codepoints the carrier was chosen for. A clean display proves the viewer drew nothing, not that the prompt was clean. Only a codepoint-level check over stored values answers it.

solid answer

~50 s

The wrong answer here is: we would have seen it in the logs. The platform team's trace viewer sits at the same disadvantage as the admin table and the analyst's notebook - it turns a codepoint stream into a glyph stream, and non-rendering codepoints contribute no glyphs. An attacker picking a carrier is selecting against every human-facing surface, and a log UI is one of them. Get the direction of the claim right: the viewer showing nothing establishes that the viewer displayed nothing, which is what it would do in both worlds. What can actually distinguish them is a machine check over the stored value - count codepoints, bucket them by Unicode general category and block, compare that count with the number of glyphs shown. If the raw log file escapes non-ASCII, reading the file rather than the UI is a second way to see bytes.

code

json · 11 lines
json
{
  "field": "account_name",
  "as_drawn_in_trace_viewer": "Northwind Retail",
  "glyphs_displayed": 16,
  "stored_codepoint_count": 63,
  "non_rendering_by_bucket": {
    "Cf format characters (e.g. U+200B, U+200D)": 12,
    "Tags block U+E0000-U+E007F": 35
  },
  "note": "[directive span elided - carried entirely in the non-rendering codepoints]"
}

go deeper

for a junior

Remember the one-liner: a log viewer displays text, so it drops the same characters every other display drops. Not seeing something there is not the same as it not being there.

for a middle

Explain why both hypotheses produce the same screen, and name the measurement that separates them - counting stored codepoints and their categories rather than reading a rendering.

for a senior

Show triage discipline: state what the observation supports, name the discriminating check, and resist the affirmative error too, since a census hit is a lead and not yet a finding.

for a principal

Be ready to correct this claim when it is made by a team that has already told stakeholders their logs are clean, and to say plainly what that assurance was worth.

## The claim being made, and what it actually supports Somebody says: a hidden-codepoint span cannot be reaching our assistant's prompts, because we would see it in the trace viewer. Take that apart. The trace viewer takes a stored record, decodes it, and paints text on a screen. It is a renderer. Every argument that disqualifies the admin table where the warehouse row was first reviewed, and the notebook cell that prints the row for an analyst, applies to it unchanged. Worse, it is a renderer the attacker can *assume* is in the loop - a platform team that runs an LLM assistant has a trace UI - so it is among the surfaces the carrier was chosen against in the first place. So the observation and its conclusion do not connect: - **Observed:** the viewer drew no unusual glyphs. - **Supported:** the viewer drew glyphs for the codepoints it renders and nothing for those it does not. - **Not supported:** the codepoint sequence in the prompt contained only what was drawn. The two worlds - clean value, and value carrying a non-rendering span - produce an *identical* observation on this instrument. An instrument that cannot separate the hypotheses provides no evidence about them, however many people look at it. ## Why this specific wrong answer is so durable Because it is normally right. Logs are the profession's default answer to 'did X happen', and for almost everything else they are the correct answer. The failure here is narrow and non-obvious: the property being investigated is *the property the display layer is specified to erase*. Nobody has to be careless for the check to fail. The most attentive engineer on the team, reading the trace line by line, gets exactly the same nothing. A second reason it persists is that the search feels thorough. Searching the viewer for suspicious text finds nothing, and 'I searched and found nothing' is psychologically much stronger than 'I looked at a summary that omits the category of thing I am looking for'. ## What does discriminate Only operations that read the stored representation rather than a rendering of it: - **A codepoint census.** For the field value or the stored prompt: total codepoints, bucketed by Unicode general category and block. Format characters (`Cf`) and the Tags block `U+E0000`-`U+E007F` are the buckets that matter, and the tell is arithmetic - a value that displays sixteen glyphs and stores sixty-three codepoints. - **Displayed length against stored length.** Cheap, and it needs no judgement about which codepoints are suspicious. - **Reading the raw log file rather than the UI.** If the pipeline writes JSON with non-ASCII escaped, the bytes on disk spell out the escapes even though the pretty-printed view does not. That is not a different level of diligence; it is a different instrument. Notice what is *not* on the list: opening the record in a second viewer, changing font or theme, asking the assistant to echo the field value back (its answer goes through a renderer too), or having a more careful person look again. All of those are the same instrument with a different operator. ## The trap in the affirmative direction as well The symmetric error is worth naming, because a triage call turns on it. If a census *does* find non-rendering codepoints in a field, that establishes the codepoints are there - not that anyone placed them deliberately, and not that the model acted on them. Format characters occur naturally in text that has been pasted through editors, spreadsheets and word processors. A finding is a codepoint census plus a reason to believe the span reads as a directive plus a path from the column to the prompt; a census alone is a lead. ## Saying it out loud in the interview The answer an interviewer is scoring for has three moves. Name the viewer as a renderer. Say what the observation supports and what it does not, without hedging. Then name the discriminating measurement and be concrete about it - counts and categories over stored values, not a more careful reading. A candidate who stops after 'logs can be misleading' has named a mood; a candidate who gets the direction of the claim right has named the mechanism.

  • What exactly would you run to answer the question the viewer cannot?
    A codepoint census over the stored value: total codepoints, grouped by Unicode general category and block, compared with the glyph count the surface showed. Format characters and the Tags block are the buckets to look at. It is a query over the column, not an inspection - which is the whole point, because it never routes through a display layer.
  • The row was reviewed by a human when it was created. Does that help?
    No, and it is the same error one hop earlier. That reviewer read the admin table, which is a renderer with the same blind spot. Two renderings by two people are not two pieces of evidence; they are one instrument used twice.
  • A census finds format characters in a field. Is that a finding?
    It is a lead. Format characters turn up naturally in text pasted through editors and spreadsheets. To call it a finding you want a reason the span reads as a directive and a demonstrated path from that column into the prompt. Reporting the census alone as an injection will not survive triage.
  • Would asking the assistant to repeat the field value back settle it?
    No. Its answer is text that lands in the same trace viewer or notebook cell, so the codepoints get dropped on the way to your eyes a second time. You would be re-running the failed instrument with an extra step, and the extra step can also alter the string.

saying these in an interview costs you the question

  • Treats a clean log view as proof the prompt was clean
  • Proposes looking again, more carefully, in the same viewer
  • Suggests a second viewer or a different font as the check
  • Asks the model to echo the value and trusts the displayed answer
  • Calls a codepoint census result an injection finding on its own

context