What must a non-rendering codepoint survive to work as a hidden carrier in stored free text?
answer
- two bars, not one
- arrive unchanged, and draw nothing anywhere
- invisibility is renderer-relative
- one viewer with a placeholder box burns it
- count codepoints, not glyphs
basics
~10 sA usable carrier clears two bars at once: it must round-trip through the form, the column and the prompt template unchanged, and draw nothing on every human-facing surface, not just the one someone checked.
solid answer
~50 sThe requirement is a conjunction, and most non-rendering codepoints fail one half of it. Survival first: the value has to pass whatever the product form accepts, sit in a text column, come back through JSON transport and get pasted into the prompt template with the same codepoints it started with. Invisibility second, and this is the strict part - it has to draw nothing in *every* surface a human might read it through: the admin table where the row was first reviewed, the platform team's trace viewer, and the notebook cell an analyst prints the row into. One of those showing a dotted box or a replacement glyph burns the carrier, because now there is a surface where a person can notice. That conjunction is why the usable set is small and why a candidate who can only recite families of invisible characters has not answered the question.
go deeper
Know that a hidden span has to both survive being stored and stay unseen when displayed, and that these are two separate things that can each fail.
Be able to walk the path - form, text column, transport, prompt template - and say where a codepoint can be rejected, transcoded or preserved, and why a placeholder glyph in any one viewer defeats the whole point.
Show you enumerate surfaces rather than sampling one. In a real loop there are usually three or four renderers with different behaviour, and the weakest assumption is that they all behave like the one you opened.
The takeaway you should be able to state to a team: a carrier's invisibility is a claim about a set of tools you happen to run today, so it degrades silently in both directions as those tools change.
## Two requirements, and they pull in different directions Asking which codepoints are invisible is a trivia question. The interesting question is which ones are *usable*, and that is a conjunction of two independent properties. A carrier fails if it loses either. ### 1. It has to arrive unchanged The path in the data-analysis setting is short and completely ordinary: somebody types a free-text value into a product form, it lands in a text column of the warehouse, and months later a query template reads that column and pastes the value into the assistant's prompt. Each hop is a place a codepoint can die: - **The form.** Free-text fields rarely reject format characters, but a field with a strict allow-list, or one that trims and collapses aggressively, can drop them. A carrier that never gets stored is not a carrier. - **The column and its encoding.** A UTF-8 text column stores what it was given. A column with a narrower encoding, or a pipeline that lossily transcodes, can turn the span into replacement characters - which are *visible*, and therefore worse than useless. - **Transport and templating.** JSON escapes non-ASCII but does not delete it; string concatenation into a prompt preserves it. These hops are usually faithful, which is exactly why the class works. The practical test an attacker applies here is round-trip fidelity: write the value, read it back at the far end, compare codepoint sequences. Not glyphs - codepoints. ### 2. It has to draw nothing on *every* surface This is the half people underestimate, and it is the half this leaf exists for. Invisibility is not a property of a codepoint alone; it is a property of a codepoint *and a renderer*. The same character can be: - absent in a proportional-font web table, - drawn as a dotted box or a hex placeholder in a developer-oriented monospace view, - expanded into an escape sequence (`\u200b`) in a raw log file that nobody pretty-prints, - reordered-looking in one control's presence and unremarkable in another's. So the carrier is chosen against the *set* of surfaces in the loop, not against one of them. In this setting there are at least three: the admin table where the value was reviewed when it was created, the trace or log viewer the platform team reads, and the notebook cell that prints the row for an analyst. A codepoint that vanishes in two of those and shows a placeholder in the third is a codepoint that will eventually be noticed by someone doing nothing special. ### Why the two requirements conflict The characters that are most reliably preserved end-to-end tend to be the ones with the most defined semantics, and defined semantics is exactly what tempts a tool to show them. The characters that are most reliably ignored tend to be the ones some layer feels entitled to strip. The usable middle is narrow, which is a much more informative thing to say in an interview than a list of blocks. ### What this looks like from the far side One underrated consequence: because invisibility is renderer-relative, the *first* surface to betray the span is usually a developer tool, not a user-facing one - and it betrays it by accident, on a day someone happened to open the raw record for an unrelated reason. That is not a control anybody designed. It is luck, and it is a poor thing to plan an assurance story around. ### The measurement, not the recipe Worth being precise about what a person on either side of this actually does: the useful operation is a codepoint census. Take the stored value, count codepoints, bucket them by Unicode general category and block, and compare that count with the number of glyphs the surface drew. That comparison is the only thing in this whole path that can distinguish a clean value from a carrying one, and it is available to anybody who queries the column rather than looking at a table of it.
- Why is a codepoint that renders as a dotted box in one developer view unusable?Because the property being relied on is not 'invisible in the main UI' but 'invisible everywhere a person might look'. A placeholder box in a monospace record view means the span is one ordinary debugging session away from being seen, and the attacker gets no warning when that happens. The carrier has to hold against the whole set of surfaces in the loop.
- Does a value passing round-trip fidelity tell you it will still be intact a year later?No. It tells you the path behaved that way once, for that value, at that time. Encodings, export jobs and downstream consumers change without anybody re-testing text fidelity, so a stored carrier can quietly degrade into replacement characters - which are visible, and which then look like data corruption rather than anything else.
- How does this differ from choosing a carrier that survives a chain of transforms?Different question with a different owner. Here the path is short and mostly faithful - form, column, template - and the hard requirement is the every-renderer half. Carrier selection against successive rewriting stages, and what a single normalising pass really covers, is a separate problem in its own right.
saying these in an interview costs you the question
- Lists blocks of invisible characters instead of naming the requirement
- Treats invisibility as a property of the codepoint alone
- Checks one viewer and calls the carrier confirmed
- Ignores whether the value survives storage and transport intact
- Compares rendered glyph counts instead of stored codepoints