How does an attacker choose a carrier against a chain of lossy passes between a fetched page and the model?
answer
- an ordered sequence of lossy passes
- each pass has a kill-set
- carrier must survive the intersection
- prefer a field a stage routes around
basics
~20 sEach pass strips some carriers and keeps others, so the attacker picks a carrier that survives the intersection of every pass on its path, not any single one. A carrier a screenshot flattens is useless even if the normaliser would have kept it. The choice is made against the whole chain.
solid answer
~50 sModel the pipeline as an ordered sequence of lossy transforms, each with its own enumerable kill-set. A scraper's HTML-to-text conversion drops markup and some attributes but keeps body text; a screenshot or re-encode flattens characters to pixels, destroying codepoint-level carriers such as zero-width runs and homoglyph distinctions while keeping the visible glyphs; a normalise pass removes its defined codepoint set. The attacker's carrier has to land in the *surviving* set of every stage on its path, which is an intersection, not a single filter. So the reasoning is: enumerate each stage's loss, find a carrier no stage removes, and prefer one in a field that some stage routes around entirely, such as a metadata field the scraper reads but the screenshot step never covers. The winning carrier is chosen against the chain, not against whichever stage the team happens to name as coverage.
code
json · 11 lines{
"carriers": [
{ "where": "body_caption", "content": "[zero-width span elided]",
"after_html_to_text": "kept", "after_screenshot": "dropped", "after_normalise": "dropped" },
{ "where": "body_caption", "content": "[homoglyph span elided]",
"after_html_to_text": "kept", "after_screenshot": "flattened_to_glyph", "after_normalise": "kept (not in table)" },
{ "where": "title_metadata", "content": "[directive span elided]",
"after_html_to_text": "routed_around", "after_screenshot": "not_covered", "after_normalise": "never_received" }
],
"note": "survivor to the model = outside every kill-set on its path"
}go deeper
Know that several passes sit between the page and the model and each can drop content. Recall that a carrier must get past all of them, not just one.
Explain the intersection: enumerate each pass's kill-set and reason about which carriers survive every pass on the path, including ones routed around a stage.
Diagnose a real pipeline by mapping its ordered passes and predicting which carrier class an attacker would pick, then say which stage change would actually shrink the survivor set.
Own the model that coverage is a property of the whole path, not of the stage a team names, and drive assurance toward enumerating survivors across the chain.
## The pipeline is a sequence of lossy transforms Between an attacker's authored page and the model that reads it, several passes run in order, and **each one is lossy in its own way.** For a research agent that fetches public pages and republishes an internal brief, a realistic chain is: 1. **Scraper / HTML-to-text conversion** — drops tags and much markup, keeps body text; may or may not carry an alt attribute or a title into the extracted text depending on how it is written. 2. **An intermediate export or paste** — moving the extracted text through a store or a document format, which can re-encode or drop things again. 3. **A screenshot or re-encode step** — if any stage rasterises text, characters become pixels: codepoint identity is gone, so zero-width runs and homoglyph-versus-Latin distinctions vanish, while the visible shapes remain. 4. **A normalise pass** — removes its defined codepoint set (a Unicode form, a control/zero-width list, a confusables table). ## Why the choice is an intersection, not a filter Each pass has an **enumerable kill-set**: the things it removes. A carrier survives a pass if and only if it is outside that pass's kill-set. To reach the model, the carrier must be outside the kill-set of **every** pass on its path. That is a set intersection. This is why reasoning about a single stage is a trap. A carrier can be perfectly invisible to the normaliser and still die at the screenshot step: - **Zero-width codepoints** survive a naive HTML-to-text scrape (they are text) but a rasterising step erases them, because pixels have no notion of an invisible codepoint, and a normalise pass that lists them also removes them. They only win on a path with no rasterise and no normaliser that targets them. - **Homoglyphs** survive the scraper and, if the confusables table misses them, the normaliser too, but a screenshot keeps only the shape, and if OCR later re-creates text the distinction may or may not come back. Their survival depends on the exact chain. - **A directive in an image** dies in any text normaliser (which never sees pixels as text) but sails through it, and survives a screenshot step that keeps images, and only becomes model-visible if a downstream stage reads the image. - **A metadata or title field** may be read into the prompt by one stage while a screenshot of the rendered body never covers it at all; it is *routed around* several passes. ## The routing-around move The strongest carriers are often ones that some stage does not merely fail to remove but **never receives.** A title or metadata field pulled straight into the prompt bypasses a body-text normaliser entirely. This is different from surviving a pass; it is not being subject to it. An attacker who has mapped the chain prefers these because they remove a stage from the intersection instead of having to survive it. ## The selection procedure, described Without writing any payload, the reasoning an attacker follows is: 1. Determine the ordered passes on the path from the page to the model. 2. For each pass, note its kill-set: what it removes and what it routes around. 3. Intersect the survivors: which carriers are outside every kill-set on the path. 4. Among those, prefer a carrier a stage never receives (routed around) over one that merely survives, because it depends on fewer stages behaving as assumed. The point the interviewer is checking is that **the attacker chooses against the chain**, so pointing at one normalise step and calling it coverage measures the wrong thing. A candidate who can enumerate the passes and reason about the intersection understands why single-stage assurances mislead; one who evaluates only the named step has missed the mechanism.
- Why is a carrier that a stage 'routes around' stronger than one that merely survives it?Surviving a pass depends on that pass behaving as the attacker assumed. A carrier in a field the pass never receives, such as a title read straight into the prompt around a body-text normaliser, removes that stage from the intersection entirely, so its survival hangs on fewer assumptions and fewer future changes to that stage.
- Why can adding a screenshot step to a pipeline both help and hurt against carriers?Rasterising text erases codepoint-level carriers like zero-width runs and homoglyph distinctions, which helps. But it preserves anything visible, and if a later stage OCRs or reads images it re-creates text and admits carriers baked into pixels. It changes the kill-set, closing some carriers and opening others; it is not a net removal.
saying these in an interview costs you the question
- Evaluates only the one stage the team calls coverage.
- Assumes a carrier that beats the normaliser therefore reaches the model.
- Thinks a screenshot step removes all carriers rather than changing which survive.
- Ignores fields that are routed around a pass entirely.