skip to content

Which properties must a span planted on a fetched page have to survive parsing, sanitising and chunking?

level: middleimportance: should knowfreq 44%

answer

  1. what does each stage keep?
  2. boilerplate removal deletes whole regions
  3. chunk boundaries are not yours to choose
  4. invariant beats tuned when you cannot observe
  5. visible to the converter, visible to a reader

basics

~20 s

It must be plain prose, sit in the region a boilerplate remover keeps as main content, depend on no structure or position, and stay whole inside one chunk — properties invariant across a chain nobody outside can observe.

solid answer

~50 s

The attacker's real problem is that the chain from a published page to a context window is unobservable from outside — they cannot see which parser, which boilerplate heuristic, which converter or which chunk size the agent uses. So the span has to be written to be invariant rather than tuned. Four properties do that. It must be ordinary prose, because every stage in the chain is defined as discard markup, keep text. It must sit inside the main content region, because boilerplate removal deletes navigation, footers and sidebars wholesale — placement, not wording, is what kills a span there. It must not depend on structure, position or adjacency, since the converter flattens all of that. And it must be self-contained in a short window, because chunk boundaries are chosen at run time by the pipeline, and half a directive span is just an odd sentence.

go deeper

for a junior

Know that the stages between a web page and a model keep readable text and throw away presentation, and that placement on the page matters because whole regions get discarded before anything is read.

for a middle

Be ready to name each stage and the property it imposes: prose for the converter, main-content placement for the boilerplate remover, self-sufficiency for the flattening, brevity for the split. Explain why an unobservable chain rewards invariance over cleverness.

for a senior

Demonstrate the diagnosis in reverse: given a span that worked once and failed elsewhere, reason from region placement and truncation before reaching for detection. Be explicit that survival is a necessary condition and not evidence of effect.

for a principal

Own the trade being made — reliability against concealment — and be able to say what that implies for how long such a construction stays useful and what would actually retire it, without turning the answer into a remediation plan.

## The problem is prediction, not wording Somebody planting a span on a public page for a research agent to read has a design constraint that shapes everything: they cannot observe the transformation chain. They do not know which HTML parser the fetcher uses, which boilerplate heuristic decides what counts as the article, whether the converter emits plain text or markdown, what the truncation budget is, or where chunk boundaries will fall. They get no error messages and no echo of the intermediate stages. Everything must therefore be *invariant* across plausible chains rather than tuned to one. That single constraint explains why the surviving constructions in this family look boring. Cleverness that depends on a specific stage behaving a specific way is a bet placed blind; plain text that every stage is contractually obliged to preserve is not a bet at all. ## Property one: it is prose, not structure Every stage between the page and the context is defined in the same direction — throw away presentation, keep the readable words. A DOM parse re-represents, a sanitiser removes executable markup, a converter drops the remaining tags. Text nodes are the fixed point of all of them. A construction that instead relies on a comment, an attribute value, an element the converter might or might not keep, or a particular encoding behaviour of one stage is a construction whose survival is unknown to its author. Those carrier-level techniques are a different family with a different cost profile; the point here is that they trade reliability for concealment, and the blind chain punishes exactly that trade. ## Property two: placement inside the kept region This is the property people miss, because it is not about the text at all. A boilerplate remover does not edit sentences; it deletes whole regions. Navigation, headers, footers, sidebars, comment blocks and related-links widgets are dropped as low-content furniture, and everything inside them goes with them. A perfectly written span placed in a footer never reaches the model — not because it was detected, but because it was in the wrong room when the room was demolished. The corollary is uncomfortable for the person triaging a report later: whether the span arrives can depend on the page's own layout, so the same wording works on one site and vanishes on another. There is a second-order effect worth naming. Boilerplate heuristics generally keep regions with dense, paragraph-shaped text. The properties that make a span survive the *converter* — being ordinary prose — are the same properties that make the region it sits in *look like content*. The two requirements point the same way, which is part of why this construction is stable. ## Property three: no dependence on position or adjacency The converter flattens a tree into a sequence. Whatever the span sat next to, whatever heading governed it, whatever visual grouping framed it — none of that reliably survives, and the order in which regions are concatenated is a pipeline decision. A span whose sense depends on `the following table` or `the instruction above` is a span whose sense is destroyed by the flattening. Self-sufficiency is what survives. ## Property four: it fits inside a chunk The last stage splits the text into passages, and boundaries are chosen at run time by a policy the author cannot see. A span spread over several paragraphs can be cut anywhere; the surviving half is a fragment that reads as noise. So it must be short enough and complete enough that a cut is unlikely to matter. This is a placement and length constraint, not a chunking technique — how a pipeline chooses its boundaries is the pipeline builder's subject, and the attacker's side of it is simply that the choice is not theirs. ## What it costs and where it stops working The cost is exposure. A span with all four properties is legible to any person who opens the page, because visibility to the converter and visibility to a reader are the same thing here. There is no version of this construction that is simultaneously invariant and hidden. That matters when estimating half-life: the construction is not patched by a model update, but it is removed by anyone who reads the page and objects. The limits are equally concrete. A truncation budget that takes only the first portion of the extracted text drops anything below the cut. An agent that never ingests page text — one that works from a summary produced elsewhere — is a different chain entirely. And arriving is not obeying: a span that survives every stage has established that it reached the context window, which is a necessary condition and nothing more. ## How to say it in a loop Answer with the constraint first — *the chain is unobservable, so the span must be invariant rather than tuned* — then give the four properties and, for each, the stage that imposes it. Finishing with the cost (it is visible to humans) and one honest limit (truncation, or the summary-only pipeline) is what separates a mechanism answer from a list.

  • Why does concealment work against the attacker here rather than for them?
    Because concealment means depending on some stage treating the span differently from ordinary prose — and the author cannot see which stages are in the chain. Every dependency is an unobservable bet. Plain prose has no dependency, so it survives by definition of what the stages do. The attacker trades away hiddenness and buys reliability, and in a blind chain reliability is the scarcer good.
  • The same span works on one site and disappears on another. What is the most likely cause?
    Region placement. Boilerplate removal deletes navigation, footers, sidebars and comment furniture wholesale, and different site layouts put the same visual position in different regions. The span was not detected or filtered on the failing site; it was inside a region the extractor discarded. Truncation is the other common cause when the page is long and the span sits far down.
  • The span survived every stage. What have you actually established?
    That the text reached the context window unchanged. Nothing more. Whether the model then preferred it over the operator's instruction is a separate, probabilistic question decided at inference, and a single obeyed run is one observation rather than a rate. Reporting survival as if it were effect is the most common overstatement in findings of this shape.

saying these in an interview costs you the question

  • Thinks the span needs an encoding trick to survive
  • Ignores that boilerplate removal deletes whole regions
  • Assumes chunk boundaries are predictable from outside
  • Writes a span that depends on nearby structure
  • Treats survival through the chain as proof of effect
  • Believes invariance and concealment can be had together

context