skip to content

Why is page-one placement of an injected span in an uploaded document not a general rule?

level: juniorimportance: should knowfreq 55%

answer

  1. the model does not read the file
  2. something cuts it up first
  3. one selected chunk, or a truncated window
  4. which path opens it decides the offset

basics

~20 s

Page one matters only if the path reading the file starts there. A chunk retriever forwards one selected passage, not the opening page; a summariser over a long file keeps only what fits its budget. Placement follows the path.

solid answer

~50 s

Because the model almost never reads the file. Something in between decides which bytes become context. In a retrieval path an extractor turns the upload into text, a splitter cuts that text into chunks, and the query selects one or a few of them — the opening page is just another chunk, and it arrives only if it is the one selected. In a summarising path over a file far longer than the context window, the pipeline truncates: some keep the head, some keep head and tail, some map over sections and merge. Only there does "the top" carry real force. So the first thing settled is which path will open the file; the offset is chosen relative to that path's window or chunk boundary. Choosing wrong is the ordinary reason a planted span simply never arrives — and it arrives as silence, not as a refusal.

go deeper

for a junior

Be ready to say that an upload becomes extracted text, then either a selected chunk or a truncated window, and that only that last artefact reaches the model.

for a middle

Explain the two paths concretely: selection by similarity versus truncation against a budget, and what each does to text sitting at a given offset.

for a senior

Show that you settle which path reads the file before arguing about the offset, and that a span which never arrives produces no signal you can distinguish from a dozen other causes.

for a principal

Own the framing that placement is a property of the pipeline rather than of the document, so any claim about it is scoped to one consumer of that corpus and does not transfer to the next one.

## Three different things get called "the document" When somebody uploads a long file — a vendor contract, a filing, a scanned appendix — at least three artefacts exist afterwards, and they are not the same object: 1. **The bytes uploaded.** What a reviewer opens and scrolls, with pages, headers and layout. 2. **The text an extractor emitted from those bytes.** A different artefact: multi-column layout is linearised, headers and footers may be repeated or dropped, tables are flattened, a cover page may or may not appear first, and for a scan the text is *created* by optical character recognition rather than read out of the file. Nothing guarantees that "page one" of the upload is the first text in the extraction. 3. **What actually enters the model's context.** Either a chunk a retriever selected, or the slice of the extracted text that survived a truncation budget. The intuition behind "put it at the top so the model reads it first" is an intuition about artefact 1, read by a human, front to back. The model reads artefact 3. ## The retrieval path In a retrieval path the extracted text is cut into chunks, each chunk is embedded, and at query time a similarity search returns the nearest few. What reaches the model is a small set of passages, usually a few hundred to a couple of thousand tokens each, stripped of their neighbours. Ordinal position in the file has no role in that selection. A chunk from page 180 and a chunk from page 1 compete on the same footing: distance in a vector space between the query's embedding and the passage's. So an authored span sitting at the top of a 200-page file has no privileged status at all — it is in the index like everything else, and on any given question it is probably not returned. (Whether a planted passage *wins* selection is a separate matter, and not this question: assume selection happens, and ask whether the span is intact in what was handed over.) ## The summarising or fetch path A summarising pipeline over the same file has the opposite shape: it is trying to read everything, and it fails only because the file exceeds the budget. What it does then varies, and each variant implies a different answer to "where": - **Head-only truncation** — the first N tokens are kept, the tail is discarded. Here "the top" really does work. - **Head-and-tail** — an opening window and a closing window are kept and the middle is dropped. The top works; so does the very end; the middle is a dead zone. - **Map-then-reduce over sections** — every section is summarised separately and the summaries are merged. Nothing is dropped outright, but each section is read in isolation, and the merge step sees only summaries. Note what varies here: not the file, but the pipeline. The same upload, handed to two consumers, has two different answers for where a span has to sit. ## Why this is the first decision, not a detail Everything else about an injected span — what it claims, what it asks for, what it gets past — is conditional on the span being present in the context at all. Placement is the arrival question, and it is settled before any of the rest matters. An engineer who reasons about a span's wording without first asking which path will read the file has skipped the step that decides whether the experiment ran. ## Silence is not evidence The failure mode when placement is wrong is *nothing observable*. No refusal, no block, no error — the model summarises the contract normally, because the span was never in front of it. This matters because the same silence is produced by several very different causes: the span sat outside the truncation window; the chunk holding it was never selected; the extractor never emitted it; the model read it and ignored it; a screening layer removed it. A candidate who reads "nothing happened" as "it was filtered" has inverted the direction of the evidence. Absence of an effect proves only that no effect was observed. ## The practical consequence Because a corpus is usually read by more than one consumer — summarised once at upload, queried chunk-wise for months afterwards — an offset that satisfies one path may be irrelevant to the other. A single "where do you put it" answer that ignores the path is not a partial answer; it is a coin flip.

  • If a planted span never arrives, what does the attacker actually observe?
    Nothing. The pipeline behaves normally. That silence is compatible with several causes: the span fell outside a truncation window, its chunk was never selected, the extractor never emitted it, the model read it and ignored it, or a screening layer removed it. Only the first three are placement problems, and they cannot be told apart from the output alone — which is why an unqualified "it didn't work" is not yet a result.
  • Does an offset in the extracted text match the page a person sees?
    Not reliably. Extraction linearises multi-column layout, may repeat or drop headers and footers, flattens tables, and for a scanned page creates the text through optical character recognition rather than reading it out of the file. The reviewer's page three and the extractor's first few thousand characters are different coordinates, and only the second one is what a truncation budget measures.
  • Why can one file need two different placements?
    Because a corpus usually has more than one consumer. The same upload may be summarised whole at ingest and then queried chunk-wise by a search front end for months. The summariser's constraint is a budget over the whole extracted text; the retriever's is a boundary around one passage. An offset chosen for one is not chosen for the other.

Putting it on page one is like taping a note to the cover of a book that a photocopier will feed page by page into two different machines — one copies only a handful of pages it picks by keyword, the other copies the first fifty and the last ten.

saying these in an interview costs you the question

  • Assumes the model reads the whole uploaded file end to end
  • Treats a retrieved chunk as if it were the whole document
  • Thinks the top of a file is always inside the context window
  • Reads silence as proof that a screening layer blocked something
  • Confuses the reviewer's page numbers with offsets in extracted text

context