skip to content

Placement Inside a Document

A retriever may never return the chunk carrying the span, while a summariser reads the whole file until its budget cuts the tail. Interviewers probe which path you asked about.

on this pageshow

explore

questions

4

Why does an attacker's span placement differ between a whole-file summariser and a chunk retriever?

level: middleimportance: must knowfreq 62%

answer

  1. the unit of survival changes
  2. file minus budget, or one chunk alone
  3. a boundary is not a budget
  4. neighbours do not travel with a chunk
  5. short and self-contained on one side

basics

~20 s

The unit of survival changes. A summariser forwards the file minus whatever truncation drops, so one occurrence anywhere inside the surviving window is enough. A retriever forwards a single chunk, so the span must sit whole inside one cut region.

solid answer

~50 s

In the summarising direction the unit is the file and the obstacle is a budget: the pipeline wants all of it and keeps only what fits, so the question is whether the offset falls inside the surviving window. The span can be long, can span paragraphs, and can lean on surrounding text, because that text travels with it. In the retrieval direction the unit is the chunk and the obstacle is a boundary: whichever passage gets selected arrives alone, stripped of its neighbours. So the span must be short enough and self-contained enough to sit entirely inside one cut region — a span straddling a cut arrives halved, and anything it depends on from the preceding paragraph is simply not there. That inversion is why a single answer to "where do you put it" is wrong: the same file, read two ways, has two different constraints.

go deeper

for a junior

Recall that a retriever hands over one passage while a summariser hands over as much of the file as fits, and that those are different amounts of text.

for a middle

Explain the inversion in mechanical terms: a budget with a moving edge on one side, a fixed cut point around an isolated passage on the other, and what each implies about length.

for a senior

Bring in the multi-stage case — an ingest summary and a later extraction pass read different slices — and be able to say which consumer a given offset actually reaches.

for a principal

Be ready to argue that a claim about placement is only meaningful once the consuming stage is named, and that a corpus with several consumers cannot be characterised by one result.

## The inversion, stated plainly | Path | What is forwarded | Obstacle | What placement optimises | |---|---|---|---| | Summarising or whole-page fetch | the extracted text minus what truncation drops | a token budget, and the window's edges | falling **inside the surviving window** | | Chunk retrieval | one passage the search returned, alone | a cut point the author cannot see | sitting **whole inside one cut region** | Both are called "getting the text into the context", and they impose nearly opposite constraints. Reasoning about one while standing in the other is the common error. ## The summarising direction: a budget problem A pipeline that summarises an uploaded file wants the whole thing. It fails only because the extracted text is longer than what fits, and it resolves that with truncation — head-only, head-and-tail, or a map over sections whose summaries are merged. Consequences for a span placed in a file the attacker authored end to end: - **Length is cheap.** The span can run to several paragraphs; nothing cuts it in the middle unless it happens to straddle the window edge. - **Context travels.** Surrounding text arrives with it, so a span can rely on the section it sits in — a heading above it, a clause before it — to do part of the work. - **The edge is the enemy, and it moves.** The surviving window is a function of the budget, and the budget is a function of everything else in that turn's context. The same offset can be inside the window one run and outside it the next. ## The retrieval direction: a boundary problem A retrieval path cuts the same extracted text into chunks ahead of time and, at query time, hands the model a small number of them. Two facts dominate: 1. **The chunk arrives alone.** Its neighbours are not in the context. A span whose force comes from the paragraph before it loses that paragraph. 2. **The cut points are not the author's to choose.** They are decided by the splitter's configured size, its unit (characters or tokens), and the exact text the extractor emitted. A span that happens to straddle one is delivered halved — and half a span is usually inert rather than half-effective. So in this direction the constraint is *self-containment within a chunk-sized region*: short, complete on its own, and not dependent on anything outside its own boundaries. Length now costs, because every extra sentence raises the chance a cut lands inside it. One thing this question does **not** turn on: whether the chunk is selected in the first place. Ranking a planted passage into the result set is its own problem. Here, assume something in that region is retrieved, and ask only whether what arrives is intact. ## The complication worth raising: stages, not just paths Real ingest pipelines are rarely one pass. A diligence workspace may summarise an upload at ingest, then run a second, narrower pass that extracts structured fields — renewal date, liability cap, counterparty — into a record other systems consume. That second pass usually reads a *different slice*: the sections a first pass flagged, or a window around each field's likely location, not the whole file. That means the honest question is not "summariser or retriever" but "which stage's window am I placing into", and the two can disagree. An offset comfortably inside the summarising window can sit outside the extraction pass's slice, in which case the prose summary a reader sees is affected and the structured record downstream is not — or exactly the reverse, which is the more interesting case, because the record is read by systems and not by a person. ## What the difference costs The retrieval direction demands brevity and self-containment; the summarising direction tolerates length but is hostage to a moving edge. Satisfying both in one document means either accepting a lower hit rate on one, or placing more than once — which is not free, and is the subject of its own reasoning about redundancy. ## The tell in an interview A candidate who answers "where do you put it" with a single offset has not asked the question that decides it. The strong answer names the path first, names the unit that path forwards, and only then talks about position — and notes that a corpus with two consumers has two answers.

  • In the retrieval direction, why does a longer span hurt?
    Because every additional sentence increases the chance a cut point lands inside it, and a span delivered in halves is usually inert rather than partially effective. The retrieval direction rewards a short, complete unit that fits comfortably inside one cut region; the summarising direction has no equivalent penalty, since nothing cuts the middle of a passage that lies inside the surviving window.
  • What does "self-contained" actually mean for a retrieved chunk?
    That the span carries everything it relies on inside its own boundaries. A retrieved chunk arrives without its neighbours, so a heading two paragraphs up, a definition earlier in the section, or a clause that set up the framing are simply absent. Anything the span leans on from outside the chunk is not there when the model reads it.
  • Where does a two-stage pipeline change the answer?
    When a first pass summarises and a second extracts structured fields, each pass has its own window over a different slice of the file. An offset inside the summariser's window can sit outside the extraction pass's, so the prose a person reads and the record a downstream system consumes can diverge. The choice is between stages, not just between paths.
  • Does the summarising direction have a stable target at all?
    Less than it looks. The surviving window depends on the budget left after everything else in that turn — conversation history, other retrieved material, tool output — so the edge moves between runs. An offset near the boundary is inside it sometimes and outside it other times, which is a common source of intermittent behaviour.

saying these in an interview costs you the question

  • Thinks a retriever hands the model the whole document
  • Assumes a chunk's neighbours travel with it into context
  • Treats a truncation budget and a chunk boundary as one obstacle
  • Believes one placement serves every consumer of the same file
  • Says a halved span still works at reduced strength

context

open as a page

Why is page-one placement of an injected span in an uploaded document not a general rule?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Page one matters only if the path reading the file starts there. A chunk retriever forwards one selected passage, not the opening page; a summariser over a long file keeps only what fits its budget. Placement follows the path.

open as a page

A filed injection finding reproduces once in five runs of the same uploaded file — is it a finding?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It can be, but not yet. First establish whether the span reached the model on the runs that worked and not on the others. Intermittent arrival is a plumbing result; intermittent obedience after arrival is a different result with a different owner.

open as a page

In a document you authored, you cannot see where the splitter cuts — what follows for placing a span?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Placement becomes probabilistic rather than exact, so the construction shifts to several short, self-contained occurrences spread across the file. Each copy raises the chance one lands whole inside a cut region, and each adds surface.

open as a page