skip to content

In RAG, why prepend an LLM-written situating summary to each chunk?

level: seniorimportance: should knowfreq 44%

answer

  1. chunks lose what their document said
  2. pronouns with no antecedent left
  3. make the chunk self-describing
  4. index-time cost, amortised over queries
  5. cache the document, not each chunk

basics

~20 s

Chunks lose their document context. A paragraph reading "revenue fell 3% against the prior period" never says which company, filing or quarter. Prepending a short model-written blurb that situates the chunk in its document makes it self-describing, lifting both embedding and keyword retrieval.

solid answer

~50 s

Take a quarterly regulatory filing split into chunks. One chunk says "revenue fell 3% against the prior period, driven by the segment reorganisation." Embedded on its own, it matches almost nothing a user would actually ask, because the entity, the period and the section are all in text that lives elsewhere in the document. Contextual retrieval fixes this at index time: for each chunk, prompt a cheap model with the document (or its relevant part) plus that chunk, ask for roughly fifty tokens that situate it, and prepend the result before embedding and before indexing for keyword search. Published results report retrieval-failure reductions of roughly a third to a half when combined with lexical search, and around two-thirds when a reranker is added. The cost is one model call per chunk at index time, made tolerable by caching the document across all its chunks — and a cheaper deterministic version, prefixing title, date and heading path, captures a good share of the benefit for free.

code

markdown · 11 lines
markdown
<!-- raw chunk: stranded once separated from its document -->

Revenue fell 3% against the prior period, driven by the segment reorganisation.

<!-- indexed chunk: situating prefix prepended before embedding -->

From the Q3 2025 quarterly filing, industrial segment MD&A section. This passage
explains the year-over-year revenue decline versus Q3 2024 following the segment
reorganisation announced in June 2025.

Revenue fell 3% against the prior period, driven by the segment reorganisation.

go deeper

for a junior

Understand the core problem: after splitting, a chunk can no longer say which document, entity or period it is about, and adding that identity back into the chunk text makes it findable.

for a middle

Explain both versions — a deterministic title-and-heading prefix versus a model-written situating blurb — and why the prefix must be embedded and lexically indexed rather than only stored as metadata.

for a senior

Reason about the pipeline as an operator: index-time cost with document caching, rebuild behaviour on document updates, prefix-length effects on intra-document ranking, and whether the generator sees the prefixed or raw text.

for a principal

Decide whether the gain justifies a permanently more complex indexing pipeline. Insist on measuring the orphan rate and retrieval-failure delta on your own corpus before adopting, and weigh a hybrid-plus-rerank baseline as the alternative spend.

## The orphan-chunk problem Splitting a document destroys anaphora. Inside the whole filing, "the segment reorganisation" and "the prior period" are perfectly well-defined; every chunk after the first inherits meaning from text it no longer sits next to. After splitting, that meaning is gone. The chunk is grammatically fine and semantically stranded. This hurts both retrieval channels. Dense retrieval fails because the embedding encodes a vague statement about unnamed revenue, which sits far from a query naming a specific company and quarter. Lexical retrieval fails because the terms the user typed — the entity name, the year, the section title — literally do not occur in the chunk. ## Two ways to fix it The deterministic version costs nothing: build a prefix out of metadata you already have. Document title, source system, effective date, the heading path, maybe the section number. Prepend it into the chunk text so it is embedded and lexically searchable, and store it as fields too. For structured corpora — well-titled docs with real headings — this alone closes much of the gap and it should always be your baseline. The generated version, popularised as contextual retrieval, goes further. For each chunk you prompt a small model with the surrounding document and the chunk, and ask for a short situating statement: what this passage is, where it sits, what the referring expressions in it point to. Something like "From the Q3 2025 filing of the industrial segment; this passage of the MD&A explains the 3% revenue decline against Q3 2024 following the segment reorganisation announced in June." That gets prepended to the chunk before embedding and before lexical indexing. Generated context beats metadata prefixes when the resolving information is *inside the prose* rather than in the headings — pronouns, comparatives, cross-references to earlier sections, tables whose subject was named three pages up. ## The economics The obvious objection is cost: one model call per chunk, over the entire corpus. Three things make it workable. It is an index-time cost, paid once and amortised over every query the corpus ever serves. A corpus queried thousands of times a day repays a one-off augmentation quickly. The document is the same across all of its chunks, so caching it as a prompt prefix means you pay full price for it once rather than once per chunk. That is the difference between viable and absurd for long documents. And it can use a small, cheap model. The task is descriptive, not analytical. The real costs are elsewhere: pipeline complexity, and rebuild latency. Every corpus update re-triggers augmentation for the affected documents, and a full re-chunk means re-augmenting everything. ## Pitfalls Prefix domination. If the blurb is long relative to the chunk, the embedding starts to represent the prefix. Every chunk from the same document then gets a similar vector, and you have destroyed the intra-document ranking you needed. Keep the blurb short — tens of tokens, not hundreds. Hallucinated context. The model is writing prose that will be indexed as if it were source text. Constrain it: describe the passage, do not interpret or extend it; if the document does not state something, omit it. Some teams keep the prefixed text for the index but hand the *raw* chunk to the generator, so no synthesized sentence can be quoted as fact. That is a defensible default in regulated domains. Staleness. The prefix encodes a claim about the document at augmentation time. If sections are inserted or the document is superseded, prefixes silently drift. Version the augmentation alongside the chunk. And it does not fix bad boundaries. A chunk that straddles two subjects is still confused, now with a paragraph explaining that it is confused about two things. ## Interaction with the rest of the stack The reported gains come from the combination, not the prefix alone: contextual embeddings plus a contextual lexical index plus a reranker over the merged candidates. The prefix specifically helps lexical retrieval a lot, because it injects exactly the proper nouns and dates a keyword index needs and the chunk body lacks. ## Deciding whether you need it Run the diagnostic before building the pipeline. Sample fifty chunks and read each cold: can you tell what document, entity and period it belongs to from the chunk alone? If most of them fail, you have the orphan problem and augmentation will pay. If most pass — because your corpus is well-titled and your headings are good — a deterministic prefix is the whole answer and the model calls buy you little. Then measure retrieval-failure rate on your own question set before and after, rather than importing someone else's percentages.

  • How do you keep the per-chunk model call from being ruinously expensive on a large corpus?
    Cache the document as a shared prompt prefix so it is paid for once rather than once per chunk, use a small cheap model since the task is descriptive, and process in batch rather than interactively. It is also a one-time index-time cost amortised across every future query, so compare it against serving cost rather than against zero.
  • Do you embed the prefixed text and also send it to the generator, or only index it?
    Index the prefixed version so retrieval benefits, but many teams pass the raw chunk to the generator. The prefix is model-written prose sitting next to source text; if it reaches the generator it can be quoted back as if it were the document. Keeping generation on the raw text preserves provenance, which matters in regulated corpora.
  • When does a deterministic metadata prefix do the job just as well?
    When the resolving context lives in structure you already have: a well-titled document, real headings, a date and an entity in the metadata. Prefix those and most chunks become identifiable for free. Generated context earns its cost when the disambiguating information is buried in prose — pronouns, comparatives, references to an earlier section — where no metadata field captures it.
  • What goes wrong if the situating blurb is long?
    The embedding starts to represent the prefix rather than the chunk. Every chunk from one document then carries a nearly identical vector, so the retriever can find the right document but cannot rank passages within it — the exact discrimination you needed. Keep prefixes to a few dozen tokens and check that intra-document ranking has not collapsed.

saying these in an interview costs you the question

  • Claims the prefix replaces good chunk boundaries
  • Writes long prefixes that dominate the chunk embedding
  • Treats the generated context as verified fact from the source
  • Assumes per-chunk generation is unaffordable without checking caching
  • Never re-generates prefixes when the source document changes

context