Why wrap each retrieved chunk in explicit delimiters in a RAG prompt?
answer
- the model needs a seam it can see
- two similar chunks blend into one
- tagged blocks with a short header
- header carries what distinguishes the sources
- delimiter must not occur in the text
basics
~20 sExplicit delimiters tell the model where one retrieved document ends and the next begins. Without them, chunks read as one continuous passage, and the model blends facts from unrelated sources into a single confident answer.
solid answer
~50 sConcatenating chunks with blank lines gives the model no boundary signal at all — it sees one long passage and has to infer document breaks from topic shifts, which it does badly when two chunks are topically similar. That is exactly the dangerous case: two suppliers' contracts, or two versions of the same policy, merge into one answer that is a fluent mixture of both and matches neither. Wrapping each chunk in a tagged block fixes the boundary and gives you a place for a short header — source, title, section, date — so the model can tell the chunks apart on the attributes that actually distinguish them. Keep the header to a line: you pay its token cost once per chunk, multiplied by k. And choose a delimiter that cannot occur inside the chunk text, or escape it on ingest, otherwise the content can manufacture a false boundary.
code
python · 8 linesdef render_context(chunks):
blocks = []
for c in chunks:
blocks.append(
f'<doc source="{c["source"]}" section="{c["section"]}" '
f'updated="{c["updated"]}">\n{c["text"]}\n</doc>'
)
return "\n".join(blocks)go deeper
Be able to say that retrieved chunks are separate documents pasted together, and that without an explicit boundary the model can merge facts from two of them into one answer.
Explain the mechanism and the header: the model infers seams from content and fails when chunks are topically similar, so the block header carries the distinguishing metadata — source, section, date — that the chunk text itself does not.
Show the ingest-side discipline: delimiters validated against the corpus so content cannot forge a boundary, a uniform block format, header fields chosen for disambiguation value, and prompts logged in a shape a human can read during an incident.
Own the prompt's structural contract across the system — one block format, one place that renders it, and a stated budget for scaffolding versus evidence — so every team's retrieval feeds a prompt shape the org can reason about and evaluate consistently.
## The failure the delimiter prevents A RAG prompt assembles k independently retrieved passages that were never written to sit next to each other. If you join them with newlines, the model receives what looks like one document. It must then infer the seams from content alone — and inference works fine when the chunks are about obviously different things, and fails precisely when they are not. The canonical failure is two topically identical chunks from different entities. Retrieve the payment-terms clause from supplier A's contract and the payment-terms clause from supplier B's, join them with a blank line, and ask "what are the payment terms?" The answer will be a fluent synthesis with A's net-30 and B's early-payment discount in one paragraph, describing an agreement that does not exist. Nothing in the model's input said those were two agreements. The same thing happens with two versions of a policy, two regions' pricing pages, or a current and a superseded procedure. ## What a delimiter actually gives you Three things: **A hard boundary.** A block structure — an XML-ish tag, a fenced region, a repeated rule line — marks unambiguously where one source ends. The model no longer needs to guess, and neither does the human reading the logged prompt during an incident. **A place for a one-line header.** The attributes that distinguish otherwise-identical chunks — which supplier, which region, which version, what date — usually live in metadata, not in the chunk text. A header attached to the block carries them into the prompt where the model can use them: `<doc source="supplier-b-msa" section="Payment" updated="2026-03-11">`. Without it, the model literally cannot know that two similar clauses belong to different parties. **A referenceable unit.** Once each block has an identifier, the rest of the prompt can talk about "the documents below" as discrete objects rather than as prose. ## Choosing the delimiter Any consistent structure works; the choice matters less than the consistency. Practical constraints: - **It must not appear inside the chunk text.** If your corpus contains HTML or markup and you delimit with tags, a chunk can close your block early and everything after it reads as prompt-level text. Escape or strip the delimiter characters at ingest, or pick a form the corpus provably does not contain, and validate that assumption in the ingest pipeline rather than assuming it. - **It should be cheap.** Every wrapper is paid k times. A dozen tokens per chunk across twenty chunks is a few hundred tokens off your evidence budget — acceptable. A verbose multi-line preamble per chunk is not, and it also dilutes the block with boilerplate the model has to read past. - **It should be uniform.** Same shape for every chunk, same attribute order, same header fields even when a field is empty or unknown. Irregular structure is itself a signal, and an unintended one — a chunk formatted differently reads as special. ## Keep the header proportionate A 100-token chunk under a 60-token metadata block is mostly metadata, and the ratio matters: you are spending evidence budget on bookkeeping and burying the actual text. Include only fields that disambiguate or that the model may need to reason about — typically source, title or section, and a date when recency is decision-relevant. Everything else stays in your logs and in the response payload rather than in the prompt. ## Where the boundary helps beyond blending Boundaries also make the retrieved region distinguishable from your own instructions. A prompt where the instructions, the question and the corpus text are all undifferentiated prose is harder for the model to navigate: it cannot cleanly tell which sentences are directives it should follow and which are material it should read. A consistently delimited evidence region gives that separation structurally rather than by wording alone. And for debugging, a delimited prompt is a readable artifact. When a RAG answer is wrong, you open the logged prompt and can see, at a glance, which documents were present and in what shape — instead of scanning an undifferentiated wall of text trying to work out where chunk 7 started. ## Common mistakes Joining with blank lines and assuming topic shift is enough of a signal. Using a delimiter that occurs in the corpus. Varying the format between chunks. Padding each block with metadata nobody uses. And formatting the blocks so elaborately that the prompt is more scaffolding than evidence — the structure exists to serve the text, not to decorate it.
- What goes wrong if your delimiter string can also appear inside the chunk text?A chunk can close the block early, and everything after that point reads as prompt-level text rather than as retrieved evidence — a false boundary manufactured by the corpus. The fix is at ingest: escape or strip the delimiter characters, or choose a form the corpus provably does not contain, and validate that assumption in the pipeline instead of assuming it holds.
- How much metadata should ride in each chunk's header?Only what disambiguates or is reasoned over — typically source, title or section, and a date when recency matters. You pay the header once per chunk, multiplied by k, so a 60-token header on a 100-token chunk is mostly bookkeeping and buries the evidence. Everything else belongs in your logs and the response payload, not the prompt.
- Does the exact delimiter format matter, or just that one exists?Consistency matters far more than the specific form. Tagged blocks, fenced regions and repeated rule lines all work. What breaks things is varying the format between chunks — irregular structure reads as significance the model was not meant to infer — or choosing a form that collides with the corpus content.
saying these in an interview costs you the question
- Joins chunks with blank lines and expects clear boundaries
- Uses a delimiter string that also occurs inside the chunk text
- Wraps a 100-token chunk in a 60-token metadata preamble
- Believes the model reliably infers document seams from topic shifts
- Formats each chunk differently depending on its source type