skip to content

In a RAG prompt, why tag each retrieved chunk with a source ID?

level: juniorimportance: must knowfreq 74%

answer

  1. assigned by your code, not the model
  2. closed set of markers
  3. URL lives in a server-side map
  4. fabricated identifiers look like verification
  5. triage axis: retrieval vs generation

basics

~20 s

Source IDs give the model a fixed vocabulary of references it can cite and give the reader a way back to the evidence. Without IDs injected alongside the text, any citation the model produces is invented rather than retrieved.

solid answer

~40 s

The IDs are assigned by your assembly code, not by the model. When you build the prompt you wrap each retrieved chunk in a short marker such as `[S3]`, usually with a tiny header (document title, section heading, date), and keep a server-side map from marker to the real record — document id, page, character range, URL. The model is told to cite only markers that appear above the answer, so citation becomes a closed-set choice rather than free generation. That matters because a model asked for a URL or a section number it was never shown will produce something well-formed and false. It also makes citations machine-checkable: any marker outside the injected set is a detectable error, and the renderer resolves the valid ones into real links for the user.

code

python · 17 lines
python
chunks = [
    {"doc": "benefits-2026.pdf", "section": "4.2 Vesting",
     "text": "Employer match vests after two years of service."},
    {"doc": "benefits-2026.pdf", "section": "4.3 Rollovers",
     "text": "Rollovers from a prior plan are accepted within 60 days."},
]

marker_map = {}
blocks = []
for i, chunk in enumerate(chunks, start=1):
    marker = f"S{i}"
    marker_map[marker] = chunk
    blocks.append(f"[{marker}] ({chunk['section']})\n{chunk['text']}")

context = "\n\n".join(blocks)
print(context)
print("resolves to:", marker_map["S1"]["doc"])

go deeper

for a junior

Be ready to say that source markers are added by your code when the prompt is built, that the model can only cite what it was shown, and that the application maps a marker back to the real document.

for a middle

Explain the closed-set idea: restrict citations to markers present in the prompt so an unknown marker is a detectable error, and keep URLs, ids and offsets in a server-side map instead of the context.

for a senior

Show how you use citations operationally — a wrong document cited is a retrieval defect, the right document with an unsupported claim is a generation defect — and name the failure modes like shotgun citation and uncited sentences.

for a principal

Own what a citation promises the user. Decide the attribution granularity the product commits to, what that commitment means in a regulated setting, and what verification spend keeps the promise honest rather than decorative.

## What attribution is for A retrieval-augmented system answers from passages fetched at query time. Two things can independently go wrong: retrieval can return the wrong passages, and the model can say something the passages do not support. Attribution — a marker on each claim pointing at the passage it came from — is the mechanism that lets a reader, and your own operations team, tell those two failures apart. It is a prompt-side contract: it only works if the prompt carries the identifiers in the first place. ## Where the identifier comes from The identifier is minted by the code that assembles the prompt, at the moment the retrieved set is turned into text. A common shape is a short opaque marker per chunk — `[S1]`, `[S2]`, … — emitted immediately before the chunk body, with a one-line header giving whatever the model needs to judge relevance: document title, section heading, effective date. Everything else stays in a server-side dictionary keyed by that marker: the document id, the URL, the page number, the character offsets, the tenant. The model never sees the URL, and it does not need to; it emits `[S2]` and your renderer swaps in a link. ## Why an opaque marker beats asking for the real reference Asking the model to output "the section number" or "the source URL" invites fabrication whenever that string is not verbatim in the context. Fabricated citations are the worst kind of error because they look like verification: a plausible-looking identifier makes readers stop checking. Restricting the citation vocabulary to markers physically present in the prompt turns the task from generation into selection, which models do far better, and it gives you a trivially cheap validity test — the set of markers in the answer must be a subset of the set you injected. ## What to put in the header Only what earns its tokens. Title and section help the model decide whether the passage is on-topic and help the user orient when the citation is rendered. A date matters when the corpus has plan years, guideline revisions or policy versions, because it lets the model prefer the current one and lets the reader see which it used. Internal record ids, access-control fields and file paths belong in the server-side map: they cost tokens, help nothing, and are exactly the sort of internal detail you do not want reflected back into an answer. ## Granularity Document-level markers are cheap but weak — "it's somewhere in this 90-page handbook" is not verification. Chunk-level is the normal choice and matches the retrieval unit. Sentence- or span-level attribution is stronger still and is what enables click-to-highlight review, but it needs offsets plumbed through chunking and more discipline in the answer format. ## The instruction that goes with it Markers alone do nothing; the prompt must state the contract. Typical wording: cite the marker of every passage you used, placed immediately after the sentence it supports; use only markers that appear in the passages above; if no passage supports a sentence, do not write that sentence. Stating placement matters — a pile of markers at the end of a paragraph tells a reader nothing about which claim came from where. ## Failure modes to expect *Citation drift*: the model cites a marker that exists but does not contain the claim. The closed-set check passes; the citation is still wrong. Catching this needs a separate support check. *Shotgun citation*: every sentence carries every marker. Technically valid, informationally worthless — worth measuring citation precision, not just presence. *Uncited claims*: sentences with no marker at all, usually the model's own background knowledge leaking in. These are the highest-risk sentences in the answer. *Markers eaten by rendering*: bracketed markers can be mangled by markdown pipelines or stripped by a summarizer downstream. Test the whole path, not just the model output. ## The operational payoff Beyond user trust, citations are your triage axis. Take a wrong answer and look at the cited chunk. If the cited chunk is the wrong document, the retrieval stage failed and you tune indexing, chunking or ranking. If the cited chunk is the right one but does not actually contain the claim, the generation stage failed and you tune the prompt contract, the model, or add a verification pass. Without attribution, every bad answer looks the same and every fix is a guess.

  • If the model cites a marker that exists but does not contain the claim, what does that tell you?
    That the closed-set check only proves the reference resolves, not that it supports the sentence. It points at the generation stage rather than retrieval: the right passage was in front of the model and it still asserted something the passage does not say. The fix lives in the prompt contract, the model choice, or a separate support check that compares the claim against the cited text.
  • Why not just include the document URL in the chunk header so the model can cite it directly?
    You can, but it buys little and costs something. The model now spends tokens copying a long string it can get wrong by a character, and a mistyped URL still looks authoritative. A short marker resolved server-side is cheaper, exactly checkable, and lets you change how a source is rendered — deep link, page anchor, internal viewer — without touching the prompt.
  • Should markers be stable across requests or renumbered each time?
    Renumber per request, because the marker's only job is to index the passages in this prompt. What must be stable is the mapping you store alongside the response, so a logged answer can still be resolved to its sources weeks later. Stable global ids inside the prompt tend to be long, waste tokens, and give the model more chance to emit one it half-remembers.

saying these in an interview costs you the question

  • Assumes the model knows real URLs or page numbers for retrieved documents
  • Thinks a well-formed citation string proves the source exists
  • Lets the model cite documents from memory rather than the injected set
  • Treats a resolvable marker as proof the claim is supported
  • Dumps internal record ids and paths into the injected chunk header

context