skip to content

In Cohere's v2 Chat API, how do you map a citation back to text and source?

level: middleimportance: must knowfreq 52%

answer

  1. start and end are offsets, not tokens
  2. end is exclusive
  3. citation echoes its own text
  4. raw output only, escape afterwards
  5. sources carry the ids you supplied

basics

~20 s

Each citation carries start and end character offsets into the generated text plus the sources that support that span. Slice the raw answer with text[start:end] and join each source to the document id you supplied. Offsets index the untouched text, so apply them before any escaping or trimming.

solid answer

~50 s

A Cohere citation is an offset pair plus an attribution. `start` and `end` are character offsets into the generated answer, with `end` exclusive, so the cited substring is exactly `answer[start:end]` — the citation also echoes that substring so you can assert the slice matches. Alongside them sits the list of supporting sources; for document-grounded answers each source resolves back to an entry in the `documents` array you sent, by the `id` you chose. (The older v1 chat shape exposed this as a flat `document_ids` list; v2 generalises it so a source can be a tool output as well as a document.) The operational rule is that offsets are indexes into the *raw* model output: HTML-escape, trim, or markdown-render the answer first and every span silently shifts. Segment on the raw string, then transform each segment.

code

python · 20 lines
python
import cohere

co = cohere.ClientV2(api_key="<key>")

res = co.chat(
    model="command-a-03-2025",
    messages=[{"role": "user", "content": "When can a customer get a refund?"}],
    documents=[{"id": "policy-1", "data": {"text": "Refunds are accepted within 30 days of delivery."}}],
)

answer = res.message.content[0].text
segments = []
cursor = 0
for citation in sorted(res.message.citations or [], key=lambda c: c.start):
    assert answer[citation.start:citation.end] == citation.text
    if citation.start > cursor:
        segments.append((answer[cursor:citation.start], None))
    segments.append((answer[citation.start:citation.end], citation.sources))
    cursor = citation.end
segments.append((answer[cursor:], None))

go deeper

for a junior

Recall that a citation gives you start and end offsets plus the supporting sources, and that slicing the answer with those offsets gives you the exact cited words.

for a middle

Explain that end is exclusive, that the citation echoes its own text so you can assert the slice, and that sources join back to the document ids you supplied.

for a senior

Demonstrate the rendering discipline: segment the raw text before escaping or rendering, do offset arithmetic in exactly one place, and handle streamed citation events against a completed buffer.

for a principal

Frame citations as an attribution and observability signal rather than a correctness proof, and define what your product does with spans that carry none.

## The three parts of a citation A citation from Cohere's grounded chat answers one question — *which part of the answer came from where* — and it does it with three pieces of data: 1. **`start`** — a character offset into the generated answer text. 2. **`end`** — the exclusive end offset, so the covered span is `answer[start:end]`. 3. **The supporting sources** — for a document-grounded call, entries that resolve back to the items you passed in `documents`, carrying the `id` you assigned. A citation also echoes the cited text itself, which is a gift: it gives you a cheap assertion. If `answer[start:end] != citation.text`, something in your pipeline has mangled the string, and you should fail loudly rather than render misaligned highlights. ## Why offsets rather than inline markers An alternative design — the one you build yourself when the API does not offer this — is to prompt the model to emit inline markers like `[1]` and then parse them out. Offsets are strictly better for a product surface: the answer text stays clean prose with no marker syntax to strip, one span can have several supporting sources without the text getting noisier, and overlapping or nested attributions are expressible. The cost is that offsets are fragile under any string transformation, which is the failure mode you will actually hit. ## The rendering trap This is the practically important part, and the one interviews probe. Offsets index the exact string the model produced. Anything you do to that string before applying the spans invalidates them: - **HTML escaping.** Turning `&` into `&amp;` grows the string; every offset after that point is now wrong by four characters. - **Trimming or collapsing whitespace.** Shifts everything after the edit. - **Markdown rendering.** The rendered DOM has no straightforward relationship to the source offsets at all. - **Concatenating multiple content blocks in the wrong order**, or dropping one, when the response has more than one text block. The correct order is always: take the raw text, walk the citations sorted by `start`, cut the string into alternating uncited and cited segments, and *then* escape or render each segment independently. Do it the other way round and your highlight lands in the middle of a word. There is a second, subtler variant of the trap: languages disagree about what a "character" is. Python indexes by Unicode code point; JavaScript strings index by UTF-16 code unit, so an emoji or an astral-plane character counts as two. If your backend computes segments and your frontend re-slices the same string, do the slicing in exactly one place and pass the segments, not the offsets, across the boundary. Otherwise multi-byte content shifts your highlights. ## Joining back to your data The source side is the easy half, provided you sent object-form documents with your own ids. The citation's sources give you those ids, and you look them up in the same map you built when assembling the request. That map is what lets the UI render "source: Refund policy, §3" with a working link, and it is what lets you log, per answer, which knowledge-base records were actually used — a genuinely useful retrieval-quality signal that costs nothing extra. If you passed bare strings as documents, you have no id of your own to join on and you are reduced to matching on position or content. That is why object form is the production default. One generalisation worth knowing: a supporting source is not always a document. When the grounding material arrived as the result of a tool call rather than through the `documents` array, the citation points at that tool output instead. Your mapping code should therefore branch on the source's type rather than assuming every source is a document — code that blindly reads a document id will break the first time someone adds a search tool to the same endpoint. ## Streaming When you stream the response, citations do not all arrive at the end; they are emitted as their own stream events interleaved with the text deltas. The safe consumer pattern is to accumulate the text deltas into a buffer and collect citations as they arrive, then apply the spans against the fully accumulated buffer once the stream completes. Applying a span against a partially received buffer will occasionally index past its end. ## What citations do not tell you A citation says the span is attributable to a document you supplied. It does not say the document is correct, current, or the best one available, and it does not say the model's paraphrase is faithful in nuance. Treat citations as an attribution mechanism and a debugging aid, not as a correctness proof.

  • Your UI HTML-escapes the answer before applying spans and the highlights drift. What is happening?
    Escaping changes the string length — one ampersand becomes five characters — so every offset after the first escaped character points at the wrong place. Offsets index the raw model output. Segment the raw string first using the citation offsets, then escape each segment independently before rendering. The same applies to trimming and markdown rendering.
  • How do citations reach you when the response is streamed?
    They arrive as dedicated citation events interleaved with the content deltas rather than in one batch at the end. Accumulate the text deltas into a buffer, collect citations as they come, and only apply the spans against the completed buffer. Slicing a partially built buffer can index past its end or land mid-token.
  • Why should your mapping code not assume every citation source is a document?
    A supporting source can also be the output of a tool call, when the grounding material came from a tool rather than the documents array. Branch on the source type and handle both, otherwise the code breaks the day someone adds a search tool to the same endpoint. Keep an internal citation type that both shapes normalise into.

saying these in an interview costs you the question

  • Thinks start and end are token indexes
  • Applies spans after HTML-escaping or trimming the answer
  • Assumes end is inclusive and slices one character short
  • Mixes Python code-point offsets with JavaScript UTF-16 slicing
  • Believes a citation proves the answer is factually correct

context