skip to content

Grounded generation and citations

Cohere's differentiator: you hand the chat endpoint your documents and it answers with citations pointing at which document supported each span. It is RAG assembled inside a single API call rather than in your own code.

on this pageshow

questions

6

In Cohere's v2 Chat API, how do you pass source documents so the reply carries citations?

level: juniorimportance: must knowfreq 62%

answer

  1. documents is its own request parameter
  2. not glued into the user message
  3. string form or id-plus-data form
  4. ids you choose come back in citations
  5. no documents means no citations

basics

~20 s

Send a top-level documents array alongside messages in the v2 chat call. Each entry is a plain string or an object with an id and a data map. The reply's message.citations then link answer spans to those documents.

solid answer

~40 s

Cohere's v2 chat endpoint takes `documents` as its own request parameter, a sibling of `messages` — you do not paste retrieved text into the user turn yourself. Each element is either a raw string or an object shaped `{"id": ..., "data": {...}}`, where `data` is a flat map of field names to string values that the model sees as a small structured record. Cohere then runs its grounded-generation path: it answers from those documents, and the response carries a `citations` list next to the generated text, each citation naming a span of the answer and the sources behind it. Because the ids are yours, a citation maps straight back to the chunk or row you retrieved. Pass no documents and you get an ordinary ungrounded completion with no citations at all.

code

python · 17 lines
python
import cohere

co = cohere.ClientV2(api_key="<key>")

res = co.chat(
    model="command-a-03-2025",
    messages=[{"role": "user", "content": "What is our refund window?"}],
    documents=[
        {"id": "policy-1", "data": {"title": "Refunds", "text": "Refunds are accepted within 30 days of delivery."}},
        {"id": "policy-2", "data": {"title": "Shipping", "text": "Orders ship within 2 business days."}},
    ],
)

answer = res.message.content[0].text
print(answer)
for citation in res.message.citations:
    print(answer[citation.start:citation.end], citation.sources)

go deeper

for a junior

Know that documents is a top-level parameter next to messages, that each entry can be a string or an object with id and data, and that the response carries a citations list.

for a middle

Be ready to explain why the id you choose matters, what data field names do for the model, and why passing no documents yields an ordinary ungrounded completion.

for a senior

Show that you cost and secure the call: documents dominate input tokens on every turn, and retrieved text is untrusted input regardless of which parameter carries it.

for a principal

Own the boundary question — how much of grounding you delegate to a vendor parameter versus keep in your own prompt assembly, and what that choice costs you if you change providers.

## What the feature is Most providers make retrieval-augmented generation your problem: you retrieve, you concatenate the results into a prompt, and you invent your own convention for making the model cite. Cohere's chat endpoint moves part of that inside the API. You hand it your retrieved material in a dedicated request field, and it hands back the answer *plus* a machine-readable list of which supplied document supported which stretch of that answer. Nothing about how you found the documents changes — that is still your retrieval pipeline — but the prompt assembly and the citation format become the API's job rather than yours. ## The request shape On the v2 chat call the parameters that matter here are: - `model` — a Command-family model id. - `messages` — the conversation, in the usual role/content form. - `documents` — a list of the sources the answer must be grounded in. `documents` is top-level. It is not a message, not a role, not a string you glue onto the last user turn. Two element forms are accepted: 1. **A plain string.** Simplest, and fine for a quick experiment, but you get no identifier of your own back in the citations, so you are left matching on position or text. 2. **An object with `id` and `data`.** `id` is any stable key you choose — a chunk id, a database primary key, a URL. `data` is a flat map of field names to string values, for example `{"title": ..., "text": ..., "url": ...}`. The model sees those fields as a small labelled record, so field names carry meaning; calling a field `text` or `snippet` beats calling it `f1`. In production you almost always want form 2, because the id is what makes a citation actionable in your UI. ## What comes back The response's assistant message carries the generated text in its content blocks and, separately, a `citations` array. Each citation identifies a span of the generated text and the sources that support it. The sources refer back to the documents you supplied, keyed by the ids you gave them. That is the whole contract: text in one place, attribution in another, joined by offsets and ids. Crucially, citations are *derived from what you supplied*. The model cannot cite a page it did not receive. If your retrieval returned the wrong three chunks, you get a confidently cited wrong answer — the citations prove the answer is consistent with the supplied documents, not that the supplied documents were the right ones. ## Grounded versus ungrounded With no `documents` (and no tool results), the same endpoint behaves like any other chat completion: the model answers from parametric knowledge and the citations list is empty or absent. So "grounded" here is not a mode flag you switch on; it is a consequence of having handed the call something to ground against. This is worth saying explicitly in an interview, because candidates often assume there is a `grounded=true` switch. ## Common mistakes - **Duplicating the documents into the user message.** Then you pay for the tokens twice and the model sees the same material in two shapes. Put it in `documents` only. - **Nesting structure inside `data`.** Keep it flat and stringly-typed; deep JSON is not what the field is for, and the model reads it worse. - **Random or per-request ids.** If the id is a UUID you generate at request time, the citation tells you nothing you can look up afterwards. Use the id of the thing in your store. - **Sending fifty documents because the context window allows it.** Every document is input tokens you are billed for on every turn, and dilution hurts answer quality. Retrieve, rerank, then send a short list. - **Assuming citations are mandatory.** The model will leave connective and general-knowledge text uncited, by design. ## Operational notes Document tokens dominate the cost of a grounded call, so cache or trim aggressively on multi-turn conversations rather than re-sending the whole document set on every turn. And treat retrieved documents as untrusted input: text you pulled from a wiki or a customer ticket can contain instructions aimed at the model, and putting it in `documents` rather than in the user turn does not sanitise it.

  • What do you get back if the question cannot be answered from the documents you passed?
    The model does not silently refuse. It will typically answer from general knowledge, or hedge, and those spans come back uncited because nothing in your documents supports them. If your product needs a hard refusal instead, you have to ask for it in the system message and then enforce it yourself by checking whether the claim-bearing spans carry citations.
  • Do the ids you attach to documents have to be meaningful?
    They are opaque to the model but echoed back in citations, so they are the only join key between an answer span and your store. Use the chunk or record id you can actually look up. If you pass bare strings instead of objects you give up that join and are left matching by position or by comparing text, which is brittle.
  • Should retrieved documents be treated as trusted content once they are in the documents array?
    No. Anything you retrieved from a wiki, a ticket, or a crawled page is untrusted input, and placing it in a dedicated parameter rather than the user turn does not neutralise embedded instructions. Constrain what the model is allowed to do in the system message, and never let a grounded answer trigger a privileged action without a separate check.

It is like handing a witness a folder of numbered exhibits instead of telling them the story: the answer comes back with exhibit numbers attached to each claim.

saying these in an interview costs you the question

  • Thinks there is a grounded=true flag to switch on
  • Pastes retrieved text into the user message and the documents array
  • Assumes citations prove the retrieved documents were correct
  • Uses random per-request ids that map to nothing
  • Sends every retrieved chunk because the context window allows it

context

open as a page

In Cohere's v2 Chat API, how do you map a citation back to text and source?

level: middleimportance: must knowfreq 52%

basics

~20 s

Each citation carries start and end character offsets into the generated text plus the sources that support that span. Slice the raw answer with text[start:end] and join each source to the document id you supplied. Offsets index the untouched text, so apply them before any escaping or trimming.

open as a page

In Cohere's Chat API, what do citation_options FAST, ACCURATE and OFF do?

level: middleimportance: should knowfreq 38%

basics

~20 s

citation_options.mode selects how much work the model spends on attribution. ACCURATE produces higher-quality citations at extra latency, FAST produces them more cheaply and quickly with somewhat coarser attribution, and OFF suppresses citations entirely so you get plain grounded text.

open as a page

A Cohere grounded answer returns spans with no citations — what does that mean?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Uncited spans are text the model did not attribute to any document you supplied — connective phrasing, framing, or claims drawn from parametric knowledge. Citations are never guaranteed to cover the whole answer, so treat an uncited claim as unsupported rather than as a bug.

open as a page

When is Cohere's in-API grounded generation the wrong choice for a RAG service?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

When attribution has to outlive the vendor. Cohere's documents-and-citations contract is specific to its chat API, so a service that must run the same grounding across several model providers, or needs citation rules the API does not expose, is better served by attribution it owns.

open as a page