In Cohere's Chat API, what do citation_options FAST, ACCURATE and OFF do?
answer
- a dial on attribution, not grounding
- uppercase FAST, ACCURATE, OFF
- latency traded against span fidelity
- OFF still reads your documents
- default varies by model, set it
basics
~20 scitation_options.mode selects how much work the model spends on attribution. ACCURATE produces higher-quality citations at extra latency, FAST produces them more cheaply and quickly with somewhat coarser attribution, and OFF suppresses citations entirely so you get plain grounded text.
solid answer
~50 s`citation_options` is a request object on Cohere's chat call whose `mode` takes the uppercase values `FAST`, `ACCURATE` or `OFF`. It is a quality-versus-latency dial on the attribution step, not on the answer: the model still reads the documents you supplied in every mode. `ACCURATE` spends more work deciding which span maps to which source and gives you the most trustworthy spans; `FAST` returns citations sooner with somewhat looser attribution; `OFF` skips citation generation altogether, which is what you want for a grounded call whose output you are not going to attribute in the UI. The default varies by model, so set it explicitly rather than inheriting whatever a new model id ships with. Note that turning citations off does not make the call cheap — you are still billed for every document token in the input.
code
python · 13 linesimport cohere
co = cohere.ClientV2(api_key="<key>")
res = co.chat(
model="command-a-03-2025",
messages=[{"role": "user", "content": "Summarise the refund policy in one sentence."}],
documents=[{"id": "policy-1", "data": {"text": "Refunds are accepted within 30 days of delivery."}}],
citation_options={"mode": "FAST"},
)
print(res.message.content[0].text)
print(res.message.citations)go deeper
Know the three uppercase values and that they trade attribution quality against latency, with OFF simply returning no citations.
Explain that the dial affects the attribution step and not whether the documents are used, and that the default has varied by model so you should set it explicitly.
Justify the choice per surface with numbers — latency at p95 versus the share of claims correctly attributed on your own evaluation set — and keep the streaming UI painting text before citations land.
Own the product policy: which surfaces owe users verifiable provenance, which merely display it, and where a second on-demand high-fidelity call is a better trade than slowing everyone down.
## What the parameter controls When you pass documents to Cohere's chat endpoint, two things happen: the model answers using them, and the API attributes spans of that answer to them. `citation_options.mode` governs the second half only. Choosing a cheaper mode does not make the answer less grounded — it makes the *bookkeeping about* the grounding cheaper. The values are uppercase strings: - **`ACCURATE`** — the highest-fidelity attribution. Spans line up tightly with the claims they cover and the source set per span is the most reliable. It costs the most time. - **`FAST`** — attribution produced with less work. You still get spans and sources; they are coarser, and the odd claim may be attributed loosely or missed. Noticeably lower added latency. - **`OFF`** — no citation generation at all. The response is grounded prose with an empty or absent citations list. Because the default has differed across model generations, treating it as "whatever the model does" is how a product quietly changes behaviour on a model upgrade. Set it in the request. ## Choosing a mode The question to ask is *what does the citation do for the user*. **Choose ACCURATE when attribution is part of the product's contract.** A compliance answer, a legal or medical assistant, an internal knowledge tool whose whole value proposition is "click through to the source" — here a citation that points at the wrong paragraph is worse than no citation, because it manufactures unearned confidence. The extra latency is the price of a defensible answer, and these surfaces are usually not the latency-critical ones anyway. **Choose FAST when citations are a supporting affordance in an interactive chat.** The user is reading streamed text; a slightly loose highlight is a minor cosmetic issue, while a pause before the first token is a felt defect. FAST is the sane default for consumer-shaped chat over a knowledge base. **Choose OFF when you are not rendering attribution.** Classification, summarisation into a downstream pipeline, an extraction step whose output is parsed by code, a background job filling a cache — none of these consume citations, so paying latency to compute them is waste. OFF is also the right answer when a first pass in a multi-step chain feeds a second call that will do the citing. ## What it does not change Three things are worth stating plainly because candidates get them wrong: 1. **It is not a grounding switch.** `OFF` does not stop the model using your documents. If you want an ungrounded answer, do not send documents. 2. **It is not a cost lever of any size.** The dominant cost of a grounded call is the document tokens in the input, which you pay on every turn regardless of mode. Citation mode buys you latency, not a materially smaller bill. 3. **It does not affect answer quality directly.** The prose the model writes is driven by the documents and the prompt. Mode changes how well that prose is annotated afterwards. ## Interaction with streaming Under streaming, citations arrive as their own events interleaved with the text. `ACCURATE` therefore shows up as citation events landing later relative to the text they cover, and possibly as a longer tail after the last content delta. If your UI blocks rendering until citations are complete, ACCURATE will look much slower than it is — render the text as it streams and decorate it with highlights when the citation events arrive. ## Evaluating the choice rather than guessing it The honest senior answer is that the FAST/ACCURATE decision should be measured on your own corpus, not assumed. Build a small evaluation set of questions with known supporting chunks, run both modes, and score two things: added latency at p50 and p95, and the fraction of claim-bearing sentences that receive a citation resolving to the correct chunk. If FAST loses nothing measurable on your documents, take the latency. If it drops attribution on the exact class of question your users ask most, that is the answer too. Making that call from a benchmark you ran, rather than from the parameter's name, is what separates a senior response here. ## A practical pattern A workable configuration for a chat product is FAST in the interactive path, with an on-demand "verify sources" action that re-runs the same question and documents at ACCURATE. Users get fast answers, and the small minority who care about provenance get precise spans on request — at the cost of a second call for that minority only.
- Does setting the mode to OFF save you money on a grounded call?Barely. The dominant cost is the document tokens you send in the input, and you pay those on every turn whatever the mode. OFF removes the attribution work, so it buys latency and a simpler response, not a materially smaller bill. If you want the call to be cheaper, send fewer or shorter documents — rerank before you send rather than after.
- How would you decide between FAST and ACCURATE for your own product?Measure it. Build an evaluation set of questions whose correct supporting chunks you know, run both modes, and compare added latency at p50 and p95 against the share of claim-bearing sentences that receive a citation resolving to the right chunk. Take FAST if the attribution loss is not visible on your corpus; take ACCURATE where a wrong pointer is worse than a slow answer.
- Your UI feels sluggish after switching to ACCURATE while streaming. What is likely wrong?Probably that the UI waits for citations before painting anything. Citations stream as their own events and, in the higher-fidelity mode, land later relative to the text they cover. Render the content deltas immediately and attach highlights when the citation events arrive, so the perceived latency is the time to first token rather than the time to full attribution.
saying these in an interview costs you the question
- Thinks OFF stops the model from using the documents
- Believes citation mode is a meaningful cost saving
- Assumes the default mode is the same on every model
- Treats ACCURATE as improving answer quality, not attribution
- Blocks the UI until citations arrive when streaming