skip to content

Cohere

Cohere is the enterprise-RAG-shaped vendor: Command models, Embed and Rerank are designed to be assembled into grounded search over your own documents. In interviews it usually surfaces through Rerank and through citation-backed generation.

on this pageshow

explore

questions

24

What does a Cohere v2/chat request require, and what shape is the response?

level: juniorimportance: must knowfreq 65%

answer

  1. Two required fields only
  2. Roles live in one array
  3. System prompt is a message
  4. content is a list of blocks
  5. finish_reason plus usage on the way back

basics

~20 s

A Cohere v2/chat call needs a model ID and a messages array of role/content objects, authenticated with a bearer API key. The response carries an assistant message whose content is a list of typed blocks, plus finish_reason and usage.

solid answer

~40 s

Chat on Cohere is `POST https://api.cohere.com/v2/chat` with `Authorization: Bearer <COHERE_API_KEY>`. Only two body fields are required: `model` (a Command model ID) and `messages`, an array of objects with a `role` of `system`, `user`, `assistant` or `tool` and a `content` value that is either a plain string or an array of typed content blocks such as `{"type": "text", "text": "..."}`. `max_tokens` is optional, unlike some other vendors. The system prompt is just the first message with `role: "system"` — there is no top-level `system` or `preamble` field in v2. The response returns an `id`, a `message` object whose `content` is an array of blocks (read `message.content[0].text` for plain answers), a `finish_reason` such as `COMPLETE`, `MAX_TOKENS`, `STOP_SEQUENCE` or `TOOL_CALL`, and a `usage` object with token counts.

code

python · 15 lines
python
import cohere

co = cohere.ClientV2(api_key="COHERE_API_KEY")

res = co.chat(
    model="command-a-03-2025",
    messages=[
        {"role": "system", "content": "Answer in one sentence."},
        {"role": "user", "content": "What is a vector index?"},
    ],
    max_tokens=200,
)

print(res.message.content[0].text)
print(res.finish_reason)

go deeper

for a junior

Be able to write the call from memory: model plus a messages array of role/content objects, a bearer API key, and reading the answer out of the content blocks rather than a flat text field.

for a middle

Explain why max_tokens is optional here, why the system prompt is a message rather than a parameter, and how you branch on each finish_reason value instead of assuming the turn completed.

for a senior

Show the operational wrapper: retry policy per status code, logging usage and the generation id per call, defensive parsing of the content array, and a history-trimming strategy for a stateless API.

for a principal

Own the abstraction boundary. Argue for an internal chat interface that hides vendor request shapes so Cohere, Anthropic and OpenAI-shaped backends stay swappable, and decide where token accounting and transcript retention live across the platform.

## The endpoint and auth Cohere's chat surface for the Command family is a single HTTP endpoint: `POST https://api.cohere.com/v2/chat`. Authentication is a bearer token — `Authorization: Bearer <COHERE_API_KEY>` — with no per-request signing, no account/project header, and no separate token exchange. The Python SDK wraps this as `cohere.ClientV2(api_key=...)` and its `chat()` / `chat_stream()` methods; there is an async twin, `cohere.AsyncClientV2`. The v2 client is the one to reach for: v1's `cohere.Client` speaks the older request shape described at the end of this explanation. ## The request body Only two fields are mandatory: - **`model`** — a Command model identifier string. Cohere IDs normally carry a date suffix (for example `command-a-03-2025`), and there are undated aliases that track the newest snapshot. - **`messages`** — the conversation, as an ordered array. Everything else is optional. Notably `max_tokens` is **not** required; omit it and the model generates until it stops naturally or hits the model's own output limit. That differs from vendors whose APIs reject a request without an output cap, and it is a common porting surprise in both directions: an un-capped Cohere call can run longer and cost more than you expected. ## The messages array Each entry is an object with `role` and `content`. The four roles are: - **`system`** — instructions/persona. In v2 this is simply a message; there is no top-level `system` string (Anthropic's shape) and no `preamble` field (Cohere v1's shape). Convention is to put it first. - **`user`** — input from the end user. - **`assistant`** — a previous model turn, replayed back so the model has conversational history. The API is stateless: you resend the whole thread every call. - **`tool`** — the result of a tool the model asked for, correlated by a tool call id. `content` accepts two forms. The simple form is a plain string. The structured form is an array of typed blocks, the basic one being `{"type": "text", "text": "..."}`. Blocks exist so a single message can carry more than prose — the array form is what multimodal and structured inputs plug into. For ordinary text chat, the string form is fine and is what most examples use. ## The response body A non-streaming response has three parts you will actually read: - **`message`** — the assistant turn, with `role: "assistant"` and a `content` array of blocks. The text answer is `message.content[0].text` in the common single-block case; treat it as a list rather than assuming one element. When the model decides to call a tool, this object carries the tool-call structures instead of (or alongside) text. - **`finish_reason`** — why generation stopped. `COMPLETE` is a normal end of turn; `MAX_TOKENS` means output was truncated at your cap; `STOP_SEQUENCE` means one of your `stop_sequences` fired; `TOOL_CALL` means the model wants a tool executed and is waiting for you to append a `tool` message and call again; `ERROR` signals a failure mid-generation. - **`usage`** — token accounting, reported both as billed units and as raw token counts for input and output. Log it per request: it is the only reliable basis for cost attribution, since your own tokenizer estimate will not match what the service counted. There is also an `id` for the generation, which is worth logging alongside your own request id when you open a support conversation or debug a bad answer. ## Error and status handling Ordinary HTTP semantics apply. `401` means a bad or missing key; `400` means a malformed body (a wrong field name, an invalid role, a model ID that does not exist for your account); `429` means you are over the rate limit and should back off; `5xx` means retry with jitter. Because the API is stateless, retries are safe as long as you resend the identical `messages` array — but note that a retry is billed again, so cap your attempts. ## What changed from v1 If you meet older Cohere code, v1's `POST /v1/chat` took a single `message` string plus a `chat_history` array and a separate `preamble` for the system prompt. v2 collapses all three into one `messages` array with roles, which is why porting is mostly mechanical: `preamble` becomes a `system` message, `chat_history` entries become `user`/`assistant` messages, and the current turn becomes the last `user` message. Response accessors changed too — v1 returned a flat `text` field, v2 returns content blocks. ## Practical checklist Send `model` and `messages`; put the system prompt in the array; read `message.content[]` defensively; branch on `finish_reason` rather than assuming success; log `usage`; and keep the whole thread client-side, because nothing is stored for you between calls.

  • How would you port a v1 Cohere chat call to v2?
    Fold the three v1 inputs into one array. The `preamble` string becomes a `{"role": "system"}` message, each `chat_history` entry becomes a `user` or `assistant` message, and the single `message` string becomes the final `user` message. On the response side, v1's flat `text` accessor becomes `message.content[0].text`, and you now branch on `finish_reason` for truncation and tool calls.
  • The API is stateless — what does that mean for a multi-turn chat feature?
    You own the transcript. Every call must resend the full `messages` array including prior assistant turns, so your service stores the conversation and decides what to trim when it grows. That also means input tokens grow with the thread, so cost per turn rises unless you truncate, summarise older turns, or cap history depth.
  • What do you do when finish_reason comes back TOOL_CALL?
    Stop treating the turn as final. The model has emitted a tool request rather than a finished answer: you execute the requested tool, append the model's assistant message and a `tool` message carrying the result to the same `messages` array, and call the endpoint again so the model can continue with the tool output in hand.

saying these in an interview costs you the question

  • Thinking max_tokens is required on Cohere's chat endpoint
  • Looking for a top-level system or preamble field in v2
  • Reading response.text instead of message.content blocks
  • Assuming Cohere stores conversation state between calls
  • Treating any response as final without checking finish_reason

context

open as a page

How do you call Cohere's v2/embed endpoint and read the vectors from the response?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Send texts plus a model and an input_type to POST /v2/embed. The response groups vectors by type: a float request returns them under embeddings.float, one vector per input text, in the order you sent them.

open as a page

In Cohere's v2 Chat API, how do you pass source documents so the reply carries citations?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Send a top-level documents array alongside messages in the v2 chat call. Each entry is a plain string or an object with an id and a data map. The reply's message.citations then link answer spans to those documents.

open as a page

What does a Cohere v2/rerank request contain, and what does the API return?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A Cohere v2/rerank call sends a model name, one query string, and a documents list of plain strings, plus an optional top_n. It returns a results array ordered by relevance, each entry carrying the document's original index and a relevance_score.

open as a page

Which models make up Cohere's Command family, and why pin a dated model ID?

level: middleimportance: must knowfreq 55%

basics

~20 s

Command A is Cohere's current flagship chat model with a 256K-token context, above the earlier Command R+ and Command R at 128K and the small Command R7B. IDs carry a date suffix; undated aliases follow the newest snapshot, so production should pin the dated ID.

open as a page

In Cohere's Embed API, why must input_type differ for documents and queries?

level: middleimportance: must knowfreq 72%

basics

~20 s

Cohere's Embed models are asymmetric: input_type tells the model what the text is for, and it conditions the vector accordingly. Corpus chunks go in as search_document, user queries as search_query, so a short question lands near the long passage that answers it.

open as a page

In Cohere's v2 Chat API, how do you map a citation back to text and source?

level: middleimportance: must knowfreq 52%

basics

~20 s

Each citation carries start and end character offsets into the generated text plus the sources that support that span. Slice the raw answer with text[start:end] and join each source to the document id you supplied. Offsets index the untouched text, so apply them before any escaping or trimming.

open as a page

Why does each Cohere rerank result carry an index instead of the document text?

level: middleimportance: must knowfreq 55%

basics

~20 s

Rerank returns a reordered list, so each result must say which input it came from. The index is the document's position in the request's documents array, which is how you join back to your own records without paying to send the text back.

open as a page

In Cohere's v2/chat API, what do the p and k parameters control?

level: middleimportance: should knowfreq 45%

basics

~20 s

Cohere spells nucleus and top-k sampling as p and k, not top_p and top_k. Both narrow the candidate pool the next token is drawn from; k defaults to 0, which disables top-k. Temperature is separate and defaults to 0.3.

open as a page

What events does Cohere's v2/chat stream emit when stream is true?

level: middleimportance: should knowfreq 50%

basics

~10 s

Cohere streams typed events rather than uniform chunks: message-start opens the turn, content-start/content-delta/content-end wrap each content block, and message-end closes it with finish_reason and usage. You accumulate the text from content-delta events.

open as a page

What do Cohere's embedding_types int8 and binary return, and how much smaller are they?

level: middleimportance: should knowfreq 48%

basics

~20 s

int8 returns one signed byte per dimension, about four times smaller than float32. binary packs one bit per dimension into bytes, so a 1024-dimension vector becomes 128 values — roughly thirty-two times smaller. You can request several types in one call.

open as a page

In Cohere's Chat API, what do citation_options FAST, ACCURATE and OFF do?

level: middleimportance: should knowfreq 38%

basics

~20 s

citation_options.mode selects how much work the model spends on attribution. ACCURATE produces higher-quality citations at extra latency, FAST produces them more cheaply and quickly with somewhat coarser attribution, and OFF suppresses citations entirely so you get plain grounded text.

open as a page

How does Cohere's v2/rerank handle a document longer than max_tokens_per_doc?

level: middleimportance: should knowfreq 42%

basics

~20 s

It truncates the document to that token budget (4096 by default) and scores only the kept portion; the overflow is invisible to the model. The v1 API instead split long documents into chunks and kept the best chunk's score, which v2 dropped.

open as a page

Which Cohere rerank model fits a corpus mixing English, French and Japanese?

level: middleimportance: should knowfreq 32%

basics

~20 s

Use a multilingual rerank model such as rerank-v3.5, which covers 100-plus languages and scores a query in one language against documents in another. You send one mixed candidate list to one call — no per-language routing and no translation step.

open as a page

Your data cannot leave your VPC — how do you run Cohere Command models?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Cohere Command models ship beyond Cohere's own API: through hyperscaler catalogues such as Amazon Bedrock, Amazon SageMaker and Azure AI Foundry, and as container images deployed inside your own VPC or on-prem under a commercial licence. Auth, model IDs and version currency all change with the channel.

open as a page

How do you embed a large document corpus with Cohere's Embed API without hitting limits?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Chunk the corpus, send batches within the documented per-call text limit of 96, keep input_type constant, set truncate deliberately, and drive concurrency with exponential backoff on 429. For very large corpora, use the asynchronous Embed Jobs API over an uploaded dataset instead.

open as a page

A Cohere grounded answer returns spans with no citations — what does that mean?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Uncited spans are text the model did not attribute to any document you supplied — connective phrasing, framing, or claims drawn from parametric knowledge. Citations are never guaranteed to cover the whole answer, so treat an uncited claim as unsupported rather than as a bug.

open as a page

How do you keep a Cohere Rerank hop from breaking a production search request?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Treat Rerank as an optional enhancement, not a dependency: give it a timeout smaller than the remaining request budget and, on timeout, 429 or 5xx, serve the vector-recall order instead of failing. Log every degradation so the silent quality loss stays visible.

open as a page

How is a Cohere Rerank call billed, and does a smaller top_n make it cheaper?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Rerank is billed in search units, where one unit covers a single query against up to 100 documents. A smaller top_n changes nothing: every document you send is still scored, so the cost lever is the candidate count, not the number of results returned.

open as a page

How do you choose between Cohere's float and binary embeddings at 100M scale?

level: principalimportance: should knowfreq 30%

basics

~20 s

Do not choose one. Request both types in a single embed call — the token bill is the same — then serve a compressed index for first-pass recall and rescore the top candidates with the higher-fidelity vectors. Validate the recall loss on your own eval set.

open as a page

How do you embed images with Cohere's Embed API, and what makes them comparable to text?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Images are passed to Cohere's embed endpoint as base64-encoded data URIs rather than as remote URLs or file uploads. Because one multimodal model produces both text and image vectors in a single shared space, a text query can retrieve images directly.

open as a page

How do you judge whether Command's multilingual coverage fits a new market?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Treat a vendor's language list as a starting hypothesis, not a guarantee. Command A is marketed at 23 languages and the earlier Command R line at ten key business languages, but only a task-specific eval with native reviewers tells you whether quality, safety behaviour and token cost hold in that market.

open as a page

When is Cohere's in-API grounded generation the wrong choice for a RAG service?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

When attribution has to outlive the vendor. Cohere's documents-and-citations contract is specific to its chat API, so a service that must run the same grounding across several model providers, or needs citation rules the API does not expose, is better served by attribution it owns.

open as a page