What does a Cohere v2/chat request require, and what shape is the response?
answer
- Two required fields only
- Roles live in one array
- System prompt is a message
- content is a list of blocks
- finish_reason plus usage on the way back
basics
~20 sA Cohere v2/chat call needs a model ID and a messages array of role/content objects, authenticated with a bearer API key. The response carries an assistant message whose content is a list of typed blocks, plus finish_reason and usage.
solid answer
~40 sChat on Cohere is `POST https://api.cohere.com/v2/chat` with `Authorization: Bearer <COHERE_API_KEY>`. Only two body fields are required: `model` (a Command model ID) and `messages`, an array of objects with a `role` of `system`, `user`, `assistant` or `tool` and a `content` value that is either a plain string or an array of typed content blocks such as `{"type": "text", "text": "..."}`. `max_tokens` is optional, unlike some other vendors. The system prompt is just the first message with `role: "system"` — there is no top-level `system` or `preamble` field in v2. The response returns an `id`, a `message` object whose `content` is an array of blocks (read `message.content[0].text` for plain answers), a `finish_reason` such as `COMPLETE`, `MAX_TOKENS`, `STOP_SEQUENCE` or `TOOL_CALL`, and a `usage` object with token counts.
code
python · 15 linesimport cohere
co = cohere.ClientV2(api_key="COHERE_API_KEY")
res = co.chat(
model="command-a-03-2025",
messages=[
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What is a vector index?"},
],
max_tokens=200,
)
print(res.message.content[0].text)
print(res.finish_reason)go deeper
Be able to write the call from memory: model plus a messages array of role/content objects, a bearer API key, and reading the answer out of the content blocks rather than a flat text field.
Explain why max_tokens is optional here, why the system prompt is a message rather than a parameter, and how you branch on each finish_reason value instead of assuming the turn completed.
Show the operational wrapper: retry policy per status code, logging usage and the generation id per call, defensive parsing of the content array, and a history-trimming strategy for a stateless API.
Own the abstraction boundary. Argue for an internal chat interface that hides vendor request shapes so Cohere, Anthropic and OpenAI-shaped backends stay swappable, and decide where token accounting and transcript retention live across the platform.
## The endpoint and auth Cohere's chat surface for the Command family is a single HTTP endpoint: `POST https://api.cohere.com/v2/chat`. Authentication is a bearer token — `Authorization: Bearer <COHERE_API_KEY>` — with no per-request signing, no account/project header, and no separate token exchange. The Python SDK wraps this as `cohere.ClientV2(api_key=...)` and its `chat()` / `chat_stream()` methods; there is an async twin, `cohere.AsyncClientV2`. The v2 client is the one to reach for: v1's `cohere.Client` speaks the older request shape described at the end of this explanation. ## The request body Only two fields are mandatory: - **`model`** — a Command model identifier string. Cohere IDs normally carry a date suffix (for example `command-a-03-2025`), and there are undated aliases that track the newest snapshot. - **`messages`** — the conversation, as an ordered array. Everything else is optional. Notably `max_tokens` is **not** required; omit it and the model generates until it stops naturally or hits the model's own output limit. That differs from vendors whose APIs reject a request without an output cap, and it is a common porting surprise in both directions: an un-capped Cohere call can run longer and cost more than you expected. ## The messages array Each entry is an object with `role` and `content`. The four roles are: - **`system`** — instructions/persona. In v2 this is simply a message; there is no top-level `system` string (Anthropic's shape) and no `preamble` field (Cohere v1's shape). Convention is to put it first. - **`user`** — input from the end user. - **`assistant`** — a previous model turn, replayed back so the model has conversational history. The API is stateless: you resend the whole thread every call. - **`tool`** — the result of a tool the model asked for, correlated by a tool call id. `content` accepts two forms. The simple form is a plain string. The structured form is an array of typed blocks, the basic one being `{"type": "text", "text": "..."}`. Blocks exist so a single message can carry more than prose — the array form is what multimodal and structured inputs plug into. For ordinary text chat, the string form is fine and is what most examples use. ## The response body A non-streaming response has three parts you will actually read: - **`message`** — the assistant turn, with `role: "assistant"` and a `content` array of blocks. The text answer is `message.content[0].text` in the common single-block case; treat it as a list rather than assuming one element. When the model decides to call a tool, this object carries the tool-call structures instead of (or alongside) text. - **`finish_reason`** — why generation stopped. `COMPLETE` is a normal end of turn; `MAX_TOKENS` means output was truncated at your cap; `STOP_SEQUENCE` means one of your `stop_sequences` fired; `TOOL_CALL` means the model wants a tool executed and is waiting for you to append a `tool` message and call again; `ERROR` signals a failure mid-generation. - **`usage`** — token accounting, reported both as billed units and as raw token counts for input and output. Log it per request: it is the only reliable basis for cost attribution, since your own tokenizer estimate will not match what the service counted. There is also an `id` for the generation, which is worth logging alongside your own request id when you open a support conversation or debug a bad answer. ## Error and status handling Ordinary HTTP semantics apply. `401` means a bad or missing key; `400` means a malformed body (a wrong field name, an invalid role, a model ID that does not exist for your account); `429` means you are over the rate limit and should back off; `5xx` means retry with jitter. Because the API is stateless, retries are safe as long as you resend the identical `messages` array — but note that a retry is billed again, so cap your attempts. ## What changed from v1 If you meet older Cohere code, v1's `POST /v1/chat` took a single `message` string plus a `chat_history` array and a separate `preamble` for the system prompt. v2 collapses all three into one `messages` array with roles, which is why porting is mostly mechanical: `preamble` becomes a `system` message, `chat_history` entries become `user`/`assistant` messages, and the current turn becomes the last `user` message. Response accessors changed too — v1 returned a flat `text` field, v2 returns content blocks. ## Practical checklist Send `model` and `messages`; put the system prompt in the array; read `message.content[]` defensively; branch on `finish_reason` rather than assuming success; log `usage`; and keep the whole thread client-side, because nothing is stored for you between calls.
- How would you port a v1 Cohere chat call to v2?Fold the three v1 inputs into one array. The `preamble` string becomes a `{"role": "system"}` message, each `chat_history` entry becomes a `user` or `assistant` message, and the single `message` string becomes the final `user` message. On the response side, v1's flat `text` accessor becomes `message.content[0].text`, and you now branch on `finish_reason` for truncation and tool calls.
- The API is stateless — what does that mean for a multi-turn chat feature?You own the transcript. Every call must resend the full `messages` array including prior assistant turns, so your service stores the conversation and decides what to trim when it grows. That also means input tokens grow with the thread, so cost per turn rises unless you truncate, summarise older turns, or cap history depth.
- What do you do when finish_reason comes back TOOL_CALL?Stop treating the turn as final. The model has emitted a tool request rather than a finished answer: you execute the requested tool, append the model's assistant message and a `tool` message carrying the result to the same `messages` array, and call the endpoint again so the model can continue with the tool output in hand.
saying these in an interview costs you the question
- Thinking max_tokens is required on Cohere's chat endpoint
- Looking for a top-level system or preamble field in v2
- Reading response.text instead of message.content blocks
- Assuming Cohere stores conversation state between calls
- Treating any response as final without checking finish_reason