How must the messages array be structured in an Anthropic Messages API request?
answer
- Two roles only, in turn
- First entry is the human
- String is sugar for one block
- No conversation id anywhere
- Echo the assistant's blocks back
basics
~20 smessages is a non-empty array of turns, each with a role of user or assistant, starting with user and alternating. Each turn's content is either a plain string or an array of typed content blocks. The endpoint is stateless, so you resend the full history every call.
solid answer
~40 s`messages` carries the conversation as an ordered array of `{role, content}` objects. Roles are `user` and `assistant`; the array must be non-empty, starts with a `user` turn, and alternates — a mismatch returns `400 invalid_request_error` with a message about roles alternating. `content` has two forms: a convenience string, or an array of typed **content blocks** (`text`, `image`, `document`, `tool_use`, `tool_result`, and so on), which is what the API actually works in — the string form is sugar for a single text block. The endpoint is **stateless**: there is no conversation id, so every request resends the whole history. When you append the previous assistant turn, append its entire `content` array verbatim rather than just the extracted text, or you silently drop the non-text blocks the next turn depends on.
code
json · 9 lines{
"model": "claude-opus-4-6",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "What is a B-tree?"},
{"role": "assistant", "content": [{"type": "text", "text": "A balanced search tree."}]},
{"role": "user", "content": [{"type": "text", "text": "And a B+ tree?"}]}
]
}go deeper
Recall the basics: an array of user and assistant turns, starting with user, alternating, each with content that can simply be a string.
Explain that content is really a list of typed blocks with the string as sugar, that the endpoint is stateless so the client resends the whole history, and what the role-alternation error looks like.
Show the operational judgment: append the assistant's full content array rather than extracted text, and trim or summarise history in a way that preserves alternation and never splits a paired exchange.
Own the conversation-state strategy — where history is persisted, how it is compacted as it approaches the context window, and the cost curve of re-sending an ever-growing transcript across a fleet of sessions.
## The array `messages` is one of the three required fields of `POST /v1/messages`. It is an ordered list of turns: ``` "messages": [ {"role": "user", "content": "What is a B-tree?"}, {"role": "assistant", "content": "A balanced tree structure..."}, {"role": "user", "content": "How does it differ from a B+ tree?"} ] ``` Rules that trip people up: - **Non-empty.** An empty array is a validation error. - **Starts with `user`.** The conversation begins with something for the model to respond to. - **Alternates.** Consecutive turns of the same role are rejected; the canonical error text is `messages: roles must alternate between "user" and "assistant"`, returned as `400 invalid_request_error`. - **Only conversational roles.** The operator instruction is not a turn — it belongs in the top-level `system` parameter. - **Order is meaning.** The array *is* the history; there is no server-side ordering or deduplication. ## Content: string or blocks Every turn's `content` may be a plain string or an array of content blocks. These are equivalent: ``` {"role": "user", "content": "Hello"} {"role": "user", "content": [{"type": "text", "text": "Hello"}]} ``` The block array is the real model; the string is shorthand for a single `text` block. Blocks are typed and heterogeneous — a single user turn can mix text with an image, a document, or the results of tools the model asked for. The response side is always blocks: `content` on a response is an array even when it holds one piece of text, which is why `content[0].text` (rather than a bare `content` string) is the idiom for reading an answer. The block design is what lets one endpoint carry text, vision, documents, reasoning and tool traffic without versioning the request shape each time a modality is added. ## Statelessness is the big one There is no session, thread or conversation id on this endpoint. Each request is self-contained: the server has no memory of the previous call. Multi-turn chat is implemented entirely by the client resending the accumulated array each time. The loop is: 1. Append the user's new message to your local array. 2. POST the whole array. 3. Append the assistant's response to your local array. 4. Repeat. Two consequences follow. First, **cost and latency grow with the conversation** — the whole history is re-read on every call, so a long chat pays for its own history repeatedly. Second, **you own truncation**: when the history approaches the model's context window, dropping or summarising old turns is your job, and doing it naively can break role alternation or orphan a tool block. ## Append the assistant's content, not its text The most common bug in real integrations: extracting the assistant's text with something like `response.content[0].text`, appending *that string* as the assistant turn, and discarding the rest of the array. If the response held anything besides that first text block, it is now gone from the history, and the next turn is missing state the model expected to still be there. The correct move is to append the response's `content` array as-is: ``` messages.append({"role": "assistant", "content": response.content}) ``` ## Trimming and summarising safely When you do have to shrink the history: - Keep the first turn a `user` turn after trimming, and keep alternation intact. - Never cut between a turn that references work and the turn that answers it — split pairs produce validation errors or a confused model. - Prefer summarising a prefix into a single `user` turn over deleting arbitrary middles, so the remaining array is still a coherent dialogue. ## Quick mental checklist A valid `messages` array: non-empty; first role `user`; roles alternate; every `content` is a string or a list of typed blocks; no operator instructions smuggled in as turns; the full history present because nothing is stored server-side.
- What error do you get if two user messages appear back to back?A `400 invalid_request_error` whose message says roles must alternate between `user` and `assistant`. It is a validation failure, not a soft warning, so any code path that appends a turn without checking the previous role — retry logic and history-trimming code are the usual culprits — will fail loudly at request time.
- Why append response.content rather than the extracted text string?Because a response's `content` is an array that may hold more than one block. Appending only the first block's text discards everything else and rewrites the history the model will read next turn. Passing the array back verbatim keeps the record faithful; extracting text is for display, not for state.
- How do you keep a long conversation inside the context window?You manage it client-side, because nothing is stored server-side. Drop or summarise the oldest turns while preserving alternation and a leading `user` turn, keep any paired turns together, and measure rather than guess — the token count of the history is what determines when you must act.
The API behaves like a stateless HTTP handler: nothing is remembered between calls, so the client carries the whole transcript in every request the way a cookie-less client carries all its state in the payload.
saying these in an interview costs you the question
- Thinks the API stores the conversation between calls
- Sends two user turns in a row and expects them merged
- Believes content must always be a plain string
- Appends only the assistant's text and drops other blocks
- Starts the array with an assistant turn