skip to content

Messages API fundamentals

The one endpoint for Claude: a model, a required max_tokens, and an alternating messages array of content blocks, answered with a stop_reason and a usage object. Interviewers like the specifics — the system prompt is a top-level parameter rather than a message role.

on this pageshow

questions

5

In Anthropic's Messages API, how do you send a system prompt to Claude?

level: juniorimportance: must knowfreq 78%

answer

  1. Not a role in the array
  2. Sits beside model and max_tokens
  3. String, or a list of text blocks
  4. Rendered before the conversation turns

basics

~20 s

The Messages API takes the system prompt as a top-level system parameter on the request body, alongside model, max_tokens and messages. It is not a message inside the messages array; that array carries only user and assistant turns.

solid answer

~40 s

Anthropic's `POST /v1/messages` puts the system prompt in a **top-level `system` field**, sibling to `model`, `max_tokens` and `messages` — not as an entry in the `messages` array. That array is for the conversation itself, whose turns are `user` and `assistant`. `system` accepts either a plain string or an array of text content blocks; the block form exists so you can attach per-block metadata such as cache breakpoints. This differs from OpenAI-shaped chat APIs, where the system prompt is simply the first message, so porting code between the two means moving the instruction out of the array rather than renaming a field. The server renders the request as tools, then system, then messages, which is also why a stable system prompt sits in the most reusable position of the prefix.

code

json · 8 lines
json
{
  "model": "claude-opus-4-6",
  "max_tokens": 1024,
  "system": "You are a terse SQL assistant. Answer with SQL only.",
  "messages": [
    {"role": "user", "content": "How many rows are in the orders table?"}
  ]
}

go deeper

for a junior

Remember the concrete shape: model, max_tokens, messages and system are all top-level keys of the request body, and the messages array holds only user and assistant turns.

for a middle

Be ready to explain that system accepts a string or an array of text blocks, and that the server renders tools, then system, then messages — which is why a constant system prompt sits in the most reusable prefix position.

for a senior

Show that you design the system prompt for stability: constant operator instructions up front, volatile per-request data pushed later, and awareness that it is re-sent and re-billed on every stateless call.

for a principal

Own the tradeoff between a large shared system prompt and per-request context: what belongs in the operator channel versus retrieved into a user turn, and how that choice drives both cost and prompt-injection exposure across tenants.

## The shape of the request Every call to Claude goes to one endpoint, `POST https://api.anthropic.com/v1/messages`, with the headers `x-api-key`, `anthropic-version: 2023-06-01` and `content-type: application/json`. The JSON body has three required fields — `model`, `max_tokens`, `messages` — and a set of optional top-level fields. **`system` is one of those top-level fields.** It is not a role, not a header, and not a special content block. ``` { "model": "...", "max_tokens": 1024, "system": "You are a terse SQL assistant.", "messages": [{"role": "user", "content": "count the orders table"}] } ``` ## Why it is not a message role The `messages` array models a dialogue: it alternates between `user` and `assistant` turns and represents things that were *said*. The system prompt is not a turn — it is operator configuration that applies to the whole exchange. Anthropic made that distinction structural rather than conventional. The practical consequences are worth knowing: - You cannot accidentally lose the system prompt by trimming the front of the history. Truncating old turns for context reasons touches `messages` only; `system` is untouched. - There is exactly one system prompt per request, so there is no ambiguity about which of several system messages wins. - Because the server renders `tools` → `system` → `messages` when building the prompt, a system prompt that never changes sits in a stable prefix position, which is what makes it reusable across requests. ## String or content blocks `system` accepts two forms. The simple form is a plain string. The structured form is an array of text content blocks: ``` "system": [ {"type": "text", "text": "You are a terse SQL assistant."}, {"type": "text", "text": "<50KB of schema documentation>"} ] ``` The array form exists so you can attach per-block metadata — most commonly a cache breakpoint on the last stable block — and so you can compose the prompt from independently-managed pieces. Only text blocks are valid here; images and documents belong in a user message's content. ## Porting from an OpenAI-shaped API In OpenAI-shaped chat APIs the system prompt is the first element of the messages list. Translating such code to Claude is not a field rename: you lift that entry out of the array entirely and assign its text to the top-level `system` parameter. If you leave it in the array as `{"role": "system", ...}`, the classic result is a `400 invalid_request_error` complaining about roles — historically the array accepted only `user` and `assistant`. Note one 2026 refinement: some current Claude models (the Opus 5 / Opus 4.8 class) additionally accept a mid-conversation `{"role": "system", "content": ...}` entry *inside* `messages`, with placement rules, as a way to inject an operator instruction later in a long conversation without disturbing the cached prefix. That is an addition for mid-conversation steering, not a change to where the primary system prompt goes — it still belongs in the top-level parameter, and support is model-dependent, so do not write it as the default. ## What belongs in it Role, tone, output-format rules, domain constraints, and any long reference material that is constant across requests. What does not belong: the user's actual question (that is a `user` message), and per-request volatile data such as a timestamp or a request id, which is better placed after the stable material so it does not destabilise the prefix. ## Failure modes to recognise - Sending `system` as a header or as `messages[0]` — rejected or ignored, and the model behaves as though it has no instructions. - Sending `system` as an array containing message objects (`{"role": ..., "content": ...}`) instead of text blocks — a validation error; the array form takes content blocks. - Rebuilding the system string on every request with an interpolated clock or UUID — it still works, but you have made a constant prefix volatile for no benefit. - Assuming the system prompt is free. It is input, it is re-sent on every request, and it is counted in the request's input token usage like everything else.

  • What happens if you put a system-role entry inside the messages array instead?
    Historically the array accepted only `user` and `assistant`, so a `system` entry came back as a `400 invalid_request_error` about roles. Current Opus 5 / Opus 4.8-class models added a narrow exception: a mid-conversation `{"role": "system"}` message may appear inside `messages` to inject an operator instruction late in a conversation, subject to placement rules. It is a steering mechanism, not the home of your primary system prompt.
  • Why would you send system as an array of text blocks rather than one string?
    Two reasons. First, per-block metadata: cache breakpoints are attached to a specific block, so the array form lets you mark where the stable prefix ends. Second, composition: a platform prompt, a tenant prompt and a schema dump can be assembled and versioned independently and concatenated at request time without string surgery.
  • Does the system prompt cost tokens on every request?
    Yes. The API is stateless, so the system prompt is re-sent and re-read on every call and counted in that request's input tokens. A large constant system prompt therefore has a per-request cost proportional to its size — which is exactly why the API offers a way to mark stable prefix content so repeated reads are billed differently.

The messages array is the transcript of a conversation; the system parameter is the briefing note handed to the participant before the conversation starts.

saying these in an interview costs you the question

  • Says the system prompt is the first message with role system
  • Thinks system is an HTTP header on the request
  • Assumes system must be a string and blocks are invalid
  • Believes the system prompt persists server-side between calls
  • Puts images or documents into the system parameter

context

open as a page

What does the required max_tokens parameter cap on an Anthropic Messages API call?

level: middleimportance: must knowfreq 72%

basics

~20 s

max_tokens is a hard ceiling on the tokens Claude generates in that one response, including any thinking tokens. It is required on every Messages API request, and hitting it truncates the output mid-stream with stop_reason "max_tokens".

open as a page

How must the messages array be structured in an Anthropic Messages API request?

level: middleimportance: must knowfreq 68%

basics

~20 s

messages is a non-empty array of turns, each with a role of user or assistant, starting with user and alternating. Each turn's content is either a plain string or an array of typed content blocks. The endpoint is stateless, so you resend the full history every call.

open as a page

Which stop_reason values can an Anthropic Messages API response return?

level: seniorimportance: should knowfreq 52%

basics

~10 s

Common values are end_turn, max_tokens, stop_sequence and tool_use, plus pause_turn, refusal and model_context_window_exceeded. Read stop_reason before touching the content array, because some outcomes return HTTP 200 with little or no content.

open as a page

In the Anthropic Messages API usage object, what does input_tokens count?

level: seniorimportance: should knowfreq 45%

basics

~20 s

input_tokens counts everything the server read for that one request — system prompt, tool definitions and the entire resent message history — not just the newest user turn. output_tokens counts what the model generated in that response.

open as a page