skip to content

Anthropic

Anthropic's Claude models behind a single Messages endpoint. Interviews here focus on the content-block request shape, tool use, the streaming event sequence, and prompt caching — the levers that keep a long-context agent affordable.

on this pageshow

explore

questions

page 1 of 2

In Anthropic's Messages API, how do you send a system prompt to Claude?

level: juniorimportance: must knowfreq 78%

answer

  1. Not a role in the array
  2. Sits beside model and max_tokens
  3. String, or a list of text blocks
  4. Rendered before the conversation turns

basics

~20 s

The Messages API takes the system prompt as a top-level system parameter on the request body, alongside model, max_tokens and messages. It is not a message inside the messages array; that array carries only user and assistant turns.

solid answer

~40 s

Anthropic's `POST /v1/messages` puts the system prompt in a **top-level `system` field**, sibling to `model`, `max_tokens` and `messages` — not as an entry in the `messages` array. That array is for the conversation itself, whose turns are `user` and `assistant`. `system` accepts either a plain string or an array of text content blocks; the block form exists so you can attach per-block metadata such as cache breakpoints. This differs from OpenAI-shaped chat APIs, where the system prompt is simply the first message, so porting code between the two means moving the instruction out of the array rather than renaming a field. The server renders the request as tools, then system, then messages, which is also why a stable system prompt sits in the most reusable position of the prefix.

code

json · 8 lines
json
{
  "model": "claude-opus-4-6",
  "max_tokens": 1024,
  "system": "You are a terse SQL assistant. Answer with SQL only.",
  "messages": [
    {"role": "user", "content": "How many rows are in the orders table?"}
  ]
}

go deeper

for a junior

Remember the concrete shape: model, max_tokens, messages and system are all top-level keys of the request body, and the messages array holds only user and assistant turns.

for a middle

Be ready to explain that system accepts a string or an array of text blocks, and that the server renders tools, then system, then messages — which is why a constant system prompt sits in the most reusable prefix position.

for a senior

Show that you design the system prompt for stability: constant operator instructions up front, volatile per-request data pushed later, and awareness that it is re-sent and re-billed on every stateless call.

for a principal

Own the tradeoff between a large shared system prompt and per-request context: what belongs in the operator channel versus retrieved into a user turn, and how that choice drives both cost and prompt-injection exposure across tenants.

## The shape of the request Every call to Claude goes to one endpoint, `POST https://api.anthropic.com/v1/messages`, with the headers `x-api-key`, `anthropic-version: 2023-06-01` and `content-type: application/json`. The JSON body has three required fields — `model`, `max_tokens`, `messages` — and a set of optional top-level fields. **`system` is one of those top-level fields.** It is not a role, not a header, and not a special content block. ``` { "model": "...", "max_tokens": 1024, "system": "You are a terse SQL assistant.", "messages": [{"role": "user", "content": "count the orders table"}] } ``` ## Why it is not a message role The `messages` array models a dialogue: it alternates between `user` and `assistant` turns and represents things that were *said*. The system prompt is not a turn — it is operator configuration that applies to the whole exchange. Anthropic made that distinction structural rather than conventional. The practical consequences are worth knowing: - You cannot accidentally lose the system prompt by trimming the front of the history. Truncating old turns for context reasons touches `messages` only; `system` is untouched. - There is exactly one system prompt per request, so there is no ambiguity about which of several system messages wins. - Because the server renders `tools` → `system` → `messages` when building the prompt, a system prompt that never changes sits in a stable prefix position, which is what makes it reusable across requests. ## String or content blocks `system` accepts two forms. The simple form is a plain string. The structured form is an array of text content blocks: ``` "system": [ {"type": "text", "text": "You are a terse SQL assistant."}, {"type": "text", "text": "<50KB of schema documentation>"} ] ``` The array form exists so you can attach per-block metadata — most commonly a cache breakpoint on the last stable block — and so you can compose the prompt from independently-managed pieces. Only text blocks are valid here; images and documents belong in a user message's content. ## Porting from an OpenAI-shaped API In OpenAI-shaped chat APIs the system prompt is the first element of the messages list. Translating such code to Claude is not a field rename: you lift that entry out of the array entirely and assign its text to the top-level `system` parameter. If you leave it in the array as `{"role": "system", ...}`, the classic result is a `400 invalid_request_error` complaining about roles — historically the array accepted only `user` and `assistant`. Note one 2026 refinement: some current Claude models (the Opus 5 / Opus 4.8 class) additionally accept a mid-conversation `{"role": "system", "content": ...}` entry *inside* `messages`, with placement rules, as a way to inject an operator instruction later in a long conversation without disturbing the cached prefix. That is an addition for mid-conversation steering, not a change to where the primary system prompt goes — it still belongs in the top-level parameter, and support is model-dependent, so do not write it as the default. ## What belongs in it Role, tone, output-format rules, domain constraints, and any long reference material that is constant across requests. What does not belong: the user's actual question (that is a `user` message), and per-request volatile data such as a timestamp or a request id, which is better placed after the stable material so it does not destabilise the prefix. ## Failure modes to recognise - Sending `system` as a header or as `messages[0]` — rejected or ignored, and the model behaves as though it has no instructions. - Sending `system` as an array containing message objects (`{"role": ..., "content": ...}`) instead of text blocks — a validation error; the array form takes content blocks. - Rebuilding the system string on every request with an interpolated clock or UUID — it still works, but you have made a constant prefix volatile for no benefit. - Assuming the system prompt is free. It is input, it is re-sent on every request, and it is counted in the request's input token usage like everything else.

  • What happens if you put a system-role entry inside the messages array instead?
    Historically the array accepted only `user` and `assistant`, so a `system` entry came back as a `400 invalid_request_error` about roles. Current Opus 5 / Opus 4.8-class models added a narrow exception: a mid-conversation `{"role": "system"}` message may appear inside `messages` to inject an operator instruction late in a conversation, subject to placement rules. It is a steering mechanism, not the home of your primary system prompt.
  • Why would you send system as an array of text blocks rather than one string?
    Two reasons. First, per-block metadata: cache breakpoints are attached to a specific block, so the array form lets you mark where the stable prefix ends. Second, composition: a platform prompt, a tenant prompt and a schema dump can be assembled and versioned independently and concatenated at request time without string surgery.
  • Does the system prompt cost tokens on every request?
    Yes. The API is stateless, so the system prompt is re-sent and re-read on every call and counted in that request's input tokens. A large constant system prompt therefore has a per-request cost proportional to its size — which is exactly why the API offers a way to mark stable prefix content so repeated reads are billed differently.

The messages array is the transcript of a conversation; the system parameter is the briefing note handed to the participant before the conversation starts.

saying these in an interview costs you the question

  • Says the system prompt is the first message with role system
  • Thinks system is an HTTP header on the request
  • Assumes system must be a string and blocks are invalid
  • Believes the system prompt persists server-side between calls
  • Puts images or documents into the system parameter

context

open as a page

What do the Haiku, Sonnet, and Opus tiers mean in Claude's model line-up?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Haiku, Sonnet and Opus are Anthropic's three Claude size tiers, ordered smallest to largest. Haiku is the fastest and cheapest, Opus is the most capable and most expensive, and Sonnet sits between them. All three are called through the same Messages API.

open as a page

How do you declare a tool for Claude in Anthropic's Messages API?

level: juniorimportance: must knowfreq 72%

basics

~10 s

Send a top-level tools array on the Messages request. Each entry needs a name, a description that tells Claude when to call it, and an input_schema: a JSON Schema object describing the parameters.

open as a page

How do you send an image to Claude's Messages API inside a user message?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Set the user message content to an array of blocks and add an image block next to your text block. The image block carries a source object that is either base64 (with media_type and data) or a url. Image blocks belong in user turns.

open as a page

What does the required max_tokens parameter cap on an Anthropic Messages API call?

level: middleimportance: must knowfreq 72%

basics

~20 s

max_tokens is a hard ceiling on the tokens Claude generates in that one response, including any thinking tokens. It is required on every Messages API request, and hitting it truncates the output mid-stream with stop_reason "max_tokens".

open as a page

How must the messages array be structured in an Anthropic Messages API request?

level: middleimportance: must knowfreq 68%

basics

~20 s

messages is a non-empty array of turns, each with a role of user or assistant, starting with user and alternating. Each turn's content is either a plain string or an array of typed content blocks. The endpoint is stateless, so you resend the full history every call.

open as a page

How does token pricing differ across Claude's Haiku, Sonnet, and Opus tiers?

level: middleimportance: must knowfreq 60%

basics

~20 s

Anthropic prices each Claude model per million tokens, with input and output billed at separate rates and output costing several times more than input. Rates climb from Haiku to Sonnet to Opus, so per-request cost depends on both the tier and the token mix.

open as a page

Where do you place cache_control in an Anthropic Messages request, and why?

level: middleimportance: must knowfreq 72%

basics

~20 s

Anthropic's prompt cache is prefix-based: a cache_control marker ends a cacheable prefix rather than caching one block. Put it after everything stable — tools first, then system, then early messages — and keep anything that varies per request after it.

open as a page

How do you accumulate Anthropic content_block_delta events into a final message?

level: middleimportance: must knowfreq 60%

basics

~20 s

Key a buffer by the event's index field and append each delta to that block. text_delta contributes its text; input_json_delta contributes a partial_json string fragment that you concatenate and parse as JSON only once the block's content_block_stop arrives.

open as a page

What event sequence does the Anthropic Messages API emit over SSE when streaming?

level: middleimportance: must knowfreq 72%

basics

~10 s

An Anthropic stream opens with message_start, then for each content block a content_block_start, a run of content_block_delta chunks, and a content_block_stop. It closes with message_delta carrying the final stop_reason and output usage, then message_stop.

open as a page

What are the tool_choice modes in Anthropic's Messages API?

level: middleimportance: must knowfreq 62%

basics

~20 s

Four values: auto lets Claude decide (the default when tools are present), any forces at least one tool call, tool with a name forces that specific tool, and none forbids calling any tool. Each can also carry disable_parallel_tool_use.

open as a page

How do you return a tool_result to Claude after a tool_use block?

level: middleimportance: must knowfreq 78%

basics

~20 s

Append the assistant's entire content array to messages, then add a new user message whose content holds a tool_result block. That block carries tool_use_id matching the call's id, plus the result content. Then call the API again.

open as a page

How does Anthropic's Messages API turn image dimensions into billed input tokens?

level: middleimportance: must knowfreq 58%

basics

~20 s

Image cost scales with pixel area, not file size: Anthropic documents roughly width times height divided by 750 tokens. Oversized images are downscaled server-side first, so you pay for the resized version, capped by the model's resolution limit.

open as a page

How do you handle an Anthropic SSE error event that arrives mid-stream?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Treat it as a failed request even though bytes already reached the user. The HTTP 200 and headers are long gone, so surface the failure in your own stream protocol, keep or discard the partial text deliberately, and retry the whole call — there is no resume point.

open as a page

Which fields in Anthropic's Messages usage object report cache hits and writes?

level: juniorimportance: should knowfreq 44%

basics

~10 s

The usage object reports cache_creation_input_tokens for tokens written to the cache on this call and cache_read_input_tokens for tokens served from it. A non-zero read count means a hit; input_tokens counts only the uncached remainder.

open as a page

In the Anthropic Python SDK, how does messages.stream() differ from create(stream=True)?

level: juniorimportance: should knowfreq 52%

basics

~20 s

create(stream=True) returns a raw iterator of stream events that you assemble yourself. messages.stream() is a context manager returning a helper that accumulates the events for you, exposes a text-only iterator, hands back the finished message, and closes the HTTP response on exit.

open as a page

How does context window size vary across Claude's Haiku, Sonnet and Opus tiers?

level: middleimportance: should knowfreq 44%

basics

~20 s

Context window is not a tier ladder: the current Claude tiers share the same standard window, so moving from Haiku to Opus buys capability and not extra room. Extended long-context has been offered as a per-model opt-in with its own pricing, not as an Opus perk.

open as a page

What is the default TTL of Anthropic's prompt cache, and what does caching cost?

level: middleimportance: should knowfreq 52%

basics

~20 s

Anthropic cache entries live five minutes by default, and every hit refreshes that window. An optional one-hour TTL is requested with ttl inside cache_control. Writes cost a premium over base input price, reads a small fraction of it.

open as a page

Which Anthropic streaming event carries the final stop_reason and output usage?

level: middleimportance: should knowfreq 48%

basics

~10 s

message_delta. It arrives after the last content block and carries delta.stop_reason, delta.stop_sequence, and a usage object with the message's output token count — none of which exist yet at message_start.

open as a page

Which image formats and size limits does Claude's Messages API accept?

level: middleimportance: should knowfreq 45%

basics

~20 s

Four media types are accepted: image/jpeg, image/png, image/gif and image/webp. Each image may be up to 5 MB through the API, and a single request may carry up to 100 images. Base64 encoding inflates the payload by about a third.

open as a page

Which stop_reason values can an Anthropic Messages API response return?

level: seniorimportance: should knowfreq 52%

basics

~10 s

Common values are end_turn, max_tokens, stop_sequence and tool_use, plus pause_turn, refusal and model_context_window_exceeded. Read stop_reason before touching the content array, because some outcomes return HTTP 200 with little or no content.

open as a page

In the Anthropic Messages API usage object, what does input_tokens count?

level: seniorimportance: should knowfreq 45%

basics

~20 s

input_tokens counts everything the server read for that one request — system prompt, tool definitions and the entire resent message history — not just the newest user turn. output_tokens counts what the model generated in that response.

open as a page

How would you route a production workload across Claude tiers instead of one model?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Split the pipeline into steps and assign a tier per step: small models for classification, extraction and routing, the large model for the one genuinely hard step. Prove each assignment with a held-out eval set, and cascade with escalation only when escalation stays rare.

open as a page

Your Anthropic prompt cache never registers a read — how do you debug it?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Check the usage counters first: persistent cache_creation with zero cache_read means the prefix is either below the minimum cacheable size, changing between calls, expiring between calls, or routed to a different model. None of these raises an error.

open as a page

How do you detect a stalled Anthropic SSE stream, and what timeout should you set?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Time the gap between events, not the whole request. Reset a watchdog on every frame including ping keepalives, and fail the stream when the idle gap exceeds your budget. A total wall-clock timeout kills legitimate long generations instead.

open as a page

Why does a Claude agent loop with tool_choice any never terminate?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Because tool_choice applies to every request you send, and any forbids a plain-text answer. Each turn is therefore obliged to return a tool_use block, so a loop that exits on stop_reason leaving tool_use never exits. Force only the first turn.

open as a page

How do you reply when Claude emits multiple tool_use blocks at once?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Execute them all, then send every tool_result back in a single user message. Splitting them across separate user turns is invalid and teaches Claude to stop batching. A failed call still needs its block, marked is_error.

open as a page

How do you combine several images with text in one Claude Messages API turn?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Put every image as its own block in the same user message content array, place them before the text that asks about them, and label them in the text so answers can refer to them unambiguously. Each image is tokenised and billed separately.

open as a page

When would you choose a URL image source over base64 in Anthropic's Messages API?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Choose a URL source when the image is already hosted somewhere Anthropic can reach and you want to keep bytes out of the request body. Choose base64 for private or locally generated images, and the Files API when the same image is reused across many calls.

open as a page

How do you manage Claude model tier migrations and snapshot deprecations at scale?

level: principalimportance: should knowfreq 36%

basics

~20 s

Pin dated Claude snapshot identifiers in one config layer keyed by role, never as literals at call sites. Aliases float to newer snapshots and can change behaviour silently. Gate every migration on a held-out eval set, shadow real traffic, and keep the old identifier available for rollback.

open as a page

showing 1–30 of 31