What does the required max_tokens parameter cap on an Anthropic Messages API call?
answer
- Required, no default
- Bounds only what is generated
- Model never sees the number
- Reasoning tokens spend it too
- Truncation shows up in stop_reason
basics
~20 smax_tokens is a hard ceiling on the tokens Claude generates in that one response, including any thinking tokens. It is required on every Messages API request, and hitting it truncates the output mid-stream with stop_reason "max_tokens".
solid answer
~50 s`max_tokens` is a **required** field on `POST /v1/messages`, and it bounds the *output* of a single response — the assistant's generated tokens, including reasoning/thinking tokens where the model produces them. It does not bound the input, and it is not a target: the model is not told the number, so it does not shape its answer to fit. If generation reaches the cap, the response comes back truncated at whatever token it was on, with `stop_reason: "max_tokens"`. Omitting the field is a `400 invalid_request_error`; so is exceeding the model's own output limit. In practice you set it generously — a low ceiling produces silently truncated answers and forces a retry — but very large values on a non-streaming request risk an HTTP timeout, which is why SDKs push you to stream when the ceiling is high.
code
json · 5 lines{
"model": "claude-opus-4-6",
"max_tokens": 64,
"messages": [{"role": "user", "content": "Explain TCP congestion control."}]
}go deeper
Know that max_tokens is required on every request and that it limits only what Claude writes back, not what you send in.
Explain that it is an enforced ceiling the model cannot see, that reasoning tokens spend it too, and that hitting it yields stop_reason max_tokens with output truncated mid-token.
Demonstrate the production habit: branch on stop_reason, never parse a truncated payload, size the ceiling against the longest plausible answer, and stream when it is large so no request dies on an HTTP timeout.
Own the policy across services — per-route ceilings tied to expected output shape, retry and degradation rules for truncation, and how output ceilings interact with latency budgets and overall spend controls.
## Required, and output-only `POST /v1/messages` has exactly three required body fields: `model`, `messages`, and `max_tokens`. The third one surprises people coming from APIs where the equivalent parameter is optional and defaults to something sensible. Here there is no default — omit it and the request fails validation with `400 invalid_request_error`. What it bounds is the **response**: the number of tokens the model is permitted to generate on this turn. It says nothing about the input. A 200,000-token conversation with `max_tokens: 100` is a perfectly valid request; it just cannot produce more than 100 tokens of answer. ## A ceiling, not a budget the model is aware of This is the distinction that separates a middling answer from a good one. `max_tokens` is enforced *by the server on the sampler*: when the count is reached, generation stops, full stop. The model is not shown the number and cannot plan around it. So a truncated response is not a shorter, well-formed answer — it is an answer cut off mid-sentence, mid-JSON, or mid-code-block. If you want a short answer, ask for one in the prompt; `max_tokens` is a safety rail, not a style control. Anthropic has a separate, opt-in mechanism for advisory pacing on long agentic runs, where the model *is* made aware of a token allowance and can wind down gracefully. That is a distinct feature; `max_tokens` remains the blunt enforced ceiling underneath it. ## Detecting the cap The response's `stop_reason` is the signal: - `"end_turn"` — the model finished on its own. - `"max_tokens"` — it was cut off at your ceiling. Any production caller should branch on this. Treating a `max_tokens` response as a complete answer is how truncated JSON reaches a parser and how half a sentence reaches a user. The usual recoveries are: raise the ceiling and retry, ask for a more compact format, or split the work. ## Thinking tokens share the budget On models that produce reasoning tokens, those tokens are generated output and count against `max_tokens` alongside the visible text. The characteristic failure is a request where reasoning consumed most of the allowance and the visible answer was truncated immediately after. The fix is to raise `max_tokens` (or reduce how much reasoning you ask for), not to assume the model misbehaved. ## Choosing a value Each model publishes its own maximum output length; asking for more than the model allows is a validation error, so the ceiling is bounded above by the model, not by your ambition. Within that: - **Classification or extraction** — a small value (a few hundred tokens) is a genuine cost and latency guard. - **General non-streaming calls** — set it comfortably above the longest plausible answer. The cost of a high ceiling is zero if the model does not use it; you are billed for tokens produced, not tokens permitted. - **Long outputs** — current Claude models support very large output ceilings (up to 128K tokens on the current Opus and Sonnet generation as of mid-2026), but a single non-streaming HTTP request that has to sit open while all of that is generated is a timeout risk. Stream instead; SDKs will actively steer you toward streaming for large ceilings. ## Edge cases worth knowing - `max_tokens: 0` is accepted. The server runs prefill and returns immediately with an empty `content` array and `stop_reason: "max_tokens"` — useful for warming a prompt prefix without paying for output. - `max_tokens` interacts with `stop_sequences`: whichever condition trips first ends the response, and `stop_reason` tells you which one it was. - Because it is a per-response ceiling, it does not accumulate across the turns of an agentic loop. Ten tool-calling turns with `max_tokens: 4096` may collectively generate far more than 4096 tokens; if you want to bound the whole loop, you must track that yourself. ## The interview framing The question is really testing whether you understand that the API is a stateless single-turn generation call with an enforced output ceiling and a status field telling you why generation stopped. Candidates who say "it limits the length of the conversation" or "it counts the prompt plus the answer" have the wrong model of the endpoint.
- How do you tell a truncated response apart from a complete one in code?Read `stop_reason` before you read the content. `"end_turn"` means the model finished; `"max_tokens"` means it was cut off at your ceiling. Never infer completeness from the text — truncated prose often ends on a plausible-looking word, and truncated JSON will simply fail to parse downstream.
- Is there a cost to setting max_tokens far higher than you need?Not directly — billing counts tokens actually produced, not tokens permitted. The real costs are operational: a very large ceiling on a non-streaming request keeps one HTTP connection open long enough to risk a client or proxy timeout, and it removes the guard that would have stopped a runaway generation early. Set it generously, and stream when it is large.
- Does max_tokens bound the whole agentic loop or just one response?Just one response. Each request in a multi-turn tool loop carries its own ceiling, so ten turns at 4096 can generate far more than 4096 tokens in total. If you need a bound on the whole run, accumulate output tokens across turns yourself, or use a mechanism designed to give the model an overall allowance it can pace against.
saying these in an interview costs you the question
- Says max_tokens limits the input or the whole conversation
- Thinks the model targets max_tokens as a length goal
- Assumes it is optional with a sensible default
- Treats a max_tokens response as a complete answer
- Forgets reasoning tokens are charged against the same ceiling