skip to content

Which stop_reason values can an Anthropic Messages API response return?

level: seniorimportance: should knowfreq 52%

answer

  1. Says why generation stopped
  2. Not all values mean success
  3. One of them arrives with empty content
  4. HTTP 200 can still be a decline
  5. Branch before indexing content

basics

~10 s

Common values are end_turn, max_tokens, stop_sequence and tool_use, plus pause_turn, refusal and model_context_window_exceeded. Read stop_reason before touching the content array, because some outcomes return HTTP 200 with little or no content.

solid answer

~40 s

`stop_reason` on the response message says *why* generation ended. `"end_turn"` means the model finished naturally; `"max_tokens"` means it hit your output ceiling and the content is truncated; `"stop_sequence"` means one of your configured stop sequences was emitted, with the matched value reported in `stop_sequence`; `"tool_use"` means it wants a tool run before continuing. Beyond those, `"pause_turn"` appears on long server-side tool loops and is resumed by re-sending the exchange unchanged; `"refusal"` means a safety classifier or the model declined — a **successful HTTP 200** that may carry an empty `content` array; and `"model_context_window_exceeded"` reports that the request outgrew the window. The operational rule: branch on `stop_reason` first, and only then read `content` — code that indexes `content[0]` unconditionally breaks the first time a refusal comes back.

code

python · 14 lines
python
response = client.messages.create(
    model="claude-opus-4-6",
    max_tokens=4096,
    messages=[{"role": "user", "content": prompt}],
)

if response.stop_reason == "refusal":
    return declined(response)          # 200, content may be empty
if response.stop_reason == "max_tokens":
    return truncated(response)         # do not parse as final
if response.stop_reason != "end_turn":
    return unexpected(response)        # unknown or loop-continuing value

return response.content[0].text

go deeper

for a junior

Know that the response carries a stop_reason field and that end_turn means the model finished while max_tokens means the answer was cut off.

for a middle

Be able to enumerate the values and say what each implies for the next request — including that stop_sequence reports the matched string in a companion field.

for a senior

Show the production reflex: branch on stop_reason before touching content, treat refusal as a 200 with possibly empty content, resume pause_turn by resending unchanged, and alert on the distribution of stop reasons.

for a principal

Own the policy for non-happy outcomes across services — which stop reasons are retryable and how, what the fallback path is for a decline, and how stop-reason rates feed capacity and prompt-quality reviews.

## What the field is Every non-streaming response from `POST /v1/messages` is a message object with `id`, `type`, `role`, `model`, `content`, `usage`, `stop_sequence` — and `stop_reason`. The last of these is the control-flow field: it tells your code why the model stopped generating, and therefore what to do next. Treating it as decoration is the difference between an integration that degrades gracefully and one that throws on a Tuesday. ## The values **`end_turn`** — the model finished on its own. The happy path: read the content and move on. **`max_tokens`** — generation hit the ceiling you set. The content is truncated at an arbitrary token, so it is not safe to parse or display as final. Raise the ceiling and retry, or ask for a shorter format. **`stop_sequence`** — the model emitted one of the strings you passed in `stop_sequences`, and generation halted there. The companion `stop_sequence` field on the response names which one matched; it is `null` for every other stop reason. Note that the stop text itself is not included in the returned content. **`tool_use`** — the model wants a tool executed and paused for you to supply the result. Your loop continues rather than terminating. **`pause_turn`** — the server ran a long server-side tool loop and reached its own iteration limit rather than finishing. The remedy is to re-send the conversation with the assistant response appended and let the server resume; you deliberately do *not* add an extra "continue" instruction, because the server detects the pending server-side work itself. **`refusal`** — a safety classifier, or the model itself, declined the request. This is the value that catches people out, because it arrives as a **200, not an error**: your HTTP layer sees success, and `content` may be an empty array. Current models pair it with a `stop_details` object describing the category; that object is `null` for every other stop reason, so branch on `stop_reason` and treat `stop_details` as informational. **`model_context_window_exceeded`** — the request, or the request plus what was being generated, outgrew the model's context window. The fix is on your side: trim or summarise history. ## The habit that matters ``` if response.stop_reason == "refusal": handle_refusal(response) elif response.stop_reason == "max_tokens": handle_truncation(response) elif response.stop_reason == "end_turn": text = response.content[0].text ``` The ordering is the point. A very common production bug is `text = response.content[0].text` executed before any check: on a refusal with an empty content array it raises an index error, at runtime, in whatever request happened to trip a classifier. Because the HTTP status is 200, no retry middleware catches it and no error metric fires — the failure looks like a crash in your own code rather than a provider outcome. ## What good handling looks like - **Exhaustive branch, explicit default.** Treat an unrecognised `stop_reason` as "do not assume completeness" rather than falling through to the happy path. New values do get added, and defaulting to success is the unsafe direction. - **Distinguish retryable from terminal.** `max_tokens` and `pause_turn` are recoverable with a modified or repeated request. A refusal is not retryable in the naive sense: replaying the identical body against the identical model gets the identical outcome; if you retry, retry differently — a different model, or a reformulated request. - **Instrument it.** `stop_reason` is a free, high-signal metric. A rising `max_tokens` rate means your ceilings are wrong for the traffic you now serve; a rising refusal rate means your prompts have drifted into territory a classifier dislikes; a `model_context_window_exceeded` at all means your history management is not doing its job. - **Never infer from the text.** Truncated prose frequently ends on a word that reads like a sentence ending. The status field is authoritative; the prose is not. ## Streaming note On a streamed response the same value arrives on the terminal delta event rather than on a whole message object, but the semantics and the branching obligations are identical — the SDK helpers that assemble a final message expose it in the same place.

  • Why is a refusal returned as HTTP 200 rather than a 4xx?
    Because the request was well-formed and the model ran — the outcome is a decline, not a protocol error. Encoding it in `stop_reason` keeps transport-level failures (auth, validation, rate limits) separate from generation outcomes. The cost is that generic HTTP error handling never sees it, so the check has to live in your application code.
  • What is the correct response to stop_reason "pause_turn"?
    Re-send the conversation with the assistant's response appended and nothing else added. The server ran a server-side tool loop up to its iteration limit and will resume where it left off, because it detects the pending server-side work in the trailing blocks. Injecting an extra "continue" user message confuses that detection rather than helping it.
  • How should an unrecognised stop_reason value be treated?
    As not-complete. New values are added over time, and a default branch that falls through to the success path will one day render or parse a response that never finished. Log the unexpected value with the request id, surface it as a distinct outcome, and fail closed rather than assuming end_turn semantics.

saying these in an interview costs you the question

  • Assumes any HTTP 200 means a usable answer
  • Indexes content[0] before checking stop_reason
  • Thinks refusals arrive as a 4xx error
  • Retries a refusal unchanged against the same model
  • Infers completeness from how the text reads

context