skip to content

Which sampling parameters does claude-sonnet-5 reject, and what controls reasoning depth instead?

level: middleimportance: should knowfreq 52%

answer

  1. one knob left, the old ones are gone
  2. not ignored — the request fails
  3. adaptive replaces the fixed budget
  4. effort lives inside output_config
  5. low, medium, high, xhigh, max

basics

~10 s

claude-sonnet-5 removed temperature, top_p and top_k — sending any of them returns a 400. The fixed thinking budget (budget_tokens) is gone too. Depth and token spend are set with output_config.effort, alongside adaptive thinking.

solid answer

~40 s

On `claude-sonnet-5` the classic sampling knobs are not merely ignored, they are removed: `temperature`, `top_p` and `top_k` each cause the request to fail with a 400 validation error. The same is true of `thinking: {"type": "enabled", "budget_tokens": N}` — the fixed-budget form of extended thinking no longer exists on that model. What you use instead is a pair of settings. Thinking is configured as `{"type": "adaptive"}`, which is also what you get if you omit `thinking` entirely; `{"type": "disabled"}` is accepted if you genuinely want it off. Depth and overall token spend are then dialled with `output_config: {"effort": ...}`, which takes `low`, `medium`, `high`, `xhigh` or `max`, defaulting to `high`. The previous generation, `claude-sonnet-4-6`, still accepts the sampling parameters, which is why a straight model-string swap can turn a working request into a 400.

code

python · 12 lines
python
from anthropic import Anthropic

client = Anthropic()

message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=4096,
    thinking={"type": "adaptive"},
    output_config={"effort": "medium"},
    messages=[{"role": "user", "content": "Summarise the CAP theorem."}],
)
print(message.content)

go deeper

for a junior

Remember that on claude-sonnet-5 you do not send temperature, top_p or top_k at all — the request fails. Leaving the thinking parameter out is fine and gives you adaptive thinking.

for a middle

Explain that these parameters are removed rather than ignored, that adaptive thinking replaces the fixed budget_tokens form, and that effort sits inside output_config with high as the default.

for a senior

Demonstrate the migration judgement: strip sampling parameters instead of translating them, move temperature-0 routes onto structured outputs, sweep effort when cost-tuning, and re-baseline spend dashboards because adaptive thinking bills differently from a fixed budget.

for a principal

Own the fleet-level story: how request builders and middleware stop injecting model-incompatible defaults, how effort levels are chosen per workload class rather than set once globally, and how a model upgrade is smoke-tested against a real request before it reaches production.

## What actually changed For years, tuning a Claude call meant reaching for `temperature` (and occasionally `top_p` or `top_k`) to trade determinism against variety. On `claude-sonnet-5` those parameters were removed from the request surface. Removed is stronger than deprecated: the API does not accept-and-ignore them, it returns HTTP 400 with a validation error. Any request builder that unconditionally attaches `temperature=0` — a very common default in wrapper code and in framework integrations — fails outright on this model. The same removal applies to the older extended-thinking shape. `thinking: {"type": "enabled", "budget_tokens": N}` was how you bought a fixed number of reasoning tokens on earlier models. On `claude-sonnet-5` that form is rejected with a 400 as well. ## The replacement: adaptive thinking Thinking on `claude-sonnet-5` is configured as `thinking: {"type": "adaptive"}`. Adaptive means the model itself decides when to reason and how much, rather than you pre-allocating a token budget for reasoning you cannot predict the need for. Omitting the `thinking` parameter entirely runs adaptive as well, so the safe minimal request is simply not to send it. `{"type": "disabled"}` is accepted if a route genuinely must skip reasoning, but on modern models turning thinking off tends to cost more quality than it saves in tokens, and lowering effort is usually the better lever. ## The replacement: effort Overall depth and token spend are controlled by `output_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}`. Two mechanical details matter. First, `effort` lives *inside* `output_config` — it is not a top-level request field, and putting it at the top level is a validation error. Second, the default is `high`, which is what you get by omitting it, so a request that never mentions effort is already spending at the high setting. As a rough mapping: `low` suits subagents and mechanical tasks and produces fewer, more consolidated tool calls with less preamble; `high` is the usual balance point; `xhigh` is the strong setting for coding and agentic work; `max` is for when correctness matters more than cost. Because effort influences both quality and spend, it is the parameter to sweep when you are cost-tuning a route — not temperature, which no longer exists here. ## Why this bites during a model swap The previous Sonnet generation, `claude-sonnet-4-6`, still accepts `temperature`, `top_p` and `top_k`, and still tolerates `budget_tokens` as a transitional escape hatch for code that needs a hard reasoning ceiling. So the failure mode is asymmetric and easy to hit: a service that runs happily on 4.6 starts returning 400 on every request the moment someone edits the model string to `claude-sonnet-5`. The error is immediate and total rather than subtle, which is a mercy, but only if you have a smoke test that exercises a real request rather than just checking that the client constructs. ## What to do when migrating Strip the sampling parameters rather than trying to translate them — there is no effort value that is "equivalent to temperature 0", because effort governs reasoning depth and spend, not output randomness. If a route depended on temperature 0 for reproducible, machine-parseable output, the right replacement is structured outputs (`output_config.format` with a JSON Schema, or `strict: true` on tool definitions), which constrain the *shape* of the response directly instead of hoping low randomness keeps it well-formed. If a route depended on high temperature for variety, vary the prompt or sample multiple completions instead. Then convert any `budget_tokens` usage: delete the budget, set `thinking: {"type": "adaptive"}` (or omit it), and pick an effort level. Re-baseline your cost dashboards afterwards, because adaptive thinking spends differently from a fixed budget — sometimes less on easy requests and more on hard ones — and a per-request average taken from the old model will mislead you. ## Diagnosing in production A 400 that appears on every request immediately after a model change is almost always an unsupported parameter, not a credentials or quota problem. Log the full error body: the validation message names the offending field. Watch particularly for parameters injected by middleware or an SDK-level default rather than by your own call site, since those are invisible at the point where you are reading the code.

  • A route relied on temperature 0 to keep JSON output parseable. What replaces that on Sonnet 5?
    Constrain the shape directly instead of leaning on low randomness. Use structured outputs — `output_config.format` with a JSON Schema on the message request, or `strict: true` with `additionalProperties: false` on a tool definition — so the response is validated against the schema rather than merely likely to match it. That is stronger than temperature 0 ever was, since low randomness never actually guaranteed well-formed JSON.
  • Does claude-sonnet-4-6 behave the same way about these parameters?
    No, and that asymmetry is the migration trap. Sonnet 4.6 still accepts `temperature`, `top_p` and `top_k`, and still tolerates `thinking.budget_tokens` as a transitional escape hatch for code needing a hard reasoning ceiling. Sonnet 5 rejects all of them with a 400, so a bare model-string swap breaks every request until the parameters are stripped.
  • Where exactly does the effort setting go in the request, and what is the default?
    It goes inside `output_config` as `output_config: {"effort": "..."}` — never as a top-level field, which fails validation. Accepted values on Sonnet 5 are low, medium, high, xhigh and max, and the default when you omit it is high. So an untouched request is already running at high effort, which matters when you are budgeting spend.

saying these in an interview costs you the question

  • Says temperature is accepted but silently ignored on Sonnet 5
  • Assumes an effort level is equivalent to a temperature value
  • Puts effort at the top level instead of inside output_config
  • Keeps budget_tokens for a hard reasoning ceiling on Sonnet 5
  • Thinks omitting thinking turns reasoning off on Sonnet 5

context