skip to content

What are the tool_choice modes in Anthropic's Messages API?

level: middleimportance: must knowfreq 62%

answer

  1. Four types, one object
  2. auto is the default
  3. any is not called required here
  4. Forcing means no plain-text answer
  5. A key inside tool_choice caps parallelism

basics

~20 s

Four values: auto lets Claude decide (the default when tools are present), any forces at least one tool call, tool with a name forces that specific tool, and none forbids calling any tool. Each can also carry disable_parallel_tool_use.

solid answer

~40 s

`tool_choice` is a top-level object on the Messages request with a `type` field taking one of four values. `{"type": "auto"}` is the default whenever `tools` is present: Claude decides between answering directly and calling a tool. `{"type": "any"}` requires Claude to call *some* tool — it may not reply with plain text. `{"type": "tool", "name": "..."}` narrows that to one named tool, which is how you use the tool-call mechanism as a structured-extraction channel. `{"type": "none"}` forbids calls while leaving the definitions in the prompt, useful for a summarisation turn at the end of a loop. Any of the four may additionally set `"disable_parallel_tool_use": true`, capping the response at a single `tool_use` block.

go deeper

for a junior

Recall the four types — auto, any, tool, none — and that auto is what you get if you send tools without a tool_choice at all.

for a middle

Explain what each mode forbids, that tool_choice is an object with a type field, and that forcing a named tool is the classic structured-extraction trick.

for a senior

Show that forcing applies per request in a loop, know that none preserves the cached prefix while withdrawing permission, and reach for a better description before reaching for any.

for a principal

Own the policy: which turns in a pipeline are forced, which are free, and how disable_parallel_tool_use trades latency for safety across a tool surface with side effects.

## Where it sits `tool_choice` is a top-level parameter on `POST /v1/messages`, a sibling of `tools`. It is an object, not a string — `tool_choice: "auto"` is not valid on Anthropic's API even though other vendors accept the bare string form. If you omit it while `tools` is present, the server behaves as `{"type": "auto"}`. ## The four modes **`auto`** — Claude decides. It may answer in plain text, or emit one or more `tool_use` blocks, or mix a text block and tool calls in the same assistant message. This is what you want for almost all agent work, because the model's judgement about whether a tool is needed is usually the point of giving it tools. **`any`** — Claude must emit at least one `tool_use` block. It cannot answer with text alone. Use this when the turn *has* to produce an action and there are several plausible actions to choose between. **`tool`** — takes a required `name` alongside the type, and forces that one tool. The dominant real-world use is not agency at all but structured extraction: declare a single tool whose `input_schema` is the record you want, force it, and read the arguments off the resulting `tool_use` block. That gives you schema-shaped output without parsing prose. **`none`** — Claude may not call anything. The tool definitions stay in the request (and stay billed as input tokens), so the prefix and its cache remain intact; only the model's permission changes. This is how you close out an agent loop: same conversation, same tools, but this turn must produce the final prose answer. ## disable_parallel_tool_use By default Claude may return several `tool_use` blocks in one assistant message so you can execute them concurrently. Adding `"disable_parallel_tool_use": true` to whichever `tool_choice` object you are sending caps it at one call per response. It is a key *inside* the `tool_choice` object, not a separate top-level parameter. Reach for it when your tools mutate shared state and concurrent execution would race, or when you want one-step-at-a-time behaviour you can approve individually — at the cost of more round trips and higher end-to-end latency. ## The trap: forcing is a per-request setting, not a session setting Forced modes describe *this* request. In an agent loop you build a fresh request each turn from the accumulated `messages`, so whatever `tool_choice` you pass is re-applied every time. Passing `any` or `tool` from a constant means every turn is required to call a tool, and the model can never reach a natural end of turn. The standard pattern is to force on the first turn only, then switch to `auto` for the rest of the loop. ## Vendor confusion OpenAI's equivalent parameter uses `"required"` where Anthropic uses `"any"`, and accepts bare strings. Gemini expresses the same idea as a `tool_config` with `AUTO` / `ANY` / `NONE` function-calling modes. Porting code across these without adjusting the literal values produces a 400 rather than a subtle behaviour change, which is at least a fast failure — but recalling the wrong literal in an interview is a visible tell. ## Interaction with tool descriptions Forcing is a blunt instrument and is often reached for to paper over a weak description. If the model is not calling a tool you believe it should, the first fix is a description that states the trigger condition, not `tool_choice: any`. Forcing changes the model's options; it does not improve its judgement about which tool fits, and with several tools declared a forced `any` can push it into calling the wrong one rather than the none it correctly wanted. ## Choosing in practice Use `auto` by default. Use `tool` for one-shot structured extraction and for the deterministic first step of a pipeline. Use `any` sparingly, when a no-op turn is genuinely invalid. Use `none` for a final summarisation turn, or to temporarily disarm a dangerous tool surface without rebuilding the request and losing the cached prefix. Add `disable_parallel_tool_use` only when concurrency is actually unsafe.

  • If you want structured JSON out of Claude, when would you force a single tool instead of using the response format controls?
    Forcing one tool whose `input_schema` is your record shape is the long-standing way to get schema-shaped output: you read `tool_use.input` instead of parsing prose. It is still the right choice when the same call may either extract a record or take a real action, or when you are on a path where response-format constraints are unavailable. For a pure extraction turn with no tools involved, the dedicated output-format controls are more direct.
  • Why leave the tools array in the request when you set tool_choice to none?
    Because removing it changes the bytes at the very front of the prompt, which invalidates the cached prefix for that request and everything after it. `none` keeps the prefix byte-identical and only withdraws permission to call. You still pay input tokens for the definitions, but on a long conversation that is far cheaper than a cache miss on the whole history.
  • When is disable_parallel_tool_use worth the extra round trips?
    When concurrent execution is unsafe — tools that mutate the same row, take a lock, or charge money — or when a human approves each call individually. It caps the response at one `tool_use` block, so a task needing four calls becomes four sequential round trips instead of one fan-out. That is a real latency cost, so make it a per-tool-surface decision rather than a global default.

saying these in an interview costs you the question

  • Calls the forcing value required, which is OpenAI's literal
  • Passes tool_choice as a bare string instead of an object
  • Thinks none removes the tool definitions from the prompt
  • Believes disable_parallel_tool_use is a top-level request parameter
  • Assumes auto must be set explicitly for Claude to call tools

context