skip to content

How do you reply when Claude emits multiple tool_use blocks at once?

level: seniorimportance: should knowfreq 48%

answer

  1. Several blocks, one reply turn
  2. All or nothing on results
  3. Independent by construction
  4. Splitting quietly kills batching
  5. One toggle serialises the fan-out

basics

~20 s

Execute them all, then send every tool_result back in a single user message. Splitting them across separate user turns is invalid and teaches Claude to stop batching. A failed call still needs its block, marked is_error.

solid answer

~40 s

Parallel tool use is on by default: one assistant message may contain several `tool_use` blocks, each with its own id. The correct reply is one `user` message whose content array holds a `tool_result` block for **every** call, in any order, each carrying its matching `tool_use_id`. Every call must be answered — you cannot answer two and defer the third, and you cannot spread them across successive user turns. Failures are answered too, with `is_error: true`. Because the calls are independent, execute them concurrently: the wall-clock win over sequential execution is the whole point of the feature. If concurrent execution is unsafe for your tools, set `"disable_parallel_tool_use": true` inside `tool_choice` so Claude only ever asks for one at a time.

go deeper

for a junior

Know that one assistant turn can carry several tool_use blocks and that all their results go back together in a single user message, each with its own tool_use_id.

for a middle

Explain why the calls can run concurrently, why every block must be answered, and how disable_parallel_tool_use inside tool_choice turns the behaviour off.

for a senior

Demonstrate the operational handling: settle-all concurrency, per-tool timeouts, partial failure reported as is_error results, and the quiet degradation that follows from splitting results across turns.

for a principal

Own the fan-out policy — which tools may run concurrently at all, idempotency and rate-limit budgets for downstream services, and the latency-versus-safety call on serialising a side-effecting surface.

## What parallel tool use is When Claude sees that several independent facts are needed to answer, it can request them all in a single assistant turn: the `content` array comes back with two, three or more `tool_use` blocks, each with a distinct server-generated `id`, its own `name`, and its own `input`. There may be a leading `text` block too. `stop_reason` is `tool_use`, exactly as for a single call — the count is not signalled anywhere except in the array itself, so your handler must iterate rather than reach for the first block. This is default behaviour on current Claude models, not something you opt into. ## The reply contract One assistant turn with N `tool_use` blocks is answered by **one** user turn containing N `tool_result` blocks. That is the whole rule, and each half of it matters. *All of them.* Every `tool_use` in the immediately preceding assistant message must have a matching `tool_result`. Answering a subset produces a 400. There is no mechanism for deferring a call to a later turn — if you genuinely cannot run one now, return a result saying so and let the model decide what to do. *In one message.* Splitting results across multiple consecutive user messages is malformed, because the API expects the results to answer the immediately preceding assistant turn. Even where a client library papers over the structure, the conversation Claude sees then implies that batched calls come back awkwardly, and the model drifts toward one-call-at-a-time behaviour on later turns — you lose the latency benefit without any error telling you why. Order within the array does not have to match the order of the `tool_use` blocks; the `tool_use_id` join is what pairs them. ## Execute concurrently The calls are independent by construction — Claude batched them precisely because none depends on another's output. Running them sequentially throws away the benefit: three 400 ms lookups become 1.2 s instead of 400 ms. Use a thread pool, `asyncio.gather`, `Promise.all`, or whatever your runtime offers, then assemble the results array once everything has settled. Use a settle-all primitive rather than a fail-fast one, because a single failure must not discard the other results — they still have to be reported. ## Partial failure Mixed outcomes are the normal case, not an edge case. Two calls succeed, one times out: send all three blocks, with the failed one carrying `"is_error": true` and a message the model can act on. Claude then typically answers from what succeeded and either retries the failure or tells the user which part is missing. Dropping the failed block breaks the request; replacing all three with an apology text message throws away work you already paid for. ## Bounding the fan-out Parallel calls multiply everything at once: concurrent load on your downstream services, tokens returned into the context, and blast radius if the tools have side effects. Three concurrent read-only lookups are free money. Three concurrent writes against the same row are a race you did not design. The lever is `"disable_parallel_tool_use": true`, set inside whichever `tool_choice` object you send; it caps each response at one `tool_use` block, trading round trips for serialisation. Apply it when the tool surface mutates shared state, when a human approves each call, or when downstream rate limits are tight enough that a burst is worse than a queue. ## Timeouts and the slowest call A batch finishes when its slowest member does, so one tool with a pathological tail latency sets the pace for the whole turn and, in a long loop, for the whole agent. Give each tool its own timeout well below the request timeout and convert a breach into an error result rather than letting it hang — a fast "timed out" that Claude can route around beats a turn that stalls. ## Idempotency Because a model may retry a call after an error result, and because your own code may retry on a transient failure, side-effecting tools should be idempotent or carry a caller-supplied key. Parallel fan-out makes duplicate execution more likely, not less, since a partially failed batch invites exactly the retry that duplicates the successful half.

  • What actually goes wrong if you return the results across two separate user messages?
    The results no longer answer the immediately preceding assistant turn, so the request is malformed and the API rejects it. Worse, where the structure survives, the conversation Claude sees models an awkward round trip for batched calls, and it drifts toward requesting one tool at a time on later turns. You lose the concurrency win with no error explaining why.
  • Two of three parallel calls succeed and one times out. What do you send?
    All three tool_result blocks in one user message: the two successes with their output, the third with `is_error: true` and a short message such as "pricing service timed out after 3s". Claude then answers from what it has and either retries or flags the gap. Dropping the failed block invalidates the request; discarding the successes wastes work already paid for.
  • When would you turn parallel tool use off entirely?
    When concurrency is unsafe rather than merely slower: tools that mutate the same record, take locks, charge money, or sit behind a tight rate limit where a burst trips a 429. Also when a human approves each call individually. Set `"disable_parallel_tool_use": true` inside `tool_choice`; you trade a single fan-out round trip for several sequential ones.
  • Why should side-effecting tools be idempotent when parallel calls are enabled?
    A partially failed batch invites a retry of the whole batch, and an error result invites the model to call again. Both paths can re-run a call that already succeeded. An idempotency key or a naturally idempotent operation turns a duplicate into a no-op instead of a second charge or a second email.

saying these in an interview costs you the question

  • Handles only response.content[0] and ignores later tool_use blocks
  • Sends one user message per tool_result
  • Runs independent parallel calls sequentially
  • Omits the failed call's block and returns only successes
  • Thinks parallel tool use must be enabled with a flag

context