skip to content

How do parallel tool calls work in OpenAI's API, and when do you disable them?

level: middleimportance: should knowfreq 50%

answer

  1. One assistant message, several calls
  2. Concurrency is yours to run and bound
  3. Every id must be answered, even failures
  4. Side effects argue for sequencing
  5. A single boolean turns it off

basics

~20 s

One assistant message can carry several entries in tool_calls; run them concurrently, then append one tool message per tool_call_id before the next request. Set parallel_tool_calls to false when calls must be ordered or have side effects.

solid answer

~50 s

By default the model may pack several independent calls into a single assistant message — three weather lookups for three cities come back as three entries in `message.tool_calls`, each with its own `id`. That is a latency win: you execute them concurrently and pay for one model turn instead of three. The obligation is that **every** `tool_call_id` gets its own `{"role": "tool", "tool_call_id": ..., "content": ...}` message appended before you call the API again; omit one and the next request fails with a 400. Their order within the array does not matter, only their presence. Set `parallel_tool_calls: false` when the calls have side effects, when one call's output should inform the next, or when a downstream system cannot absorb the concurrency — that caps the assistant message at one call and forces a strictly sequential loop.

code

python · 17 lines
python
import json
from concurrent.futures import ThreadPoolExecutor

def run_call(call, handlers):
    try:
        args = json.loads(call.function.arguments)
        return call.id, json.dumps(handlers[call.function.name](**args))
    except Exception as exc:                      # never drop a call
        return call.id, json.dumps({"error": str(exc)})

def answer_all(assistant_msg, handlers, messages):
    calls = assistant_msg.tool_calls or []
    with ThreadPoolExecutor(max_workers=4) as pool:
        results = list(pool.map(lambda c: run_call(c, handlers), calls))
    for call_id, content in results:
        messages.append({"role": "tool", "tool_call_id": call_id, "content": content})
    return messages

go deeper

for a junior

Know that message.tool_calls is a list and can hold more than one call, and that you loop over it appending one tool message per entry.

for a middle

Explain the default, the 400 you get when an id goes unanswered, and the parallel_tool_calls flag — plus why batching independent lookups saves a whole model turn.

for a senior

Demonstrate the operational discipline: bounded concurrency in the dispatcher, per-call timeouts that still emit a message, and errors returned as content rather than raised out of the loop.

for a principal

Own the safety argument — which tool surfaces are allowed to run in parallel at all, where mutual exclusion belongs in code rather than prompt text, and how the latency win trades against auditability of the trace.

## What the model can do in one turn Since tool calling replaced the older single-`function_call` formulation, an assistant message is allowed to contain a *list* of calls rather than one. Ask "compare the weather in Oslo, Lisbon and Cairo" against a single-city tool and a capable model will emit three entries in `message.tool_calls`, each with a distinct `id` and its own `arguments` string. `parallel_tool_calls` defaults to true, so this is the behaviour you get unless you opt out. ## Why it is worth having Each model turn costs a full round trip plus the price of re-reading the whole conversation. Three sequential turns for three independent lookups means three times that overhead and three times the wall-clock latency. Batching them into one turn means one inference, one prompt re-read, and — if you dispatch with a thread pool, `asyncio.gather` or your language's equivalent — the tool latency of the slowest call instead of the sum. On tool-heavy agents this is one of the largest single latency wins available. ## The obligation it creates The transcript rule is unforgiving: an assistant message with N entries in `tool_calls` must be followed by N tool messages, one per `tool_call_id`, before the conversation is valid again. Miss one and the API rejects the next request with a 400 rather than proceeding with partial information. Two consequences follow. First, error handling must be total. If one of three concurrent calls throws, you cannot simply skip it — you still append a tool message for its id whose content describes the failure. The model then sees two results and one error and can decide whether to retry, work with what it has, or tell the user. Skipping it breaks the request outright. Second, timeouts must be per-call and must still produce a message. A hung tool that never returns is not just a slow turn; it is a transcript you can never legally continue. Wrap each dispatch in a timeout and synthesize `{"error": "timed out after 10s"}` as the content when it fires. The *order* of the tool messages relative to each other is not significant — the ids do the binding — but they must all sit after the assistant message that requested them. ## When to turn it off `parallel_tool_calls: false` limits the model to at most one call per assistant message. Reach for it when: - **The tools mutate state.** Two concurrent writes to the same record, or a `create` and a `delete` fired together, are a correctness problem the model has no way to reason about. Sequencing them means each call's effect is visible in the transcript before the next is chosen. - **Calls are logically dependent.** If the model needs an id from the lookup before it can call the update, forcing one call per turn stops it from guessing the id and calling both at once. - **The downstream cannot take the concurrency.** A rate-limited third-party API or a small connection pool may prefer the serialized shape over your own semaphore. - **You want a deterministic, auditable trace.** One call per turn makes the transcript a readable sequence for debugging and for compliance review. The cost is real: you re-pay the prompt on every step, so a five-tool task becomes five model turns. Disable deliberately, not by default. ## Executing them safely Concurrency belongs on your side, and it is your responsibility to bound it. Dispatch each call through a name-to-handler map, run them under a semaphore sized to what your dependencies tolerate, and collect results into the same order-independent structure keyed by `tool_call_id`. Nothing about the model's output tells you the calls are safe to run together — the model is guessing at independence from the phrasing of your tool descriptions. If two of your tools must never run concurrently, either disable parallel calls or enforce mutual exclusion in your dispatcher; do not assume the model will respect an instruction in a description. ## Duplicate calls A parallel batch can contain the same function twice with identical arguments — a mild failure mode that wastes a call. Both entries still need answers, but you can execute once and copy the result into both tool messages. If it happens routinely, the tool description is probably ambiguous about whether it accepts multiple items; a tool that takes a list of cities is usually better than three parallel single-city calls.

  • Does the order of the tool messages have to match the order of the tool_calls array?
    No. The tool_call_id is what binds a result to its request, so the tool messages can appear in any order among themselves — which is convenient when you resolve them as concurrent futures complete. What is required is that all of them appear, and that they follow the assistant message that requested them. Interleaving an unrelated user message before they are all answered breaks the transcript.
  • One tool in a parallel batch times out. What do you send back?
    Still send a tool message for that id, with content describing the timeout — an object like {"error": "timed out"} is enough. The request is invalid without it, so a hung tool would otherwise strand the conversation permanently. Enforce the timeout on your side rather than waiting on the dependency, and let the model decide whether to retry the call, proceed with partial data, or report the failure to the user.
  • How would you stop the model from calling a mutating tool alongside a read tool?
    Do not rely on wording in the tool description — set parallel_tool_calls to false so at most one call can appear per assistant message, or enforce mutual exclusion in your dispatcher so conflicting handlers serialize regardless of what arrives. The model infers independence from your descriptions and has no visibility into locks, transactions, or idempotency, so the guarantee has to live in your code.

saying these in an interview costs you the question

  • Thinking each tool call always arrives in its own assistant message
  • Skipping the tool message for a call that failed
  • Assuming OpenAI runs the calls concurrently for you
  • Believing tool messages must be ordered to match the calls
  • Leaving parallel calls on for tools that mutate shared state

context