In LLM tool calling, how are parallel tool results returned and paired to their calls?
answer
- Independent requests, one turn
- Answer all of them, or none continues
- The join key is not the index
- Completion order lies
- Carry the id, not the position
basics
~20 sOne assistant turn can carry several independent tool-call blocks. All of them are answered in a single following turn holding one result block per call, each matched by the call id it echoes — never by array position or completion order.
solid answer
~50 sWhen requests are independent, a model may emit several tool-call blocks in one assistant turn — a restaurant-chain inventory assistant asking `check_stock` for twelve stores at once, for example. Your code may run them concurrently, but the contract is strict on the way back: **every** outstanding call must be answered, and all the results go into one following turn as twelve result blocks. Pairing is by the id each call carried, not by position and not by the order the tools finished. That distinction is where the classic bug lives — a runtime that appends results in completion order and lets the model align them by index will silently report the Leeds stock level under Brighton, and the answer looks perfectly confident. Assume nothing about execution order, either: parallel calls carry no dependency information, so if one must run before another, the model should have emitted them in separate turns.
code
json · 12 lines[
{ "role": "assistant", "content": [
{ "type": "tool_use", "id": "c_brighton", "name": "check_stock",
"input": { "store": "brighton", "sku": "UMB-01" } },
{ "type": "tool_use", "id": "c_leeds", "name": "check_stock",
"input": { "store": "leeds", "sku": "UMB-01" } }
] },
{ "role": "user", "content": [
{ "type": "tool_result", "tool_use_id": "c_leeds", "content": "{\"units\": 3}" },
{ "type": "tool_result", "tool_use_id": "c_brighton", "content": "{\"units\": 14}" }
] }
]go deeper
Know that one assistant turn can request several tools at once and that all the results go back together in a single turn, each tagged with the id of the call it answers.
Explain why id pairing beats positional pairing, and describe the concrete bug: results appended in completion order get aligned by index and the wrong store's number is reported with no error anywhere.
Show the operational habits — bounded concurrency, per-call timeouts, a failure result rather than a dropped call, and an assertion that call ids and result ids match before the turn is sent. Mention staggered-latency tests.
Own where fan-out belongs. Decide when parallel conversational calls are the right instrument versus a fixed workflow or a sandbox loop, and set the tool-design rule that mutating operations are never split into calls whose order the protocol cannot express.
## Why a turn can hold several calls A model that has decided it needs stock levels for twelve stores has no reason to ask for them one at a time. When requests are independent — none of them needs another's answer to be formed — the model can emit several tool-call blocks in a single assistant turn. This collapses twelve round trips into one, which is usually the single largest latency win available in an agent loop, because each avoided round trip removes a full model inference from the critical path. Whether a model does this depends on the model, the tools, and sometimes a provider switch that disables the behaviour. It is a capability, not a guarantee: correct code handles one call and twelve identically. ## The return contract Three rules govern the way back, and all three are frequently broken: 1. **All calls must be answered.** You cannot answer eight of twelve and continue. The model's turn is a set of outstanding requests, and the provider generally rejects a continuation that leaves any of them dangling. 2. **The answers go in one turn.** You do not make a model call per result. One following turn carries all twelve result blocks; then the model runs once, seeing every answer at the same time. 3. **Pairing is by id.** Each result block repeats the id of the call it answers. Position is not the join key. Completion order is not the join key. ## The id-pairing bug The failure worth remembering is a runtime that fires twelve `check_stock` calls concurrently, appends each result to an array as it returns, and sends that array back. The results are now in completion order — fastest store first — while the calls are in the order the model emitted them. If the ids are also written from the loop index, or omitted where the provider tolerates it, the Leeds stock level is presented as Brighton's. What makes this vicious is that nothing errors. The shapes are valid, the token counts are normal, and the model produces a fluent, confident, wrong answer. It surfaces as "the assistant occasionally reports the wrong store", which reads like a model quality problem and is actually a plumbing bug. The fix is mechanical: carry the id from the call into the result on the same object, never through a parallel array, and assert before sending that the set of result ids equals the set of call ids. ## What you may not assume about ordering - **Not execution order.** Parallel calls carry no sequencing information. If the model emitted a write and a read of the same record in one turn, there is no defined order between them, and there is no way to express one. - **Not dependency safety.** The model asked for these together because it judged them independent. If they are not — because two of them mutate shared state — that is a tool-design and concurrency problem on your side, not something the protocol resolves. - **Not result ordering in context.** The model reads all the results together; do not encode meaning in their sequence. ## Concurrency is your choice Nothing obliges you to run parallel calls in parallel. You may execute them sequentially and still return them in one turn; the protocol only describes the message shape. Running them concurrently is where the wall-clock win comes from, so most runtimes do — with a bounded worker pool, per-tool timeouts, and a rule that one failing call does not abort the batch. A call that fails still needs a result block; a missing one breaks rule 1. ## When parallelism is the wrong shape Parallel calls are for fan-out over independent items. They are not a substitute for a plan. If the work is genuinely sequential — look up an order, then refund it, then email the customer — the model should be emitting one call per turn so it can react to each result. And if the fan-out is very wide (hundreds of items), a conversational turn carrying hundreds of results is often the wrong instrument entirely; that is where invoking tools from code inside a sandbox starts to win, because the loop never enters the context window. ## What good implementations do Keep call and result together as one object through the whole pipeline. Validate the id sets match before sending. Log the pairing, not just the payloads, so a mismatch is visible in traces. And test with deliberately staggered latencies so completion order differs from call order in your test suite — the bug does not reproduce when every stub returns instantly.
- One of twelve parallel calls times out. What do you send back?A result block for that call id as well, describing the failure, alongside the eleven successes in the same turn. Leaving it out breaks the contract that every outstanding call is answered, and the model then has no idea why one store is missing. Given the failure in context it can retry that one call, answer with eleven stores and say so, or escalate — all better than a dangling call or an aborted batch.
- Can you rely on a model emitting parallel calls when the work is parallelizable?No. Whether a turn fans out depends on the model, how the tools are described, and provider settings that can disable the behaviour, so it varies run to run. Write the loop to handle one call and twenty identically, and if wide fan-out is a hard requirement rather than an optimization, drive it from your own code — a fixed workflow or sandbox loop — instead of hoping the model batches.
- Two parallel calls in one turn write to the same record. What is your position?That the tools are badly scoped. The protocol expresses no ordering between calls in a turn, so there is no way to make the sequence deterministic once the model has emitted them. The fix sits in tool design: make the mutating operation a single tool that takes the whole change, or guard it so concurrent invocations are serialized or rejected, rather than trying to influence emission order through prompting.
saying these in an interview costs you the question
- Matching results to calls by array index
- Sending one model call per parallel result
- Assuming parallel calls execute in emitted order
- Dropping a failed call instead of returning a failure result
- Believing every model batches parallelizable work