When should a ReAct agent dispatch several tool calls in one action step?
answer
- independent reads, one step
- arguments must be knowable up front
- chained calls are inherently sequential
- no ordering, no rollback on writes
- cap fan-out in code, not in the prompt
basics
~20 sFan out only when the calls are independent reads whose arguments do not depend on each other's results. Keep writes sequential: parallel writes have no ordering guarantee, and a partial failure leaves inconsistent state with nothing to roll back.
solid answer
~50 sThe test is dependency. Five stock-level lookups for five different SKUs share nothing — no call's arguments come from another's result — so issuing them in one step costs one model round-trip and roughly the wall-clock of the slowest call, instead of five round-trips serialised end to end. That is the case parallel dispatch exists for. The moment one call's arguments depend on another's output, the work is inherently sequential in ReAct, because the next action is conditioned on the observation that has not arrived yet. Writes are the other boundary: parallel writes have no defined order, no transaction spanning them, and a partial failure leaves you with three of five applied and no rollback — so unless each write is idempotent and touches a disjoint key, dispatch them one at a time. Independently of correctness, cap the fan-out in the runtime; a model that decides to issue forty calls at once is a load-test against your own service.
code
json · 5 lines[
{"tool": "get_stock_level", "call_id": "c1", "args": {"sku": "A-100"}},
{"tool": "get_stock_level", "call_id": "c2", "args": {"sku": "B-200"}},
{"tool": "get_stock_level", "call_id": "c3", "args": {"sku": "C-300"}}
]go deeper
Know that several tool calls can be issued in one step only when none of them needs another's result, and that reads are the safe case while writes are not.
Explain what fan-out saves — one model round-trip and the wall-clock of the slowest call rather than the sum — and why a chained lookup is inherently sequential in a ReAct loop.
Demonstrate the write reasoning in production terms: no ordering, no transaction, partial failure with no rollback, unsafe retries without idempotency keys, and per-call attribution of results so a mixed outcome stays representable.
Own fan-out as a platform limit — concurrency caps and per-tool ceilings enforced in the runtime — because the blast radius of an unbounded agent against internal services scales with every user running it at once.
## What fanning out actually buys A ReAct turn is expensive in two currencies. There is the model call — latency to first token plus generation, on every step — and there is the tool call itself. Sequential dispatch pays both, per tool, in series: five lookups means five model round-trips and five tool latencies stacked end to end. Issuing all five in one action step collapses that to one model round-trip and, if the runtime executes them concurrently, roughly the wall-clock of the slowest call. On a fan-out of five reads against a service with a 200 ms tail, this is the difference between a few seconds and a few hundred milliseconds, and it removes four rounds of generated reasoning tokens. On an agent doing dozens of independent lookups per task, it is the difference between usable and abandoned. ## The dependency test Calls can go together when none of them needs another's result to be *formed*. Five `get_stock_level` calls, one per SKU, qualify: the arguments were all knowable before any call ran. Reading three unrelated config files qualifies. Querying two independent services for a status qualifies. The opposite is a chain: look up a customer to get their account id, then read that account's balance. The second call's arguments literally do not exist until the first returns. This is not a limitation of parallel dispatch — it is inherent to the ReAct shape, where the next thought is conditioned on the observation. Trying to fake it (guessing the id and issuing both) is how agents produce confidently wrong work. A subtler dependency is *order sensitivity without argument dependency*. Two calls may both be formable up front and still be wrong to interleave, because one changes what the other observes. That is really the write case. ## Why writes are different Parallel writes fail in ways parallel reads do not. **No ordering.** Concurrent calls complete in whatever order the network and the service decide. If the outcome depends on order — two updates to the same record, a create and its dependent update — you have a race with no arbiter. **Partial failure with no rollback.** Five writes, three succeed, one times out, one is rejected. There is no transaction spanning them, so you are left in a state no one designed, and the agent must now reason about a mixed outcome. Sequential dispatch does not eliminate partial failure, but it makes the failure point unambiguous and lets the loop stop at the first one. **Retry hazard.** Recovery from a timeout means resending, and a timeout does not tell you whether the write landed. Without idempotency keys, a retried parallel write can double-apply. **Shared-resource contention.** Simultaneous writes are exactly what triggers lock contention, deadlocks and per-key rate limits — the failures that only appear under concurrency, and only in production. The practical rule: parallelise reads freely, serialise writes by default, and allow parallel writes only when each one is idempotent and touches a disjoint key — five per-SKU cache entries, say, rather than five ledger entries. ## Fan-out is a runtime concern, not a prompt one You can ask the model to be restrained about how many calls it issues, and it will mostly comply, and then one day it will not. Bound it in code: a maximum number of calls accepted per action step, a semaphore capping in-flight concurrency, and per-tool limits for anything backed by a rate-limited dependency. An agent that fans out forty calls against an internal API is a self-inflicted load test, and the blast radius grows with every user running the agent at once. The same reasoning applies to cost. Fan-out multiplies tool-side spend within a single step, where a per-step review would have caught it. ## Reporting a batch honestly When several calls go out together, each result must come back attributable to the call that produced it — a stable identifier per dispatched call — so that "three succeeded, one timed out, one was rejected" is representable rather than collapsing into one blurred blob. Losing that attribution is the most common implementation bug in fan-out, and it turns a recoverable partial failure into an agent that cannot tell which SKU it failed to price. ## Judgement, honestly stated There is no universal fan-out policy. It depends on how expensive your tools are, how tolerant your downstreams are, and how much a wrong-order side effect costs in your domain. The defensible position in an interview is a *rule* with a *reason*: reads fan out, writes serialise unless proven idempotent, concurrency is capped in the runtime, and results stay individually attributable. Candidates who answer "parallelise everything, it's faster" have not operated an agent against a system that can be damaged.
- Three of five parallel calls succeed and two fail. What does the runtime owe the agent?Individually attributable results. Every dispatched call needs a stable identifier so each outcome — success, timeout, rejection — comes back tied to the call that produced it. Collapsing a batch into one blended result is the classic fan-out bug: the agent knows something went wrong but not which SKU it failed to price, and it can neither report accurately nor act on the gap.
- When are parallel writes actually acceptable?When each write is idempotent and touches a disjoint key, so order cannot matter and a retry cannot double-apply. Writing five independent per-SKU cache entries is fine; posting five ledger entries is not. Idempotency keys are what make the retry safe, and disjoint keys are what make ordering irrelevant — you need both properties, not either one alone.
- How do you stop a model from fanning out forty calls in one step?In the runtime, not the prompt. Cap the number of calls accepted per action step, put a semaphore on in-flight concurrency, and add per-tool limits for anything behind a rate-limited dependency. Prompt guidance shapes typical behaviour but is not an enforcement mechanism, and the day it is ignored is the day you load-test your own internal service from production.
- Does fanning out always reduce latency?No. If the tools share a bottleneck — one database, one rate-limited upstream — concurrent calls queue behind each other and you have converted a predictable series into contention, sometimes with worse tail latency than sequential execution. The win is real when the calls hit genuinely independent capacity; measure rather than assume, especially where a connection pool or a per-key limit sits underneath.
saying these in an interview costs you the question
- Parallelising calls whose arguments come from an earlier result
- Fanning out writes because reads parallelise safely
- Treating a partial batch failure as if the whole step failed
- Relying on prompt instructions to bound the fan-out
- Assuming concurrency always helps, even behind a shared bottleneck