Gemini returned three FunctionCall parts in one turn — how must your app respond?
answer
- Independent requests, one candidate
- Answer all of them, none skipped
- One turn, several response parts
- Match each response by function name
- Fan out concurrently, bound the fan-out
basics
~10 sExecute all three, then send one turn containing three FunctionResponse parts, each naming the function it answers. Splitting them across turns or answering only some leaves calls unresolved and the model typically re-issues them.
solid answer
~50 sGemini can emit multiple `function_call` parts in a single candidate when the requests are independent. The contract is symmetric: append the model's turn once, then append a single `Content` whose `parts` list holds one `types.Part.from_function_response` per call, matched by function name (and by `id` when the call carries one). Because they are independent, you can dispatch them concurrently and cut wall-clock latency to the slowest tool rather than the sum — that is the main reason parallel calls exist. Two operational rules matter: never drop a call, even a failed one, because an unanswered request leaves the history incomplete and the model will usually ask again; and cap the fan-out, since three concurrent tools that each hit a database can amplify load. If two calls target the same non-idempotent function, treat that as a signal to deduplicate rather than to execute twice.
code
python · 23 linesimport asyncio
from google.genai import types
async def answer_calls(calls, dispatch, max_parallel=4):
sem = asyncio.Semaphore(max_parallel)
async def run(call):
async with sem:
try:
value = await asyncio.wait_for(
dispatch[call.name](**dict(call.args)), timeout=10
)
return types.Part.from_function_response(
name=call.name, response={"result": value}
)
except Exception as exc:
return types.Part.from_function_response(
name=call.name,
response={"error": type(exc).__name__},
)
parts = await asyncio.gather(*(run(c) for c in calls))
return types.Content(role="user", parts=list(parts))go deeper
Know that Gemini can ask for several functions at once and that every one of them needs a matching response before the conversation continues.
Explain the shape: one model turn holding several function_call parts, answered by one turn holding one function_response part per call, matched by name.
Show that you dispatch concurrently with per-call timeouts, always return a payload for failures, and reason about idempotency and fan-out load.
Own the design tradeoff between fine-grained tools that fan out and coarser tools that guarantee ordering or atomicity, plus the load and observability policy around fan-out.
## Why one turn can hold several calls When a user asks "compare the weather in Oslo and Lisbon", nothing about the second lookup depends on the first. Gemini can recognise that and emit two `function_call` parts in one candidate rather than doing them one after another. Each part is a complete, independent request with its own `name` and `args`. This is distinct from *sequential* tool use, where the second call's arguments depend on the first call's result. Sequential chains still take one round trip each; parallel calls collapse a fan-out into a single round trip. ## The response contract The reply shape mirrors the request shape: - Append `response.candidates[0].content` — the single model turn carrying all the calls — exactly once. - Build one function-response part per call and put **all of them in one `Content`**, whose `parts` list has the same cardinality as the calls. - Match each response to its call by `name`; echo the `id` when the call supplied one. A common bug is to loop over the calls and append a separate `Content` per response. Some of the time the model copes; the shape you actually want is one turn in, one turn out, so keep the batching symmetric. A second common bug is partial answering — skipping a call whose tool threw. Return a structured error payload for it instead: `{"error": "timeout"}`. An error is an answer; silence is not, and the model will typically re-issue the unanswered call on the next turn, wasting a round trip. ## Concurrency is the point If you execute the calls serially you have kept the round-trip saving but thrown away the latency saving, which is usually the larger prize. Dispatch them with a thread pool, an async gather, or whatever your runtime offers, with a per-call timeout so one slow tool cannot hold the whole turn hostage. In an async application, an `asyncio.gather` over the calls with `return_exceptions=True` maps neatly onto "every call gets a payload, successes and failures alike". With concurrency come the ordinary concurrency concerns: - **Fan-out amplification.** A model that emits eight calls, each hitting the same downstream service, multiplies your load by eight for that turn. Cap the number you execute concurrently, and consider rejecting calls beyond a limit with an explanatory error payload. - **Idempotency.** Read-only lookups parallelise safely. Writes do not. If a turn contains two calls to a non-idempotent function — the same charge issued twice, say — deduplicate or reject rather than executing both; the model is not a transaction manager. - **Ordering.** Do not rely on execution order between the calls. If order matters, they were not independent, and the safer design is to declare a single coarser tool that performs the ordered work internally. - **Shared state.** Concurrent handlers touching the same cache or connection need the same discipline as any other concurrent code path. ## Keeping the results manageable Parallel calls multiply payload size in one go. Three tools each returning a large JSON document land in the history together, and stay there for the rest of the conversation. Truncate per-call, flag truncation in the payload, and think about summarising results before they enter the history when the raw form is not needed later. ## Debugging Trace each call individually — name, arguments, latency, outcome — and tag them with the turn index. Aggregate metrics hide the interesting failure, which is usually "one of the three timed out and the model quietly changed its answer". Also log the count of calls per turn; a sudden rise often means a prompt or declaration change made the model fan out where it used to chain. ## Provider differences The capability is near-universal, the encoding is not. Gemini expresses it as several `function_call` parts inside one content, paired by function name. OpenAI-shaped APIs return several entries in a `tool_calls` array and expect one `role: "tool"` message per call, keyed by `tool_call_id`. Anthropic returns several `tool_use` blocks and expects the matching `tool_result` blocks in one user message. If you write an adapter, the invariant to preserve is "every request gets exactly one response, in the vendor's expected grouping".
- One of the three tools times out. Do you still send a response part for it?Yes. Send a structured error payload such as {"error": "timeout", "retryable": true} for that call while sending real results for the others. Omitting it leaves a request unanswered, and the model typically re-issues it, costing an extra round trip. It also lets the model tell the user which part of the answer is missing.
- When would you prefer one coarser tool over letting Gemini fan out?When the calls are not truly independent — when order, a shared transaction, or a consistent snapshot matters. Fan-out gives you no ordering or atomicity guarantees, so encapsulate the sequence behind a single declaration that performs the ordered work internally and returns one consolidated result.
- How do you stop parallel calls from amplifying load on a shared backend?Bound the concurrency you actually execute rather than trusting the model's count: cap parallel dispatch, apply per-call timeouts, and reject calls beyond the cap with an explanatory error payload the model can reason about. Alert on calls-per-turn, since a rising fan-out usually traces to a prompt or declaration change.
saying these in an interview costs you the question
- Sending each function response in its own turn
- Silently skipping a call whose tool failed
- Executing the parallel calls one after another
- Assuming the calls run in the order listed
- Running non-idempotent duplicated calls twice