How do you return a function result to Gemini and continue the tool-use turn?
answer
- The endpoint is stateless — resend history
- Append the model's call turn verbatim
- from_function_response with the same name
- Payload must be an object, not a scalar
- Return errors, do not raise out
basics
~20 sAppend the model's own turn containing the FunctionCall to the conversation, then append a turn whose part is a FunctionResponse carrying the same function name and a JSON-object result. Re-send the whole history with the same tools and loop until no more calls appear.
solid answer
~40 sThe continuation is history-based: Gemini has no server-side thread, so you rebuild `contents` yourself. Take `response.candidates[0].content` — the model turn holding the `function_call` — and append it verbatim, then append a new `Content` whose part is `types.Part.from_function_response(name=..., response={...})`; the SDK examples use role `"user"` for that turn. Two rules bite in practice: the `name` must match the call exactly, and `response` must be a JSON **object**, so wrap scalars as `{"result": value}`. Then call `generate_content` again with the full `contents` and the same tool declarations still attached — dropping them mid-loop degrades the model's grasp of what it just did. Repeat until a turn contains no function calls, and bound the loop with a maximum iteration count so a stubborn model cannot spin forever.
code
python · 31 linesfrom google import genai
from google.genai import types
client = genai.Client()
config = types.GenerateContentConfig(
tools=[types.Tool(function_declarations=[get_weather_decl])],
automatic_function_calling=types.AutomaticFunctionCallingConfig(disable=True),
)
contents = [types.Content(role="user", parts=[types.Part(text="Weather in Oslo?")])]
for _ in range(5):
response = client.models.generate_content(
model="gemini-2.5-flash", contents=contents, config=config
)
calls = response.function_calls
if not calls:
print(response.text)
break
contents.append(response.candidates[0].content)
parts = []
for call in calls:
try:
result = DISPATCH[call.name](**dict(call.args))
payload = {"result": result}
except Exception as exc:
payload = {"error": type(exc).__name__, "message": str(exc)[:200]}
parts.append(
types.Part.from_function_response(name=call.name, response=payload)
)
contents.append(types.Content(role="user", parts=parts))go deeper
Remember the sequence: append the model's call turn, append a function-response turn with the same function name, then call the model again with the full history.
Explain that generateContent is stateless, that the response payload must be a JSON object, and that the tool declarations must stay on the continuation request.
Demonstrate production habits: structured error payloads instead of raised exceptions, truncated tool results, iteration ceilings, and per-call tracing.
Own the loop's cost and safety envelope — context growth per iteration, budget caps, timeouts, and how a shared abstraction translates this shape to other vendors' tool protocols.
## Why you rebuild the history The Gemini `generateContent` endpoint is stateless. Whatever the model is allowed to know about the conversation must be present in the `contents` list you send. That has an immediate consequence for tool use: the turn where the model asked for a function is not remembered for you. You must append it, and then append your answer to it, so that the next request reads as a coherent four-beat exchange — user asks, model calls, tool answers, model concludes. ## The exact four steps 1. **Detect the call.** Iterate the response parts (or use `response.function_calls`) and collect every `function_call`. 2. **Append the model turn.** Push `response.candidates[0].content` onto your `contents` list unchanged. Do not reconstruct it by hand — round-tripping the object the SDK gave you avoids subtle mismatches. 3. **Append the tool turn.** Build a part with `types.Part.from_function_response(name=call.name, response={...})` and wrap it in a `Content`; the documented examples use role `"user"` for this turn. 4. **Re-send.** Call `generate_content` again with the full `contents` and the same `config` (tools included), then check the new response for further calls. When the turn contained several calls, all their responses go into that one tool turn together, one part each. ## The two shape rules people get wrong **The name must match.** The `name` on the response is how the model pairs the result with its request. A typo, a namespaced variant, or a renamed tool between turns leaves the call effectively unanswered and the model will often just ask again. If the call carried an `id`, echo it too. **The response payload is an object.** `response` is a structured value, and a bare string or number is not a valid payload — wrap it: `{"result": 21.5}`, `{"orders": [...]}`, `{"error": "..."}`. This also gives you room to add metadata the model can use, such as units, freshness or truncation flags. ## Returning failures instead of raising The worst thing you can do when a tool throws is to abort the whole request. The model is perfectly capable of recovering — retrying with corrected arguments, choosing another tool, or telling the user the data is unavailable — but only if it hears about the failure. Catch the exception and return a structured error as the function response: `{"error": "unknown_customer", "message": "No customer with id cus_9999"}`. Keep error payloads short, machine-readable and free of stack traces or internal identifiers you would not want echoed to the user. The same reasoning applies to large results. A tool that returns fifty kilobytes of JSON will consume context on every subsequent turn of the loop, because the history keeps growing. Truncate, paginate, or summarise, and say in the payload that you did (`{"items": [...], "truncated": true}`) so the model does not conclude the list was complete. ## Bounding the loop A tool-use loop is a `while` loop driven by a probabilistic process, so give it a hard ceiling — typically five to ten iterations for interactive work. On each pass, log the call name and duration; those traces are the only way to debug "it called `search_orders` eleven times" after the fact. If the ceiling is hit, break out and produce a graceful answer rather than continuing to burn tokens. Cost grows quadratically-ish across the loop, because every iteration re-sends the entire accumulated history plus all declarations. That is why trimming tool results and keeping the tool list small matter more here than in a single-shot call. ## Where this differs from other vendors The *idea* is universal, the *shape* is not. Gemini uses content parts (`function_call` / `function_response`) matched by function name inside a stateless `contents` list. OpenAI-shaped APIs use a message with a `tool_calls` array and reply with separate `role: "tool"` messages keyed by `tool_call_id`; Anthropic uses `tool_use` and `tool_result` blocks keyed by a block id. Code that abstracts over several providers must translate all three, and the pairing key differs in each — name here, ids there. ## A working shape The loop is small enough to write once and reuse: detect calls, dispatch through a name-to-callable map, wrap outcomes as objects, append both turns, re-send, and stop on either "no calls" or the iteration ceiling.
- What happens if you send the function response but forget to append the model's function-call turn?The history becomes incoherent: a result appears with nothing that requested it. In practice the model either ignores the orphaned response and re-issues the call, or answers from thin air. Always push the candidate content object you received, unmodified, before appending your response turn.
- Your tool throws a timeout. What do you send back?A structured error as the function response — for example {"error": "timeout", "retryable": true} — rather than aborting the request. The model can then retry with a narrower query, fall back to another tool, or tell the user the service is unavailable. Keep the payload short and free of internal stack detail.
- Should the tool declarations stay attached on the continuation request?Yes. Because the request is stateless, dropping tools mid-loop means the model no longer sees the definition of the function it just called, which degrades its interpretation of the result and prevents any follow-up call. Keep the same config across the loop, and change the tool set deliberately, not accidentally.
- How do you keep a long tool loop from exhausting the context window?Trim what you return, not just what you send: truncate or summarise large results, flag truncation in the payload, drop superseded tool results from the history once the model has used them, and cap loop iterations. Every pass re-sends the accumulated history, so unbounded results compound quickly.
saying these in an interview costs you the question
- Sending a bare string as the function response payload
- Dropping the model's call turn from the history
- Raising the tool exception out of the loop
- Expecting the API to remember the conversation
- Looping without any iteration ceiling