skip to content

What is automatic function calling in the Gemini Python SDK, and when do you disable it?

level: middleimportance: should knowfreq 40%

answer

  1. Callables in tools, not declarations
  2. Runs in your process, not Google's
  3. Docstrings become the descriptions
  4. maximum_remote_calls bounds the hops
  5. Side effects want a human seam

basics

~20 s

When you pass plain Python callables as Gemini tools, the SDK derives the declarations from their signatures and then executes them locally on your behalf, looping until the model produces text. Disable it with AutomaticFunctionCallingConfig(disable=True) whenever you need approval, custom error handling or tracing.

solid answer

~40 s

Automatic function calling (AFC) is a **client-side** convenience in the `google-genai` SDK, not a server feature. If `config.tools` contains Python functions rather than `types.FunctionDeclaration` objects, the SDK builds declarations from their signatures and docstrings, and when the model responds with a call it invokes the function in your process, appends the response, and re-requests — returning only the final answer. The number of automatic round trips is bounded by `maximum_remote_calls` (10 by default), and you switch it off with `automatic_function_calling=types.AutomaticFunctionCallingConfig(disable=True)`, which makes the raw `function_call` parts surface to you again. Disable it when you need a human approval gate before side effects, per-tool timeouts and retries, tracing and metrics per call, argument validation beyond the schema, or concurrent dispatch of parallel calls. Passing `FunctionDeclaration` objects instead of callables also leaves execution entirely to you.

code

python · 27 lines
python
from google import genai
from google.genai import types

client = genai.Client()

def celsius_to_fahrenheit(celsius: float) -> float:
    """Convert a temperature in Celsius to Fahrenheit."""
    return celsius * 9 / 5 + 32

# AFC on: the SDK derives the schema and runs the function locally
auto = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="What is 21.5 C in Fahrenheit?",
    config=types.GenerateContentConfig(tools=[celsius_to_fahrenheit]),
)
print(auto.text)

# AFC off: raw function_call parts come back to you
manual = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="What is 21.5 C in Fahrenheit?",
    config=types.GenerateContentConfig(
        tools=[celsius_to_fahrenheit],
        automatic_function_calling=types.AutomaticFunctionCallingConfig(disable=True),
    ),
)
print(manual.function_calls)

go deeper

for a junior

Know that passing Python functions as Gemini tools lets the SDK build the declarations and run the functions for you, and that a config flag turns that off.

for a middle

Explain that AFC is client-side, that declarations come from signatures and docstrings, and that maximum_remote_calls bounds the automatic hops.

for a senior

Argue when to disable it: approval gates for side effects, per-call tracing, timeouts and retries, argument validation, and concurrent dispatch of parallel calls.

for a principal

Own the policy split — which application paths may use the convenience and which must run an explicit loop — plus the portability cost of depending on an SDK-only behaviour.

## What AFC actually is The manual loop — detect calls, dispatch, append a function response, re-request — is mechanical, so the `google-genai` SDK will write it for you. Pass a plain Python function in `config.tools` and two things happen. First, **declaration derivation**: the SDK inspects the callable's signature, type annotations and docstring, and synthesises the `FunctionDeclaration` that would otherwise be hand-written. Your annotations become the parameter schema; your docstring becomes the description the model uses to decide when to call it. That makes annotations and docstrings load-bearing production artefacts rather than documentation. Second, **automatic execution**: when the model responds with a function call, the SDK calls your function in your process with the model's arguments, wraps the return value as a function response, appends both turns to the history, and issues the next request — repeating until the model answers in text. What comes back to your code is the final response, with the intermediate hops having happened invisibly. Crucially, this is *client-side*. Nothing about your function reaches Google. The execution model has not changed at all; only the authorship of the loop has. ## The controls AFC is configured through `automatic_function_calling` on the generation config, built from `types.AutomaticFunctionCallingConfig`: - `disable=True` turns it off entirely, so raw `function_call` parts come back to you and you write the loop. - `maximum_remote_calls` bounds how many automatic round trips the SDK will perform before it stops — 10 by default. This is the SDK's guard against a model that keeps calling forever, and it is not a substitute for your own budget checks. Note also that AFC only engages for **callables**. If you pass `types.FunctionDeclaration` objects, there is nothing for the SDK to execute, so you get the calls back regardless of the flag. ## When the convenience is right - Prototypes, notebooks and internal scripts, where the tools are pure reads and the failure mode is a stack trace you will read anyway. - Small, side-effect-free helpers — arithmetic, formatting, a lookup in an in-process dict. - Demos and evaluation harnesses where loop code would obscure the point. ## When to disable it **Side effects need a gate.** If a tool sends an email, charges a card, or writes to a database, you almost certainly want a decision point between "the model asked" and "it happened" — a policy check, an allow-list, or a human confirmation. AFC removes that seam. **Observability.** Production systems need per-call traces: which tool, which arguments, how long, what outcome. Writing the loop yourself is where those spans are emitted. Debugging a system whose tool hops are invisible is miserable. **Error semantics.** You will want timeouts, bounded retries, circuit breakers, and structured error payloads that let the model recover gracefully. Owning the loop is what makes those possible. **Concurrency.** When Gemini emits several independent calls in a turn, you want them dispatched concurrently with a bounded fan-out. Hand-rolled loops make that explicit. **Validation.** The schema constrains shape, not intent. Arguments derived from user text may be adversarial, so you want an explicit validation and authorization step before dispatch. **Cost control.** Each automatic hop is a full billed request carrying the whole accumulated history. Your own loop can enforce a token or currency budget and stop early, which `maximum_remote_calls` only approximates. **Streaming and UX.** If you want to show the user "looking up your orders…" between hops, you need to see the hops. ## A middle path You are not forced to choose globally. Because the setting lives on the request config, you can enable AFC for a read-only assistant path and disable it for the path that can mutate data, sharing the same client. Some teams also keep the callables as the single source of truth — for schema derivation — while disabling automatic execution, so declarations stay in sync with the code without giving up the loop. ## Portability note AFC is a property of the SDK, not of the Gemini API. Raw REST calls, and other language SDKs, do not necessarily give you the same behaviour, so code that depends on it does not transfer cleanly. If you are building an abstraction over several providers, write the loop explicitly — it is the only shape all of them share.

  • If AFC executes the function, does any of your code run on Google's servers?
    No. AFC is purely client-side: the SDK writes the tool loop and calls your function in your own process. Google's servers only ever see the function declaration and whatever you send back as a function response. The execution and trust boundary is unchanged — only the authorship of the loop moved into the SDK.
  • What happens to your docstrings and type annotations when you pass callables as tools?
    They become the declaration. The SDK derives the parameter schema from annotations and the model-facing description from the docstring, so vague prose or missing annotations directly degrade tool selection and argument quality. Treat them as production artefacts and review them the way you would review a hand-written schema.
  • How would you keep AFC's schema convenience without its automatic execution?
    Keep the callables as the source of truth but set automatic_function_calling with disable=True, or derive declarations from your typed models at startup. You still get schemas that cannot drift from the implementation, while the raw function_call parts come back to you so you can validate, gate, trace and dispatch concurrently.

saying these in an interview costs you the question

  • Thinking the SDK executes tools on Google's servers
  • Leaving AFC on for tools with real side effects
  • Assuming REST behaves the same as the Python SDK
  • Treating maximum_remote_calls as a cost budget
  • Expecting AFC when passing FunctionDeclaration objects

context