skip to content

In Google ADK, how do you decide which OpenAPIToolset operations an agent gets?

level: principalimportance: should knowfreq 34%

answer

  1. One tool per operation, not per endpoint
  2. Declarations ride on every request
  3. Filter before you ship
  4. Split the API across agents
  5. Hand-write the journeys that matter

basics

~20 s

OpenAPIToolset generates one callable REST tool per operation in the spec, so a large API becomes hundreds of tools. Curate deliberately: filter the toolset down to the operations an agent actually needs, and split the rest across separate agents rather than loading one agent with everything.

solid answer

~50 s

`OpenAPIToolset(spec_str=..., spec_str_type="json")` parses a spec and produces one REST tool per operation, named from its `operationId` and described from the operation's summary and description; auth is configured once on the toolset with `auth_scheme` and `auth_credential` rather than per tool. That generosity is the problem — every tool's declaration is serialized into every model request, so a 200-operation spec means a large fixed prompt cost on every turn plus a much harder choice for the model. The levers are: a `tool_filter` naming the operations this agent is allowed to use; splitting the API by domain across several agents, each reached through `AgentTool` or as a sub-agent; and hand-writing a small function tool for the two or three journeys that matter, where a purpose-written docstring beats a generated one. Generated coverage is a starting point, not a shipping configuration.

code

python · 21 lines
python
from google.adk.agents import Agent
from google.adk.tools.openapi_tool.openapi_spec_parser.openapi_toolset import (
    OpenAPIToolset,
)

with open("billing_openapi.json") as f:
    spec = f.read()

billing_tools = OpenAPIToolset(
    spec_str=spec,
    spec_str_type="json",
    tool_filter=["list_invoices", "get_invoice", "download_invoice_pdf"],
)

billing_agent = Agent(
    name="billing_agent",
    model="gemini-2.0-flash",
    description="Answers questions about invoices and billing history.",
    instruction="Use the billing tools to answer invoice questions. Never guess amounts.",
    tools=[billing_tools],
)

go deeper

for a junior

Know that pointing the toolset at a spec produces one callable tool per operation, and that you normally hand an agent a filtered subset rather than the whole thing.

for a middle

Explain the generation details — operation ids as names, spec prose as descriptions, auth configured on the toolset — and why every exposed declaration costs tokens on every turn.

for a senior

Show the curation plan: filter to the journeys an agent supports, split domains across agents reached via AgentTool, fence destructive operations, and hand-write the hot paths.

for a principal

Own the tradeoff and the process — generated coverage versus hand-written fit, exposed tool lists treated as versioned configuration, spec prose quality as prompt quality, and evaluation updated whenever the surface changes.

## What the toolset generates Give `OpenAPIToolset` a spec string and a spec type and it walks every path and method, producing one callable REST tool per operation. Each tool's name comes from the operation's `operationId`, its parameters from the operation's path, query and body parameters, and its description from the operation's summary and description text. Authentication is configured once on the toolset — an `auth_scheme` plus an `auth_credential` — instead of being threaded through each generated tool. Then, at call time, ADK builds and issues the HTTP request for you. For a ten-endpoint internal service that is close to magic: your API becomes agent-callable in four lines. The question is what happens at two hundred. ## Why "just add the whole spec" fails A toolset is not a lazily-consulted catalog. The declarations of the tools an agent holds are serialized into the request sent to the model on **every turn**. So a 200-operation spec means: - **A fixed token tax per turn.** Hundreds of names, parameter schemas and descriptions ride along with every user message, before any actual conversation. That is money and latency on turns where no tool is called at all. - **A much harder decision.** The model must pick one operation from a list where a dozen entries have near-identical descriptions inherited from an API written for humans with a reference manual open, not for a model choosing blind. - **Blast radius.** Every exposed operation is an operation the agent can invoke. If the spec includes destructive endpoints, they are now reachable through natural language. - **Descriptions you did not write.** Generated descriptions are as good as the spec's prose. Most internal specs have terse or missing summaries, and the model reads exactly that. ## The levers, in the order I reach for them **1. Filter to the job.** A toolset accepts a tool filter, so the same generated spec can be narrowed per agent to the operations that agent's job requires. Start from the user journeys the agent must support, list the operations those journeys need, and expose only those. It is normal for a 200-operation API to yield an eight-tool agent. **2. Split by domain across agents.** When several genuinely distinct jobs live in one API — billing, provisioning, reporting — build one agent per domain, each with its own filtered toolset, its own instruction, and its own description. Expose them to a coordinating agent as `AgentTool`s, or arrange them as sub-agents when a hand-off is the right shape. The coordinator now chooses between three well-described specialists instead of 200 near-identical operations, and each specialist's prompt stays small. **3. Hand-write the hot paths.** For the two or three flows that carry most of the traffic, a purpose-written function tool usually beats the generated one: you can name it after the user's intent rather than the endpoint, write a docstring aimed at the model, collapse a three-call sequence into one, apply defaults and validation, and trim the response to what matters instead of returning whatever the API's full payload happens to be. **4. Fence the destructive operations.** Keep writes and deletes out of the default set. Where they are genuinely needed, put them behind a human gate — the long-running-approval shape — or on a separate agent that is only reachable in an explicit flow. ## Generated versus hand-written: the honest tradeoff Generated tools win on **coverage and drift**: when the API changes, regenerating from the spec is free, and there is no hand-written wrapper to fall out of date. Hand-written tools win on **fit**: intent-shaped naming, model-facing documentation, composed multi-call flows, trimmed responses, and validation. The mature answer is both — generation as the substrate so nothing is unreachable, curation on top so the agent's visible surface is small and well-described — with a bias towards hand-writing anything the agent gets wrong twice. The same reasoning applies to any imported toolset. When tools come from an MCP server, ADK's MCP toolset takes connection parameters for the server plus a tool filter, and manages the connection's lifecycle for you; the curation question is identical — a server offering forty tools does not mean your agent should see forty. Pick per agent. ## Operational consequences to plan for - **Toolsets are dynamic.** A toolset resolves the tools it contributes when the request is being built, not once at import, so the visible surface can legitimately differ between agents and over time. That is power and a hazard: an agent whose tool list changes underneath it is harder to reason about and to evaluate. - **Evaluation must follow the surface.** If you change which operations an agent can see, your tool-trajectory expectations change with it. Treat the exposed tool list as versioned configuration, not an incidental detail. - **Spec quality becomes prompt quality.** Once generated descriptions are what the model reads, improving the OpenAPI summaries is prompt engineering. That is often the cheapest available fix, and it improves the human docs at the same time. ## What interviewers listen for That you know generation is per operation and that declarations cost tokens on every turn; that your first instinct is to curate rather than to expose everything; that splitting across agents is a real architectural lever, not a workaround; and that you can state the generated-versus-handwritten tradeoff without pretending one side always wins.

  • When is a hand-written function tool worth the maintenance over a generated one?
    When the fit matters more than the coverage: a flow that needs three API calls in sequence, an endpoint whose spec description is useless to a model, a response that must be trimmed before it enters the conversation, or an intent that does not map one-to-one onto an endpoint. Keep generation underneath so nothing is unreachable, and hand-write the handful of journeys that carry the traffic or that the agent keeps getting wrong.
  • How does splitting one API across several agents actually help?
    It replaces one impossible choice with two easy ones. The coordinator picks between a few specialists described in intent terms, and each specialist picks from a short list within its own domain. Each prompt stays small, each agent can have an instruction tuned to its job, and permissions can differ per agent so destructive operations live somewhere reachable only by an explicit flow.
  • What changes when the tools come from an MCP server instead of an OpenAPI spec?
    The curation question does not change: a server offering forty tools does not mean your agent should see forty, so you filter per agent the same way. What changes is lifecycle — the toolset holds a connection to a running server, configured with that server's connection parameters, and ADK manages establishing and closing it. Availability of the server becomes an operational dependency of the agent.

saying these in an interview costs you the question

  • Assumes unused generated tools cost nothing until called
  • Points a single agent at an entire enterprise spec
  • Thinks one tool is generated per path rather than per operation
  • Treats generated descriptions as good enough for a model to choose from
  • Exposes destructive operations by default because they are in the spec

context