What does adding Gemini's Google Search grounding tool change in the request and response?
answer
- One entry in the tools array
- Search runs server-side, no second turn
- Model decides whether to search at all
- groundingMetadata may be absent
- Recency and attribution, not truth
basics
~20 sYou add a google_search tool to the request's tools list. The model then decides on its own whether to search, and the returned candidate carries groundingMetadata: the queries it issued, the web sources it used, and the mapping from answer spans to those sources.
solid answer
~50 sOn the request side it is one entry in the `tools` array — `Tool(google_search=GoogleSearch())` in the google-genai SDK. Unlike function calling, you never see a call to execute: search runs inside Google's service, and the model decides per request whether it needs it. On the response side the candidate gains `groundingMetadata`, containing `webSearchQueries` (the searches the model actually ran), `groundingChunks` (each with a source `uri` and `title`), `groundingSupports` (spans of the answer text mapped to the chunk indices that back them), and `searchEntryPoint.renderedContent` — pre-built HTML for the Search Suggestions chip you are required to display. Note the metadata is *absent* when the model chose not to search, so treat it as optional. Grounding is billed per grounded request on top of tokens and adds latency, and it improves recency and attribution — it does not guarantee the answer is correct.
code
python · 21 linesfrom google import genai
from google.genai import types
client = genai.Client()
response = client.models.generate_content(
model="gemini-2.5-flash",
contents="What changed in the EU AI Act timeline this month?",
config=types.GenerateContentConfig(
tools=[types.Tool(google_search=types.GoogleSearch())],
),
)
candidate = response.candidates[0]
metadata = candidate.grounding_metadata
if metadata is None:
print("model answered without searching")
else:
print("queries:", metadata.web_search_queries)
for chunk in metadata.grounding_chunks or []:
print(chunk.web.title, chunk.web.uri)go deeper
Know that grounding is switched on by listing a google_search tool in the request and that the answer comes back with metadata naming the web sources the model consulted.
Explain that retrieval happens server-side with no tool-call round trip, that the model decides whether to search so the metadata is optional, and name the metadata's parts: queries, chunks, supports, search entry point.
Show the operating view: per-request grounding cost and latency, gating which turns get the tool, using webSearchQueries to debug bad answers, and treating grounding as attribution rather than a correctness guarantee.
Weigh managed grounding against a retrieval stack you own — control over sources, redaction, caching and data residency versus zero infrastructure — and set the policy for which product surfaces are allowed to answer from the open web at all.
## Turning it on Grounding with Google Search is enabled by listing it as a tool on the generation config: `config=types.GenerateContentConfig(tools=[types.Tool(google_search=types.GoogleSearch())])` On the REST body that is `"tools": [{"google_search": {}}]`. There is nothing else to configure in the simple case — no API key for search, no result count, no site filter. ## It is not function calling This is the point interviewers probe. With ordinary function declarations, the model emits a function call, *your* code executes it, and you send the result back in a second turn. With the search tool none of that happens: the retrieval executes server-side inside Google's infrastructure, and you get a single finished response. Your loop stays one round trip, and you cannot intercept, filter, or cache the search results. If you need control over retrieval — your own index, your own ranking, your own redaction — you want function calling or a retrieval pipeline you own, not this tool. The model also decides *whether* to search. A question about arithmetic or about the content already in the prompt typically triggers no search at all, and the response then carries no grounding metadata. Defensive code treats `grounding_metadata` as possibly `None`. ## What comes back When the model did search, the candidate carries `groundingMetadata` with four things worth knowing: - **`webSearchQueries`** — the actual query strings issued. This is the single best debugging field on the whole feature: when a grounded answer is wrong, the query the model chose usually explains why. - **`groundingChunks`** — the sources. Each web chunk has a `uri` and a `title`. The URI is a Google redirect link rather than the publisher's own URL, and it is documented as valid only for a limited window, so it is not a permanent citation. - **`groundingSupports`** — the attribution map: each entry has a `segment` (with start and end indices plus the text) and `groundingChunkIndices` pointing into the chunk list. This is what lets you render footnote markers against specific sentences rather than dumping a source list at the bottom. - **`searchEntryPoint`** — with `renderedContent`, an HTML/CSS snippet for the Google Search Suggestions chips. Displaying it is a condition of using the feature, not a styling suggestion. ## Version history worth knowing On the Gemini 1.5 generation the tool was named `google_search_retrieval` and accepted a `dynamicRetrievalConfig` with a mode and a `dynamicThreshold` — a number you tuned to control how eager the model was to search. From Gemini 2.0 onward the tool is simply `google_search` and the eagerness decision is internal to the model; there is no threshold to tune. If you are reading older sample code, that renaming is the reason it will not run against a current model, and it is a fair thing to be asked about. ## Cost, latency and limits Grounded requests are billed as a separate line item per request in addition to normal token cost, and there is a free daily allowance before that kicks in — so a chat product that grounds *every* turn spends materially more than one that grounds only when the user asks about current events. Latency rises too, because retrieval happens inside the request. A common design is to route: a cheap classification or an explicit user affordance ("search the web") decides whether to attach the tool for that call, rather than attaching it globally. There are also product constraints to check against current documentation rather than memory: availability varies by model family, and combining search grounding with other tools in one request has historically been restricted on some models. Verify for the model you actually deploy. ## What grounding does and does not buy you It buys **recency** — the model can answer about events after its training cutoff — and **attributability**, because you get sources you can show and a mapping from claims to those sources. It does not buy correctness. The model can still misread a retrieved page, blend two sources, or cite a source that does not actually support the sentence it is attached to. Grounded answers still need the same scepticism as ungrounded ones; what changes is that a reviewer now has a link to check against, which is exactly why teams in regulated domains value it. Treat the presence of grounding metadata as evidence the model consulted the web, not as a certificate that the answer is right.
- Why can't you inspect or filter the search results before the model uses them?Because retrieval executes inside Google's service as part of the same request — there is no tool-call round trip back to your code. You receive only the finished answer plus metadata about which sources were consulted. If you need to vet, redact, or restrict sources, use ordinary function calling against a retrieval backend you control, where you decide what goes back into the prompt.
- A grounded call returns no groundingMetadata at all. Is that an error?No. The model decides per request whether searching is warranted, and for arithmetic, rewriting, or questions answerable from the prompt it simply does not search. The candidate is a normal successful response with no metadata attached. Code must treat the field as optional; failing to do so turns a perfectly good answer into a null-reference crash on exactly the cheap requests.
- How does the older google_search_retrieval tool differ from google_search?google_search_retrieval was the Gemini 1.5-era tool and accepted a dynamicRetrievalConfig with a mode and a dynamicThreshold that let you tune how readily the model searched. From Gemini 2.0 the tool is named google_search and that decision moved inside the model, so there is no threshold parameter. Older sample code fails against current models purely because of the rename.
- How would you keep grounding cost under control in a chat product?Do not attach the tool to every turn. Grounded requests carry a per-request charge on top of tokens and add retrieval latency, so gate them: an explicit user affordance, a cheap classifier over the turn, or a rule that only questions mentioning current events or named entities get the tool. Log the grounded-request share so the cost line is attributable to a feature rather than to the whole assistant.
saying these in an interview costs you the question
- Thinks the search tool returns a tool call you must execute
- Assumes groundingMetadata is present on every response
- Believes grounding guarantees the answer is factually correct
- Expects to filter or cache the underlying search results
- Treats grounded requests as costing the same as ungrounded ones