skip to content

Google Gemini

Google's Gemini API, reached through the GenAI SDK. Its distinctive interview material is native multimodality over image, audio and video, a very large context window with explicit caching, and answers grounded in Google Search.

on this pageshow

explore

questions

page 1 of 2

In Google's Gemini API, what does the model return when it decides to call a declared function?

level: juniorimportance: must knowfreq 70%

answer

  1. The model asks, it never runs
  2. Look inside the candidate's parts list
  3. name plus args, not a string
  4. args arrives already parsed
  5. Several calls can share one turn

basics

~20 s

Gemini returns an ordinary response whose candidate content holds a FunctionCall part: a function name plus an args object of already-parsed arguments. The model never runs your code — your application executes the call and sends the result back.

solid answer

~40 s

Function calling in Gemini is a *request*, not an execution. You call `generate_content` with your `FunctionDeclaration` tools, and if the model wants a tool it emits a part whose `function_call` field carries `name` and `args`. In the `google-genai` SDK, `args` is already a parsed dict — you do not parse a JSON string as you would with some other vendors — and `response.function_calls` is a convenience list of every call in the turn. A turn's `parts` list can mix text and function calls, and Gemini may emit several calls at once, so always iterate `response.candidates[0].content.parts` rather than assuming a single part. `response.text` may be empty when the turn is a pure tool request. Nothing runs on Google's servers: you dispatch to your own code and continue the conversation with a FunctionResponse.

code

python · 27 lines
python
from google import genai
from google.genai import types

client = genai.Client()

get_weather = types.FunctionDeclaration(
    name="get_weather",
    description="Get the current weather for a city.",
    parameters=types.Schema(
        type=types.Type.OBJECT,
        properties={"city": types.Schema(type=types.Type.STRING)},
        required=["city"],
    ),
)

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="What is the weather in Oslo?",
    config=types.GenerateContentConfig(
        tools=[types.Tool(function_declarations=[get_weather])],
        automatic_function_calling=types.AutomaticFunctionCallingConfig(disable=True),
    ),
)

for part in response.candidates[0].content.parts:
    if part.function_call:
        print(part.function_call.name, dict(part.function_call.args))

go deeper

for a junior

Be able to say that Gemini returns a FunctionCall part with a name and an args object, and that your own code runs the function and sends the result back.

for a middle

Explain the traversal: candidates, content, parts, and the function_call field; note that args is already decoded and that a turn may contain text plus several calls.

for a senior

Show that you treat model-chosen arguments as untrusted input and that your dispatcher handles unknown names, multiple calls and missing text without crashing.

for a principal

Frame the boundary: the model proposes, your process authorizes and executes. Argue where validation, rate limiting and audit logging of tool invocations belong in the system.

## The core idea When you attach tools to a Gemini request, you are not giving the model the ability to run code. You are giving it a vocabulary. The model can answer in words, or it can answer with a structured message that means "please run `get_weather` with `{"city": "Oslo"}` and tell me what you got". Executing that call is entirely your application's job, and so is deciding whether the call is safe to make at all. ## Where the call lives in the response A Gemini response contains `candidates`, each with a `content` object, and that content has a list of `parts`. A part is a union: it may hold `text`, `inline_data`, a `function_call`, a `function_response`, and so on. A tool request appears as a part whose `function_call` field is populated, with two fields that matter: - `name` — the exact `name` you gave the matching `FunctionDeclaration`. - `args` — the arguments the model chose, shaped by the parameter schema you declared. In the `google-genai` Python SDK, `args` arrives as a Python mapping of already-decoded values, not as a JSON string that you must `json.loads`. That is a real difference from OpenAI-shaped APIs, where tool-call arguments arrive as a string, and it is a frequent source of copy-paste bugs when porting code between vendors. Some function calls also carry an optional `id`. When one is present you should echo it back on the corresponding function response so the model can pair them unambiguously. ## Reading it without assuming shape Three assumptions break in production: 1. **"The function call is `parts[0]`."** Gemini can put a short piece of narration text before the call, so iterate the whole list. 2. **"There is at most one call."** Gemini can emit several `function_call` parts in a single turn when the requests are independent. 3. **"`response.text` always works."** When the turn is purely a tool request, there may be no text at all; accessing `.text` gives you nothing useful and can warn about non-text parts. The SDK gives you `response.function_calls`, a flattened list of every `FunctionCall` in the turn, which is the shortest correct way to check "did the model ask for a tool?". ## The model does not execute anything This is worth saying plainly because juniors often assume the provider runs the function. It does not. Google's servers never see your function body, never open a socket on your behalf, and never learn what your tool returned unless you send it. Everything the model knows about the outcome comes from the FunctionResponse you choose to put in the next turn. That is also your security boundary: validate `args` before you use them, exactly as you would validate any untrusted input, because the values were produced by a language model and may be wrong, out of range, or adversarially influenced by the user's text. There is one nuance for the Python SDK: if you pass plain Python callables as tools instead of `FunctionDeclaration` objects, the SDK's automatic function calling will execute them for you *client-side* and hand you only the final text. That is still your process running your code — the execution model has not changed, only who writes the loop. ## What happens next Once you have `name` and `args`, the loop is the familiar one: dispatch to your implementation, then append two things to the conversation — the model's own turn containing the function call, and a new turn containing a function response part carrying your result. You then call the model again with the whole history and the same tool declarations. The model reads the result and either answers in text or asks for another call. ## Common failure modes - Treating `args` as a JSON string and calling `json.loads` on a dict. - Dropping the model's function-call turn from the history, so the response you send has nothing to attach to. - Trusting `args` blindly — passing a model-chosen file path or SQL fragment straight into a privileged operation. - Assuming an empty `response.text` means the request failed, when it simply means the turn was a tool request. ## Mental model Think of Gemini as a colleague who can only speak. When it needs the weather it writes you a note that says `get_weather(city="Oslo")`. It cannot pick up the phone. You make the call, write the answer on the note, and hand it back.

  • How does reading a Gemini function call differ from reading an OpenAI-style tool call?
    Shape and encoding both differ. Gemini puts the request in a content part's `function_call` field, with `args` already decoded into a mapping. OpenAI-shaped APIs return a `tool_calls` array on the message where the arguments are a JSON *string* you must parse yourself, and each call carries an id you must echo on the reply. Porting code between them means changing both the traversal and the argument decoding.
  • If a turn contains both text and a function call, do you show the text to the user?
    Usually you can, but treat it as optional narration rather than the answer. The authoritative answer comes after the tool result is returned, so the safest UX is to render such text as a transient status line and let the final post-tool turn produce the message you persist.
  • What should you do before passing args into your implementation?
    Validate them as untrusted input. The values were generated by a model that may have been steered by user text, so enforce types, ranges, allow-lists and authorization on your side even though you declared a schema. A schema constrains shape, not intent.

saying these in an interview costs you the question

  • Believing Google's servers execute your declared function
  • Calling json.loads on the args mapping
  • Assuming the call is always parts[0]
  • Thinking an empty response.text means an error
  • Using model-supplied args without any validation

context

open as a page

How do you make a basic Gemini generateContent call with the google-genai SDK?

level: juniorimportance: must knowfreq 80%

basics

~10 s

Create a client with genai.Client(), which picks up the GEMINI_API_KEY environment variable, then call client.models.generate_content(model=..., contents=...). You get back a GenerateContentResponse whose .text property joins the text parts of the first candidate.

open as a page

How do you send an image to Gemini's generateContent as inline data or a file URI?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Gemini takes media as a Part inside contents. Small files ride along as inline_data (raw bytes plus a mime_type); larger or reused files are uploaded to the Files API first and referenced as file_data with a file_uri and mime_type.

open as a page

In the Gemini API, what does a safetySettings entry with a category and threshold control?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Each safetySettings entry pairs a harm category (harassment, hate speech, sexually explicit, dangerous content, civic integrity) with a block threshold. The threshold sets how confident the filter must be before Gemini refuses to return content in that category.

open as a page

What does task_type change in a Gemini embedContent request?

level: middleimportance: must knowfreq 70%

basics

~20 s

task_type tells Gemini's embedding model what the text is for — a stored passage, a search query, a classification input — and the model places the same text differently in vector space for each. Index and query sides must use the matching pair.

open as a page

How do you define a FunctionDeclaration for the Gemini API?

level: middleimportance: must knowfreq 62%

basics

~20 s

A FunctionDeclaration carries a name, a description, and a parameters schema written in the OpenAPI subset Gemini accepts — object type, properties, required, enum, items. Declarations are grouped into a Tool and passed on the request config's tools list.

open as a page

How do you return a function result to Gemini and continue the tool-use turn?

level: middleimportance: must knowfreq 58%

basics

~20 s

Append the model's own turn containing the FunctionCall to the conversation, then append a turn whose part is a FunctionResponse carrying the same function name and a JSON-object result. Re-send the whole history with the same tools and loop until no more calls appear.

open as a page

What does finishReason on a Gemini generateContent candidate tell you?

level: middleimportance: must knowfreq 70%

basics

~20 s

finishReason states why generation of that candidate stopped: STOP for a natural end or a stop sequence, MAX_TOKENS for truncation at the output ceiling, SAFETY or RECITATION for a filtered candidate, OTHER for anything else. Only STOP means the answer is complete.

open as a page

How is Gemini context caching billed, and what does the cache TTL cost?

level: middleimportance: must knowfreq 58%

basics

~20 s

Gemini charges cached input tokens at a reduced per-token rate on each request, plus a separate storage fee for holding the cache — priced per million cached tokens per hour. A longer TTL means a bigger storage bill, whether or not anyone reads the cache.

open as a page

When must a Gemini request use the Files API instead of inline media bytes?

level: middleimportance: must knowfreq 62%

basics

~20 s

Inline bytes only work while the whole generateContent request stays under about 20 MB, so anything larger — and effectively all video — must go through the Files API, which accepts up to 2 GB per file and keeps it for 48 hours.

open as a page

How can you tell a Gemini response was blocked for safety rather than simply empty?

level: middleimportance: must knowfreq 58%

basics

~20 s

Read the metadata, not the text. A rejected prompt returns zero candidates and a block reason under promptFeedback; a filtered generation returns a candidate whose finish reason is SAFETY. Both carry safetyRatings naming the category that fired.

open as a page

How do you handle 429 RESOURCE_EXHAUSTED from the Gemini API in production?

level: seniorimportance: must knowfreq 58%

basics

~20 s

429 RESOURCE_EXHAUSTED means a quota was hit — requests per minute, tokens per minute, or requests per day. Retry per-minute breaches with exponential backoff and jitter and a capped attempt count; a daily breach will not clear by retrying, so shed or queue the work.

open as a page

How do you embed text with the Google GenAI SDK's embed_content method?

level: juniorimportance: should knowfreq 58%

basics

~10 s

Create a genai.Client, call client.models.embed_content with a model name and contents, and optionally pass an EmbedContentConfig for task type and dimensionality. Read the floats from response.embeddings[i].values, one entry per input text.

open as a page

In the Gemini API, what does CachedContent store and how do you call it?

level: juniorimportance: should knowfreq 45%

basics

~10 s

CachedContent stores a large unchanging request prefix — system instruction, documents, uploaded files, tool declarations — server-side under a name like cachedContents/abc123. Later generateContent calls reference that name instead of resending those tokens.

open as a page

How does output_dimensionality work for gemini-embedding-001 vectors?

level: middleimportance: should knowfreq 50%

basics

~20 s

output_dimensionality truncates the model's full-length embedding to a shorter prefix instead of running a smaller model. The model is trained so early dimensions carry most of the signal, and truncated vectors should be re-normalised to unit length before storage.

open as a page

What is automatic function calling in the Gemini Python SDK, and when do you disable it?

level: middleimportance: should knowfreq 40%

basics

~20 s

When you pass plain Python callables as Gemini tools, the SDK derives the declarations from their signatures and then executes them locally on your behalf, looping until the model produces text. Disable it with AutomaticFunctionCallingConfig(disable=True) whenever you need approval, custom error handling or tracing.

open as a page

How do you consume a streaming Gemini generateContent response and assemble the chunks?

level: middleimportance: should knowfreq 66%

basics

~10 s

Call client.models.generate_content_stream() instead of generate_content(). It yields GenerateContentResponse chunks, each holding an incremental slice of the answer, so you concatenate them. The finish reason and final usage metadata arrive on the last chunk.

open as a page

How does Gemini's implicit caching differ from explicit CachedContent caching?

level: middleimportance: should knowfreq 42%

basics

~20 s

Implicit caching is automatic on Gemini 2.5 models: repeat a long common prefix across requests and any discount is applied for you, with no handle and no storage fee. Explicit CachedContent is one you create, name, pay storage for, and control via TTL.

open as a page

How does Gemini turn a minute of video or audio into billable tokens?

level: middleimportance: should knowfreq 52%

basics

~20 s

Gemini bills time-based media per second: roughly 32 tokens per second of audio (about 1,900 per minute) and about 263 tokens per second of video at default resolution (roughly 16,000 per minute), because video is sampled at one frame per second and its audio track counted too.

open as a page

What does adding Gemini's Google Search grounding tool change in the request and response?

level: middleimportance: should knowfreq 52%

basics

~20 s

You add a google_search tool to the request's tools list. The model then decides on its own whether to search, and the returned candidate carries groundingMetadata: the queries it issued, the web sources it used, and the mapping from answer spans to those sources.

open as a page

How do you set a system instruction in the Gemini API, and how does it differ from a user message?

level: middleimportance: should knowfreq 48%

basics

~20 s

Gemini takes the system instruction as a separate top-level field on the request, set through the generation config — not as a role inside contents. Roles in contents are only user and model, and the instruction applies to the whole conversation rather than one turn.

open as a page

How do you embed a large document corpus with Gemini's embedding API?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Chunk documents to fit the model's input token limit, send many chunks per request via the batch embedding endpoint, run bounded concurrency with backoff on rate-limit errors, and make the job resumable by chunk ID so an interruption does not restart it.

open as a page

Your Gemini-embedded RAG index needs a new embedding model. What breaks?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Vectors from two different embedding models are not comparable, so the entire corpus must be re-embedded — there is no conversion. Plan a backfill into a second index and a cutover, not an in-place, gradual replacement.

open as a page

Gemini returned three FunctionCall parts in one turn — how must your app respond?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Execute all three, then send one turn containing three FunctionResponse parts, each naming the function it answers. Splitting them across turns or answering only some leaves calls unresolved and the model typically re-issues them.

open as a page

What does Gemini's tool_config mode ANY do, and when does it trap your loop?

level: seniorimportance: should knowfreq 48%

basics

~20 s

ANY forces Gemini to emit a function call on that turn instead of free text, optionally restricted to allowed_function_names. Left on for the whole conversation it prevents the model from ever writing the final answer, so the loop keeps calling tools until you switch back to AUTO.

open as a page

Why can a Gemini generateContent call return a response with no text at all?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Four causes: the prompt was blocked so candidates is empty, the candidate was filtered and carries no parts, the candidate holds only non-text parts such as a tool call, or a thinking model spent its whole output budget on reasoning. Never trust response.text alone.

open as a page

When does filling Gemini's million-token window beat retrieval over the same corpus?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Loading everything wins when the corpus genuinely fits, is queried many times per session, and questions need whole-document reasoning. Retrieval wins on corpora larger than the window, per-document access control, freshness, and cost at high query volume with low reuse.

open as a page

Why does a Gemini call fail immediately after uploading a video via the Files API?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Large uploads come back in PROCESSING state, and referencing the file before it turns ACTIVE is rejected with a 400 FAILED_PRECONDITION. The fix is to poll files.get on the file name until state is ACTIVE, and to treat FAILED as a re-upload.

open as a page

How do you cut Gemini's token cost for a long video without dropping content?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Send less of the video rather than a smaller file: clip with video_metadata start_offset and end_offset, lower the frame sampling rate, and set media_resolution to MEDIA_RESOLUTION_LOW. Re-encoding to a smaller file changes bytes, not tokens.

open as a page

How do you turn Gemini's groundingMetadata into inline citations in your UI?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Walk groundingSupports: each has a text segment with start and end indices and a list of chunk indices. Slice the answer by those offsets to place markers, resolve the indices against groundingChunks for the source titles and URIs, and render the searchEntryPoint HTML as required.

open as a page

showing 1–30 of 32