skip to content

generateContent API

The main call: contents in, candidates out, with a streaming twin for incremental delivery. Interviewers look for familiarity with finish reasons, why a response can come back with no text at all, and how you handle quota errors and retries.

on this pageshow

questions

5

How do you make a basic Gemini generateContent call with the google-genai SDK?

level: juniorimportance: must knowfreq 80%

answer

  1. one client, then models.generate_content
  2. model plus contents are the payload
  3. knobs live in a config object
  4. stateless: resend the whole history
  5. response.text is only a shortcut

basics

~10 s

Create a client with genai.Client(), which picks up the GEMINI_API_KEY environment variable, then call client.models.generate_content(model=..., contents=...). You get back a GenerateContentResponse whose .text property joins the text parts of the first candidate.

solid answer

~40 s

In the `google-genai` SDK you build a client once — `client = genai.Client()`, which reads `GEMINI_API_KEY` (or `GOOGLE_API_KEY`) from the environment, or takes `api_key=` explicitly — and then call `client.models.generate_content(model="gemini-2.5-flash", contents="...")`. `contents` accepts a plain string, a list of parts, or full `Content` objects with `role` and `parts` for multi-turn history. Everything else — temperature, `max_output_tokens`, `stop_sequences` — goes inside `config=types.GenerateContentConfig(...)`, not as top-level keyword arguments; that trips people coming from other SDKs. The response is a `GenerateContentResponse` with `candidates`, `usage_metadata` and a `.text` convenience property. Under the hood this is `POST v1beta/models/{model}:generateContent` with an `x-goog-api-key` header, so the Python field names are snake_case while the REST body is camelCase (`generationConfig`, `maxOutputTokens`).

code

python · 16 lines
python
from google import genai
from google.genai import types

client = genai.Client()  # reads GEMINI_API_KEY from the environment

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Explain HTTP caching in two sentences.",
    config=types.GenerateContentConfig(
        max_output_tokens=200,
        temperature=0.2,
    ),
)

print(response.text)
print(response.usage_metadata.total_token_count)

go deeper

for a junior

Be able to write the four lines from memory: import, client, generate_content with model and contents, print response.text. Say plainly that the API key comes from the GEMINI_API_KEY environment variable, not a hardcoded string.

for a middle

Explain that the call is stateless, that history is a list of Content objects with roles user and model, and that generation parameters live inside GenerateContentConfig rather than as top-level arguments.

for a senior

Show the production hygiene: one reused client, an explicit max_output_tokens ceiling, usage_metadata logged per call for cost attribution, and awareness that the same SDK targets Vertex with different credentials.

for a principal

Own the boundary decision — whether calls go through the public Gemini endpoint or Vertex, how keys are provisioned and rotated per environment, and how the stateless request shape interacts with your conversation-storage and cost-attribution design.

## The shape of the call Gemini's core text-generation entry point is `generateContent`. In Python you reach it through the `google-genai` SDK (`from google import genai`), which is a thin, generated wrapper over the REST surface. Three things make up a call: a **client** (auth and transport), a **model id**, and **contents** (what you are sending). ``` client = genai.Client() response = client.models.generate_content(model=..., contents=...) ``` The client is constructed once per process and is safe to reuse; constructing one per request wastes connection setup. ## Authentication `genai.Client()` with no arguments reads an API key from the environment — `GEMINI_API_KEY`, falling back to `GOOGLE_API_KEY`. You can also pass `api_key="..."` explicitly, which is convenient in tests but should never be a literal in committed code. The same SDK can target Vertex AI instead of the public Gemini API by constructing `genai.Client(vertexai=True, project=..., location=...)`, in which case auth switches to Google Cloud application-default credentials rather than an API key. That dual mode is the reason a Gemini snippet you copied from a Vertex tutorial may fail with a permission error: the two backends authenticate differently even though the method names are identical. Over plain HTTP the request is `POST https://generativelanguage.googleapis.com/v1beta/models/{model}:generateContent` with the key in an `x-goog-api-key` header. Knowing the raw shape matters because errors and quotas are reported at that level. ## contents `contents` is deliberately flexible. The simplest form is a string, which the SDK wraps into a single user `Content` with one text `Part`. For a multi-turn conversation you pass a list of `Content` objects, each with a `role` (`"user"` or `"model"` — note it is `model`, not `assistant`) and a `parts` list. There is no separate history object in the stateless call: you resend the whole conversation every turn, and every turn's tokens are re-charged as input. The SDK also exposes `client.chats.create(...)` as a convenience that keeps the list for you, but it is still the same stateless endpoint underneath. ## config Generation knobs do not live at the top level of the Python call. They go in `config=types.GenerateContentConfig(...)`: `temperature`, `top_p`, `top_k`, `max_output_tokens`, `stop_sequences`, `candidate_count`, `response_mime_type`, `response_schema`, and so on. The REST body calls the same object `generationConfig` with camelCase keys. A frequent beginner bug is passing `temperature=0.2` directly to `generate_content`, which raises a type error rather than silently working. `max_output_tokens` is worth calling out: unlike some other vendors' APIs it is optional and defaults to the model's maximum, but leaving it unset in production means an unexpectedly long generation can cost far more than you budgeted. ## The response You get a `GenerateContentResponse`. Its important members are: - `candidates` — a list, normally of length one unless you asked for more via `candidate_count`. Each candidate has `content` (with `parts`), a `finish_reason`, and `safety_ratings`. - `usage_metadata` — `prompt_token_count`, `candidates_token_count`, `total_token_count`, plus `cached_content_token_count` and `thoughts_token_count` where they apply. This is what you log for cost accounting; do not estimate tokens client-side when the server tells you exactly. - `.text` — a convenience property that concatenates the text parts of the first candidate. It is fine for demos and unsafe as the only path in production, because it can be `None` when the candidate carries no text parts. ## Async and streaming twins The same method exists in three variants that share the request shape: `client.models.generate_content` (blocking, one response), `client.models.generate_content_stream` (an iterator of partial responses), and `client.aio.models.generate_content` / `client.aio.models.generate_content_stream` for asyncio. Choosing between them is purely a delivery decision — the request body and the billing are the same. ## What interviewers listen for A junior answer that names the client, the model id, `contents`, and `response.text` is enough. What separates a stronger answer is knowing that the call is stateless, that generation parameters live in a config object, and that `.text` is a convenience over a candidate list rather than the response itself.

  • The call is stateless — what does that mean for cost as a conversation grows?
    Every turn resends the full history, so input tokens grow roughly linearly with conversation length and the cost per turn climbs even if the user's message is short. Practical mitigations are trimming or summarising old turns, or moving a large stable prefix into explicit context caching so the repeated part is billed at the cached rate.
  • What is the assistant role called in Gemini's contents array?
    `model`. Gemini uses `user` and `model` as the two roles inside `contents`; there is no `assistant` role and no `system` role in the array — a system instruction is a separate configuration field, not a message. Code ported from an OpenAI-shaped payload will be rejected until the roles are renamed.
  • How do you request more than one candidate, and what does it cost?
    Set `candidate_count` in `GenerateContentConfig`. The response then carries several entries in `candidates`, each with its own `finish_reason`. Output tokens are charged for all candidates, so asking for four samples costs roughly four times the output of one; the prompt is charged once.

saying these in an interview costs you the question

  • Passing temperature directly to generate_content instead of config
  • Assuming the API keeps conversation state between calls
  • Using role 'assistant' in a Gemini contents array
  • Thinking response.text is the whole response object
  • Creating a new genai.Client for every request

context

open as a page

What does finishReason on a Gemini generateContent candidate tell you?

level: middleimportance: must knowfreq 70%

basics

~20 s

finishReason states why generation of that candidate stopped: STOP for a natural end or a stop sequence, MAX_TOKENS for truncation at the output ceiling, SAFETY or RECITATION for a filtered candidate, OTHER for anything else. Only STOP means the answer is complete.

open as a page

How do you handle 429 RESOURCE_EXHAUSTED from the Gemini API in production?

level: seniorimportance: must knowfreq 58%

basics

~20 s

429 RESOURCE_EXHAUSTED means a quota was hit — requests per minute, tokens per minute, or requests per day. Retry per-minute breaches with exponential backoff and jitter and a capped attempt count; a daily breach will not clear by retrying, so shed or queue the work.

open as a page

How do you consume a streaming Gemini generateContent response and assemble the chunks?

level: middleimportance: should knowfreq 66%

basics

~10 s

Call client.models.generate_content_stream() instead of generate_content(). It yields GenerateContentResponse chunks, each holding an incremental slice of the answer, so you concatenate them. The finish reason and final usage metadata arrive on the last chunk.

open as a page

Why can a Gemini generateContent call return a response with no text at all?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Four causes: the prompt was blocked so candidates is empty, the candidate was filtered and carries no parts, the candidate holds only non-text parts such as a tool call, or a thinking model spent its whole output budget on reasoning. Never trust response.text alone.

open as a page