skip to content

Why can a Gemini generateContent call return a response with no text at all?

level: seniorimportance: should knowfreq 52%

answer

  1. 200 does not mean text
  2. empty list versus empty parts
  3. not every part is text
  4. reasoning draws on the same budget
  5. write one unwrapping helper

basics

~20 s

Four causes: the prompt was blocked so candidates is empty, the candidate was filtered and carries no parts, the candidate holds only non-text parts such as a tool call, or a thinking model spent its whole output budget on reasoning. Never trust response.text alone.

solid answer

~40 s

A 200 response with no usable text has distinct causes and they need different handling. First, the prompt can be rejected before generation: `candidates` is empty and `prompt_feedback.block_reason` explains why — indexing `candidates[0]` here raises an IndexError. Second, the candidate can be generated then suppressed, arriving with a non-`STOP` `finish_reason` and no text parts. Third, the candidate may hold only non-text parts — a tool call, for instance — so there is nothing for `.text` to concatenate. Fourth, on thinking-enabled models the reasoning tokens are drawn from the same `max_output_tokens` budget, so a tight ceiling can be exhausted before any answer text is emitted; you see `finish_reason` `MAX_TOKENS`, a non-zero `thoughts_token_count` in `usage_metadata`, and empty text. Robust code checks candidates, then finish reason, then iterates parts explicitly.

code

python · 13 lines
python
from google.genai import types

def extract_text(response) -> str:
    if not response.candidates:
        reason = getattr(response.prompt_feedback, "block_reason", None)
        raise RuntimeError(f"prompt rejected before generation: {reason}")

    candidate = response.candidates[0]
    if candidate.finish_reason != types.FinishReason.STOP:
        raise RuntimeError(f"incomplete candidate: {candidate.finish_reason}")

    parts = (candidate.content.parts if candidate.content else None) or []
    return "".join(part.text for part in parts if part.text)

go deeper

for a junior

Know that response.text can be None and that you should check the candidates list before reading it. Recognise that a successful HTTP call does not guarantee any text came back.

for a middle

Be able to name the distinct causes — prompt rejected, candidate suppressed, non-text parts only — and describe the order in which you check candidates, finish reason, then parts.

for a senior

Show that you have debugged this in production: read usage_metadata to spot reasoning tokens consuming the output budget, classify causes separately in logs, and route each to a different remedy instead of one generic retry.

for a principal

Own the policy — which empty-response classes are user-visible errors, which are retried and at what cost, and how output-budget sizing on reasoning models is standardised across services so one team's default does not silently break another's.

## The failure that looks like success `response.text` is a convenience property: it walks the first candidate's parts, concatenates the text ones, and hands you a string. When there is nothing to concatenate it yields `None` rather than raising, so a naive integration turns a rich diagnostic response into a `NoneType` error three frames away from the API call. Every serious Gemini integration eventually writes a response-unwrapping helper, and knowing the four distinct causes is what that helper encodes. ## Cause 1: the prompt never generated anything The input can be rejected before the model runs. The HTTP status is still 200, but `candidates` is an **empty list** and `prompt_feedback` carries a `block_reason`. Because the list is empty, the very first thing your code does — `response.candidates[0]` — throws. The fix is an explicit guard that reads `prompt_feedback` and surfaces the reason to the caller, rather than a try/except that swallows it. Operationally this class should be counted separately: a spike in prompt-side rejection usually means an upstream data source started feeding you content you did not expect, not that the model regressed. ## Cause 2: the candidate was generated and then suppressed Here a candidate object exists but its `content` may be `None` or its `parts` list empty, with `finish_reason` set to something other than `STOP`. This is why the unwrapping order matters: check the list is non-empty, then check the finish reason, then defensively handle `content` being absent. Writing `candidate.content.parts` without the `content is None` guard is a second common crash. ## Cause 3: there is content, just not text A candidate's `parts` list is heterogeneous. A part may carry text, or it may carry something else entirely — a function call, inline binary data, and so on. `.text` only sees the text ones. If your prompt supplied tools and the model chose to call one, the natural, correct response contains no prose at all, and treating that as an error means your agent loop can never make progress. The lesson is structural: iterate `parts` and dispatch on what each part actually holds, rather than asking for text and being surprised by its absence. ## Cause 4: the thinking budget ate the answer This one catches experienced engineers. On thinking-enabled Gemini models, the internal reasoning tokens are billed as output and are drawn from the **same** `max_output_tokens` allowance as the visible answer. Set that ceiling low — say a couple of hundred tokens, a habit carried over from non-reasoning models — and the model can spend the entire allowance thinking, hit the ceiling, and return a candidate with `finish_reason` `MAX_TOKENS` and zero visible text. The tell is in `usage_metadata`: `thoughts_token_count` is large while `candidates_token_count` is near zero. The remedies are to raise `max_output_tokens` so it covers reasoning *plus* the answer, or to constrain reasoning explicitly through the thinking configuration (including reducing the thinking budget to zero on models that support disabling it). Sizing the ceiling as "expected answer length" is simply wrong arithmetic on these models. ## The unwrapping order Encode it once, in one place: 1. `candidates` non-empty? If not, report `prompt_feedback.block_reason`. 2. `finish_reason == STOP`? If not, classify: truncation, filtering, or unknown. 3. `content` present with a `parts` list? If not, you have a suppressed candidate. 4. Iterate parts; collect text, and dispatch non-text parts to their handlers. 5. Only then decide whether "no text" is an error for *this* call site. ## Why interviewers ask it Because it separates people who have run this API in production from people who have only run the quickstart. The quickstart's `print(response.text)` works every time on a friendly prompt; the four failure modes above are exactly what a real traffic mix produces, and each one wants a different response — surface to the user, retry with a bigger ceiling, feed into the tool loop, or page someone.

  • How do you tell a thinking-budget exhaustion apart from an ordinary truncation?
    Both report `MAX_TOKENS`, so read `usage_metadata`. Ordinary truncation shows a large `candidates_token_count` with visible text that stops mid-sentence. Budget exhaustion shows a large `thoughts_token_count` and a `candidates_token_count` at or near zero — the model never got to the answer. The fixes differ: shorten the task versus raise the ceiling.
  • What does response.text do when the first candidate has no text parts?
    It returns `None` rather than raising, because it simply has nothing to concatenate. That silence is the hazard: the error surfaces later as a NoneType failure in your own code, far from the API call. Guarding at the boundary — and logging the finish reason and usage alongside — keeps the diagnosis at the point where the information still exists.
  • Should an empty text response be retried automatically?
    Only for the causes that are actually transient. A prompt-side rejection will reproduce identically, so retrying wastes quota. Budget exhaustion needs a changed request, not a repeat. A tool-call-only response is not an error at all. Blanket retry on 'no text' converts four distinct signals into one expensive loop.

saying these in an interview costs you the question

  • Calling response.text and assuming it is a string
  • Indexing candidates[0] without checking the list is empty
  • Treating a tool-call-only response as a model failure
  • Sizing max_output_tokens as just the expected answer length
  • Retrying every empty response regardless of cause

context