What does finishReason on a Gemini generateContent candidate tell you?
answer
- per-candidate, not per-response
- only one value means complete
- truncation has its own value
- empty candidates is a different field
- open enum — handle the unknown case
basics
~20 sfinishReason states why generation of that candidate stopped: STOP for a natural end or a stop sequence, MAX_TOKENS for truncation at the output ceiling, SAFETY or RECITATION for a filtered candidate, OTHER for anything else. Only STOP means the answer is complete.
solid answer
~50 sEach entry in `candidates` carries its own `finish_reason`, and it is the field that tells you whether the text you got is a whole answer. `STOP` is the good case — the model ended naturally or hit one of your `stop_sequences`. `MAX_TOKENS` means it was cut off at the `max_output_tokens` ceiling, so the text is truncated mid-thought and any JSON in it is probably unparseable; you either raise the ceiling or shorten the task. `SAFETY` and `RECITATION` mean the candidate itself was suppressed, typically leaving no text parts at all. `OTHER` is the catch-all. The critical distinction is scope: `finish_reason` is per-candidate and describes generation. A prompt rejected *before* generation produces no candidates at all — the reason then lives in `prompt_feedback.block_reason` instead. Production code should branch on `finish_reason != STOP` rather than assuming success from a 200 response.
go deeper
Know that finishReason lives on each candidate and that STOP is the only value meaning the answer is complete. Recognise MAX_TOKENS as truncation caused by the output-token ceiling.
Explain each common value and, crucially, the difference between a candidate stopped mid-generation and a prompt blocked before generation, which yields an empty candidates list and promptFeedback instead.
Demonstrate the operational habit: branch on the reason before parsing, log it with token usage, alert on a rising MAX_TOKENS rate, and handle unknown enum values as generic failures rather than crashing.
Frame it as a contract between the model and downstream systems — decide where truncated or filtered output is allowed to flow, what the retry and degradation policy is, and which finish-reason rates become service-level indicators.
## Why the field exists An HTTP 200 from `generateContent` means the request was well-formed and the service ran it. It says nothing about whether the model produced a complete answer. `finishReason` is the field that carries that second, semantic verdict, and ignoring it is the single most common way a Gemini integration ships a bug that only shows up on long or unusual inputs. ## Where it lives `finishReason` sits on each **candidate**, not on the response. In the Python SDK that is `response.candidates[0].finish_reason`, typed as `types.FinishReason`. If you requested several candidates with `candidate_count`, each one has its own value — one may finish cleanly while another is truncated, and your selection logic should prefer the clean one. ## The values you will actually see **STOP** — normal termination. The model emitted its end-of-turn signal, or it matched one of the `stop_sequences` you configured. Note the ambiguity baked in here: a stop-sequence match and a natural ending both report `STOP`, so if you rely on stop sequences and also need to know whether the model was finished, you must inspect the tail of the text yourself. **MAX_TOKENS** — generation hit the `max_output_tokens` ceiling from your config (or the model's own limit) and was cut mid-stream. The text is a prefix of an answer. This is the value that most often causes downstream crashes, because a truncated JSON object or code block parses as garbage. Treat it as a failure, not a partial success: retry with a higher ceiling, split the task, or ask for a shorter output format. **SAFETY** — the candidate was filtered by content safety after generation. The candidate typically comes back with no usable text; the per-category detail lives in the candidate's safety ratings. **RECITATION** — generation was halted because the output was reproducing recognisable source material at length. It is rare, and it surprises people because the prompt looks innocuous; it tends to appear when you ask a model to reproduce lyrics, long standard texts, or well-known code verbatim. **OTHER** — an unclassified stop. Treat it exactly like a failure and log the whole candidate so you can characterise it later. The enum also carries values for tool-related and policy-specific terminations, so your handling must be *open*: match the values you know and fall through to a generic "incomplete" branch for anything else. Hard-coding an exhaustive match is a maintenance trap, because Google adds enum members over time and an SDK that receives an unknown value will hand it to you rather than crashing. ## finishReason versus promptFeedback This is the distinction interviewers probe. There are two different rejection points: - **Input-side.** The prompt is rejected before any generation happens. The response comes back with an **empty `candidates` list** and `promptFeedback.blockReason` set (values include `SAFETY`, `BLOCKLIST`, `PROHIBITED_CONTENT`, `OTHER`). There is no `finishReason` to read, because no candidate was ever produced. Code that does `response.candidates[0]` blindly throws an IndexError here — a real, common production crash. - **Output-side.** Generation started and was then cut short. You get a candidate whose `finishReason` explains why, possibly with no text parts. So the correct defensive order is: check the candidate list is non-empty first (and read `prompt_feedback` if it is not), then check `finish_reason`, then read the parts. ## In streaming When you consume `generate_content_stream`, intermediate chunks carry no meaningful finish reason; the terminal value arrives on the final chunk. That means a stream can deliver text you have already rendered to the user and *then* end with a non-`STOP` reason. Any UI that streams must be able to mark a message as truncated or withdrawn after the fact. ## What to do with it Log it on every call, alongside token usage — the ratio of `MAX_TOKENS` finishes is a direct health signal that your output ceiling is mis-sized for the task. Alert on a rising rate rather than on individual events. And never let a non-`STOP` candidate flow into a JSON parser, a database write, or a tool invocation without an explicit branch: a truncated answer that looks plausible is worse than an error.
- You get finishReason MAX_TOKENS while asking for structured JSON. What do you do?Treat the output as invalid rather than trying to repair it. Raise `max_output_tokens`, or reduce what you asked for — fewer fields, fewer items per call, or the work split across requests. Repairing truncated JSON by appending brackets guesses at content the model never produced, which is how silently wrong records reach a database.
- How does finishReason behave when you request multiple candidates?Each candidate carries its own value, so a single response can mix a clean `STOP` with a truncated or filtered sibling. Selection logic should filter to candidates that finished cleanly before ranking them, and fall back explicitly if none did — not just take `candidates[0]`.
- Does a non-STOP finish reason mean you were not billed?No. Tokens generated before the stop are counted in `usage_metadata.candidates_token_count`, and the prompt is always billed. A run of `MAX_TOKENS` finishes is therefore paid-for output you cannot use, which is exactly why the truncation rate is worth monitoring as a cost signal, not just a quality one.
saying these in an interview costs you the question
- Treating HTTP 200 as proof the answer is complete
- Reading finishReason off the response instead of the candidate
- Parsing JSON from a MAX_TOKENS candidate anyway
- Confusing a blocked prompt with a blocked candidate
- Assuming the enum has a fixed, exhaustive set of values