skip to content

How can you tell a Gemini response was blocked for safety rather than simply empty?

level: middleimportance: must knowfreq 58%

answer

  1. Two blocks, two different fields
  2. Input side versus candidate side
  3. HTTP 200 either way
  4. Refusal text is not a block
  5. Check metadata before reading text

basics

~20 s

Read the metadata, not the text. A rejected prompt returns zero candidates and a block reason under promptFeedback; a filtered generation returns a candidate whose finish reason is SAFETY. Both carry safetyRatings naming the category that fired.

solid answer

~50 s

There are two distinct blocks and they surface in different places. If the *input* is rejected, Gemini returns no candidates at all and sets `promptFeedback.blockReason` (values include `SAFETY` and `OTHER`), with `promptFeedback.safetyRatings` showing which category tripped. If the *output* is filtered, you do get a candidate, but its `finishReason` is `SAFETY` and its content is empty or truncated, with per-category `safetyRatings` on the candidate. Neither case is an HTTP error — the call returns 200, so code that only checks status codes sees success. Convenience accessors make this worse: reading the SDK's `.text` property on a blocked response yields nothing useful and hides why. The defensive read order is: check `prompt_feedback.block_reason` first, then `candidates[0].finish_reason` against `SAFETY`, then use the text. Remember that a model's polite "I can't help with that" is *not* a block — it is normal text with a `STOP` finish reason.

code

python · 23 lines
python
from google import genai
from google.genai import types

client = genai.Client()

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=user_text,
)

feedback = response.prompt_feedback
if feedback and feedback.block_reason:
    raise RuntimeError(f"prompt blocked: {feedback.block_reason}")

if not response.candidates:
    raise RuntimeError("no candidates and no block reason")

candidate = response.candidates[0]
if candidate.finish_reason == types.FinishReason.SAFETY:
    fired = [(r.category, r.probability) for r in candidate.safety_ratings or []]
    raise RuntimeError(f"output blocked: {fired}")

print(candidate.content.parts[0].text)

go deeper

for a junior

Know that a blocked call still returns HTTP 200 and that you must look at promptFeedback and the candidate's finish reason rather than trusting that empty text means the model had nothing to say.

for a middle

Explain the two block paths precisely: promptFeedback.blockReason with zero candidates for input, finishReason SAFETY with per-candidate safetyRatings for output, and be able to write the guarded read order.

for a senior

Demonstrate the operational side: no blind retries, structured logs of category and probability, one shared helper used by interactive and batch paths alike, and a UI story for a block that arrives mid-stream.

for a principal

Frame refusals as a product surface. Decide what users see, which refusals get human review, how false-refusal rate is measured per feature, and how that evidence feeds threshold and prompt policy across teams.

## Two blocks, two places to look Gemini can refuse at either end of the call, and the two refusals look nothing alike in the response object. Knowing which is which is the difference between a five-minute fix and an afternoon of guessing. **Input block.** The classifier scores your prompt before generation. If it exceeds the threshold in force for a category, the API does not generate at all: `candidates` is empty (or absent), and `promptFeedback` is populated with a `blockReason` — `SAFETY` when a harm category fired, `OTHER` for reasons the API does not enumerate to you — plus a `safetyRatings` array describing the prompt's scores. Code written as `response.candidates[0]` throws an `IndexError` here, which is why blocked prompts so often show up in production as a mysterious crash rather than a handled refusal. **Output block.** The model generated something, and the filter stopped it. Now there *is* a candidate, but `candidate.finishReason` is `SAFETY` and `candidate.content` is empty or cut short. The candidate's own `safetyRatings` tell you which category and at what probability. ## The status code lies to you Both blocks come back as HTTP 200 with a well-formed body. This surprises people who wrap the SDK in a generic "if it didn't raise, it worked" helper. A safety block is a *successful* API call that produced no usable content, and it has to be handled as a domain outcome — surfaced to the user, logged with the category, possibly routed to a fallback — rather than as a transport error to retry. Blind retries on a blocked prompt burn quota and produce the identical block every time, because the classifier is deterministic enough that the same input keeps failing. ## The third case: the model just declined The subtlest confusion is between a filter block and an ordinary refusal. When the model itself decides not to help, it *generates text* saying so. The finish reason is `STOP`, the ratings are unremarkable, `usageMetadata` shows output tokens billed, and there is no block metadata anywhere. You cannot detect that case from the metadata at all — only from the content. Interviewers like this distinction because it separates people who have actually read a blocked response object from people who have read a blog post about safety filters. Related but different: `finishReason: RECITATION` means generation stopped because the output was reproducing memorised material, not because of a harm category, and `MAX_TOKENS` means you simply ran out of budget. Treating every non-`STOP` finish reason as "safety" will mislabel your own truncation bugs as policy problems. ## A defensive read path The order matters, because the earlier checks guard the later ones: 1. `response.prompt_feedback` — if `block_reason` is set, the input was rejected; report it with the ratings and do not retry the same text. 2. `response.candidates` — if the list is empty and there was no block reason, treat it as an anomaly rather than assuming safety. 3. `candidate.finish_reason` — compare against `FinishReason.SAFETY`; also branch on `RECITATION` and `MAX_TOKENS` so those get their own message. 4. Only then read the parts' text. Wrap that in one helper used by every call site. The single most common production defect here is having the checks in one code path (the interactive endpoint) and not in another (the batch job), so the batch job silently writes empty strings into the database. ## Streaming changes the timing, not the mechanism When you stream, a prompt block shows up on the very first chunk you receive — the stream is effectively over before content starts. An output block can arrive mid-stream: you will have already emitted partial text to the user, and then a chunk arrives whose candidate carries the safety finish reason. Your UI needs a way to retract or annotate what it already rendered; a stream that just stops looks identical to a network drop unless you inspect that final chunk. Accumulate the finish reason and ratings from the last chunk you see rather than assuming the stream ending means success. ## What to log Log the block reason or finish reason, every returned category/probability pair, a hash or truncated form of the prompt, and your own request id. That gives you a false-refusal dataset. Without the categories you cannot tell whether to adjust a threshold, rewrite a system instruction, or accept that the request genuinely should have been refused — and "the model keeps returning nothing" is not a bug report anyone can act on.

  • Why is retrying a safety-blocked request usually the wrong response?
    Because it is not a transient failure. The classifier scores the same input the same way, so the retry reproduces the block, consumes quota and adds latency. The right moves are to report the refusal, log the category, and — if the refusal is a false positive on legitimate content — change the input, the system instruction, or that one category's threshold. Retry only if you actually altered the request.
  • How does a mid-stream safety block reach the client?
    You receive normal content chunks first, then a chunk whose candidate carries the safety finish reason and ratings; nothing more follows. Because partial text is already on screen, the UI must be able to retract or annotate it. Code that treats "stream ended" as success cannot distinguish a filtered generation from a dropped connection, so accumulate the finish reason from the final chunk.
  • What distinguishes a RECITATION finish reason from a SAFETY one?
    RECITATION means generation stopped because the output was reproducing memorised source material, not because a harm category fired; the safety ratings will look unremarkable. It usually signals a prompt that pushes the model to quote a well-known text verbatim. The fix is prompt-side — ask for a summary or paraphrase, or supply the source yourself — not a safety-threshold change.
  • How would you distinguish a filter block from the model simply declining?
    Only by the metadata. A decline is a normal generation: finish reason STOP, output tokens billed in usageMetadata, no block reason, and readable refusal text in the parts. A filter block has empty or truncated content plus either promptFeedback.blockReason or a SAFETY finish reason. If you need to detect declines too, you are doing content classification on the text, not metadata inspection.

saying these in an interview costs you the question

  • Expects a 4xx error or an exception on safety blocks
  • Indexes candidates[0] without checking the list is non-empty
  • Treats a blocked call as transient and retries it
  • Calls the model's refusal sentence a safety block
  • Assumes any empty response means the safety filter fired

context