skip to content

Safety & Grounding

Gemini returns safetyRatings per harm category and can refuse to answer at a threshold you configure, and it can ground responses in Google Search with attributions attached. Interviewers ask how you distinguish a safety block from an ordinary empty answer, and what the grounding metadata is good for.

on this pageshow

questions

6

In the Gemini API, what does a safetySettings entry with a category and threshold control?

level: juniorimportance: must knowfreq 62%

answer

  1. Per-request, per-category knob
  2. Category plus threshold, five categories
  3. Thresholds describe confidence, not severity
  4. LOW_AND_ABOVE strictest, BLOCK_NONE loosest
  5. Core filters ignore your settings

basics

~20 s

Each safetySettings entry pairs a harm category (harassment, hate speech, sexually explicit, dangerous content, civic integrity) with a block threshold. The threshold sets how confident the filter must be before Gemini refuses to return content in that category.

solid answer

~40 s

Safety settings are a per-request list you pass in the generation config. Each entry has a `category` (one of `HARM_CATEGORY_HARASSMENT`, `HARM_CATEGORY_HATE_SPEECH`, `HARM_CATEGORY_SEXUALLY_EXPLICIT`, `HARM_CATEGORY_DANGEROUS_CONTENT`, `HARM_CATEGORY_CIVIC_INTEGRITY`) and a `threshold` from the ladder `BLOCK_LOW_AND_ABOVE` (strictest) → `BLOCK_MEDIUM_AND_ABOVE` → `BLOCK_ONLY_HIGH` → `BLOCK_NONE` (most permissive; newer models also expose `OFF`). Gemini scores each request and each candidate for every category and blocks when the estimated probability reaches the threshold you chose. The settings apply to both the prompt and the model's output, and any category you do not list keeps the model's default. They are a tuning knob, not an off switch: a core set of non-configurable filters (for example child-safety content) applies regardless of what you send.

code

python · 24 lines
python
from google import genai
from google.genai import types

client = genai.Client()

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Explain the pharmacology of common overdose antidotes.",
    config=types.GenerateContentConfig(
        safety_settings=[
            types.SafetySetting(
                category=types.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
                threshold=types.HarmBlockThreshold.BLOCK_ONLY_HIGH,
            ),
            types.SafetySetting(
                category=types.HarmCategory.HARM_CATEGORY_HARASSMENT,
                threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE,
            ),
        ],
    ),
)

for rating in response.candidates[0].safety_ratings or []:
    print(rating.category, rating.probability)

go deeper

for a junior

Know the shape: a list of entries, each with a harm category and a block threshold, passed in the request config. Be able to name a couple of the categories and say that the threshold decides how aggressively that category is blocked.

for a middle

Explain that thresholds encode classifier confidence (NEGLIGIBLE to HIGH), that the ladder runs from BLOCK_LOW_AND_ABOVE to BLOCK_NONE, that omitted categories keep defaults, and that the settings cover both the prompt and the generated candidate.

for a senior

Show that you tune one category at a time against measured false-refusal rates, log the returned ratings for evidence, and know that a permissive setting still leaves non-configurable filters and the model's own refusal behaviour in place.

for a principal

Own the policy: which product surfaces get which thresholds, who signs off on loosening a category, how the decision is recorded for audit, and what compensating controls (human review, output moderation, logging) justify running a surface at the permissive end.

## What a safety setting actually is Gemini runs a classifier over the prompt you send and over each candidate it generates. For every supported harm category it produces an estimated *probability* that the content is unsafe — reported on a coarse scale of `NEGLIGIBLE`, `LOW`, `MEDIUM`, `HIGH`. A `safetySettings` entry is your instruction about where on that scale the API should stop handing content back. It is a **request-level** parameter, not an account setting, so two calls made a second apart can use different strictness. In the google-genai Python SDK the list lives on the generation config: `config=types.GenerateContentConfig(safety_settings=[types.SafetySetting(category=..., threshold=...)])` On the wire it is the `safetySettings` array on the `generateContent` request body, each element `{"category": ..., "threshold": ...}`. ## The categories The configurable harm categories for text generation are harassment, hate speech, sexually explicit content, dangerous content, and civic integrity. They are named with the `HARM_CATEGORY_` prefix (`HARM_CATEGORY_HARASSMENT`, `HARM_CATEGORY_HATE_SPEECH`, `HARM_CATEGORY_SEXUALLY_EXPLICIT`, `HARM_CATEGORY_DANGEROUS_CONTENT`, `HARM_CATEGORY_CIVIC_INTEGRITY`). You set each one independently — a medical-information product typically loosens `DANGEROUS_CONTENT` while leaving harassment and hate speech at their defaults, because loosening everything is both unnecessary and harder to defend in a review. Any category you omit from the array simply keeps the service default for that model. There is no "inherit but stricter" mode; the entry you send fully replaces the default for that one category on that one request. ## The threshold ladder Thresholds are cumulative bands, read as "block content whose probability is at or above this level": - `BLOCK_LOW_AND_ABOVE` — strictest; blocks anything rated LOW, MEDIUM or HIGH. - `BLOCK_MEDIUM_AND_ABOVE` — blocks MEDIUM and HIGH. - `BLOCK_ONLY_HIGH` — blocks only HIGH. - `BLOCK_NONE` — the configurable filter for that category stops blocking. - `OFF` — exposed on newer models to disable that category's filter outright. The common mistake is reading the names as "how bad the content is". They describe the classifier's *confidence*, not severity. Content the model is only mildly suspicious about is exactly what `BLOCK_LOW_AND_ABOVE` catches and what `BLOCK_ONLY_HIGH` lets through, which is why over-strict settings produce so many false refusals on benign domain vocabulary (self-harm helplines, security research, oncology). ## Where the settings apply They govern both directions. If your prompt itself trips a threshold, the whole request is refused before generation: you get no candidates and a `promptFeedback` block reason instead. If the generated candidate trips a threshold, that candidate stops with a safety finish reason. Either way the response carries `safetyRatings` — category plus probability — so you can see which classifier fired rather than guessing. (Vertex AI's variant of the API additionally returns severity fields; the Gemini Developer API's core rating is category + probability, plus a blocked flag.) ## What you cannot turn off Setting every category to the most permissive value does **not** produce an unfiltered model. A separate, non-configurable safety layer always applies to core harms such as child sexual abuse material, and the model's own alignment training still makes it decline requests it considers harmful — that refusal arrives as a normal, successful response whose text is a polite "I can't help with that", not as a filter block. Confusing those two failure modes is the classic interview trap: one is an API-level block you detect from metadata, the other is ordinary generated text. ## Practical guidance Pick thresholds per feature, not per company. A moderation-review console reading user reports needs permissive settings simply to display what it is reviewing; a children's tutoring product wants the strict end. Log the returned ratings alongside your own request id so you can measure the false-refusal rate before you loosen anything, and keep the loosening scoped to the single category that is actually firing. Document the choice: "we set `HARM_CATEGORY_DANGEROUS_CONTENT` to `BLOCK_ONLY_HIGH` because clinical dosage questions were being blocked at MEDIUM" is a defensible sentence in a design review; "we set everything to `BLOCK_NONE`" is not.

  • What happens to the categories you leave out of the safetySettings array?
    They keep the model's default threshold for that category. The array is a set of overrides, not a complete specification — omitting harassment while overriding dangerous content leaves harassment exactly as the service ships it. Because defaults can differ between model families, teams that care about reproducibility list every category explicitly rather than relying on what the default happens to be today.
  • If you set every category to BLOCK_NONE, will the model answer anything?
    No. Two things still stop it. A non-configurable filter layer covers core harms such as child sexual abuse material regardless of your settings, and the model's own training still leads it to decline requests it judges harmful — that arrives as an ordinary successful response containing a refusal sentence, with a normal stop finish reason and no block metadata at all.
  • Do safety settings apply to the prompt as well as the output?
    Yes, to both. A prompt that trips a threshold is rejected before generation, so you receive zero candidates and a block reason in promptFeedback. A generated candidate that trips one stops with a safety finish reason. Handling only the second case is a common bug: the code indexes candidates[0] and throws an IndexError on exactly the requests that were blocked at the input.

saying these in an interview costs you the question

  • Thinks BLOCK_NONE makes the model answer anything
  • Reads thresholds as severity levels rather than classifier confidence
  • Believes safety settings are an account-wide console toggle
  • Assumes filters only inspect model output, never the prompt
  • Confuses a model's polite refusal text with an API safety block

context

open as a page

How can you tell a Gemini response was blocked for safety rather than simply empty?

level: middleimportance: must knowfreq 58%

basics

~20 s

Read the metadata, not the text. A rejected prompt returns zero candidates and a block reason under promptFeedback; a filtered generation returns a candidate whose finish reason is SAFETY. Both carry safetyRatings naming the category that fired.

open as a page

What does adding Gemini's Google Search grounding tool change in the request and response?

level: middleimportance: should knowfreq 52%

basics

~20 s

You add a google_search tool to the request's tools list. The model then decides on its own whether to search, and the returned candidate carries groundingMetadata: the queries it issued, the web sources it used, and the mapping from answer spans to those sources.

open as a page

How do you set a system instruction in the Gemini API, and how does it differ from a user message?

level: middleimportance: should knowfreq 48%

basics

~20 s

Gemini takes the system instruction as a separate top-level field on the request, set through the generation config — not as a role inside contents. Roles in contents are only user and model, and the instruction applies to the whole conversation rather than one turn.

open as a page

How do you turn Gemini's groundingMetadata into inline citations in your UI?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Walk groundingSupports: each has a text segment with start and end indices and a list of chunk indices. Slice the answer by those offsets to place markers, resolve the indices against groundingChunks for the source titles and URIs, and render the searchEntryPoint HTML as required.

open as a page

Your Gemini service keeps hitting safety blocks on legitimate user content — how do you diagnose it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Instrument first: log the block reason, the finish reason and every category/probability pair per request. That shows which single category fires and where. Then fix in order — prompt and system instruction, then that one category's threshold, never a blanket loosening.

open as a page