In the Gemini API, what does a safetySettings entry with a category and threshold control?
answer
- Per-request, per-category knob
- Category plus threshold, five categories
- Thresholds describe confidence, not severity
- LOW_AND_ABOVE strictest, BLOCK_NONE loosest
- Core filters ignore your settings
basics
~20 sEach safetySettings entry pairs a harm category (harassment, hate speech, sexually explicit, dangerous content, civic integrity) with a block threshold. The threshold sets how confident the filter must be before Gemini refuses to return content in that category.
solid answer
~40 sSafety settings are a per-request list you pass in the generation config. Each entry has a `category` (one of `HARM_CATEGORY_HARASSMENT`, `HARM_CATEGORY_HATE_SPEECH`, `HARM_CATEGORY_SEXUALLY_EXPLICIT`, `HARM_CATEGORY_DANGEROUS_CONTENT`, `HARM_CATEGORY_CIVIC_INTEGRITY`) and a `threshold` from the ladder `BLOCK_LOW_AND_ABOVE` (strictest) → `BLOCK_MEDIUM_AND_ABOVE` → `BLOCK_ONLY_HIGH` → `BLOCK_NONE` (most permissive; newer models also expose `OFF`). Gemini scores each request and each candidate for every category and blocks when the estimated probability reaches the threshold you chose. The settings apply to both the prompt and the model's output, and any category you do not list keeps the model's default. They are a tuning knob, not an off switch: a core set of non-configurable filters (for example child-safety content) applies regardless of what you send.
code
python · 24 linesfrom google import genai
from google.genai import types
client = genai.Client()
response = client.models.generate_content(
model="gemini-2.5-flash",
contents="Explain the pharmacology of common overdose antidotes.",
config=types.GenerateContentConfig(
safety_settings=[
types.SafetySetting(
category=types.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
threshold=types.HarmBlockThreshold.BLOCK_ONLY_HIGH,
),
types.SafetySetting(
category=types.HarmCategory.HARM_CATEGORY_HARASSMENT,
threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE,
),
],
),
)
for rating in response.candidates[0].safety_ratings or []:
print(rating.category, rating.probability)go deeper
Know the shape: a list of entries, each with a harm category and a block threshold, passed in the request config. Be able to name a couple of the categories and say that the threshold decides how aggressively that category is blocked.
Explain that thresholds encode classifier confidence (NEGLIGIBLE to HIGH), that the ladder runs from BLOCK_LOW_AND_ABOVE to BLOCK_NONE, that omitted categories keep defaults, and that the settings cover both the prompt and the generated candidate.
Show that you tune one category at a time against measured false-refusal rates, log the returned ratings for evidence, and know that a permissive setting still leaves non-configurable filters and the model's own refusal behaviour in place.
Own the policy: which product surfaces get which thresholds, who signs off on loosening a category, how the decision is recorded for audit, and what compensating controls (human review, output moderation, logging) justify running a surface at the permissive end.
## What a safety setting actually is Gemini runs a classifier over the prompt you send and over each candidate it generates. For every supported harm category it produces an estimated *probability* that the content is unsafe — reported on a coarse scale of `NEGLIGIBLE`, `LOW`, `MEDIUM`, `HIGH`. A `safetySettings` entry is your instruction about where on that scale the API should stop handing content back. It is a **request-level** parameter, not an account setting, so two calls made a second apart can use different strictness. In the google-genai Python SDK the list lives on the generation config: `config=types.GenerateContentConfig(safety_settings=[types.SafetySetting(category=..., threshold=...)])` On the wire it is the `safetySettings` array on the `generateContent` request body, each element `{"category": ..., "threshold": ...}`. ## The categories The configurable harm categories for text generation are harassment, hate speech, sexually explicit content, dangerous content, and civic integrity. They are named with the `HARM_CATEGORY_` prefix (`HARM_CATEGORY_HARASSMENT`, `HARM_CATEGORY_HATE_SPEECH`, `HARM_CATEGORY_SEXUALLY_EXPLICIT`, `HARM_CATEGORY_DANGEROUS_CONTENT`, `HARM_CATEGORY_CIVIC_INTEGRITY`). You set each one independently — a medical-information product typically loosens `DANGEROUS_CONTENT` while leaving harassment and hate speech at their defaults, because loosening everything is both unnecessary and harder to defend in a review. Any category you omit from the array simply keeps the service default for that model. There is no "inherit but stricter" mode; the entry you send fully replaces the default for that one category on that one request. ## The threshold ladder Thresholds are cumulative bands, read as "block content whose probability is at or above this level": - `BLOCK_LOW_AND_ABOVE` — strictest; blocks anything rated LOW, MEDIUM or HIGH. - `BLOCK_MEDIUM_AND_ABOVE` — blocks MEDIUM and HIGH. - `BLOCK_ONLY_HIGH` — blocks only HIGH. - `BLOCK_NONE` — the configurable filter for that category stops blocking. - `OFF` — exposed on newer models to disable that category's filter outright. The common mistake is reading the names as "how bad the content is". They describe the classifier's *confidence*, not severity. Content the model is only mildly suspicious about is exactly what `BLOCK_LOW_AND_ABOVE` catches and what `BLOCK_ONLY_HIGH` lets through, which is why over-strict settings produce so many false refusals on benign domain vocabulary (self-harm helplines, security research, oncology). ## Where the settings apply They govern both directions. If your prompt itself trips a threshold, the whole request is refused before generation: you get no candidates and a `promptFeedback` block reason instead. If the generated candidate trips a threshold, that candidate stops with a safety finish reason. Either way the response carries `safetyRatings` — category plus probability — so you can see which classifier fired rather than guessing. (Vertex AI's variant of the API additionally returns severity fields; the Gemini Developer API's core rating is category + probability, plus a blocked flag.) ## What you cannot turn off Setting every category to the most permissive value does **not** produce an unfiltered model. A separate, non-configurable safety layer always applies to core harms such as child sexual abuse material, and the model's own alignment training still makes it decline requests it considers harmful — that refusal arrives as a normal, successful response whose text is a polite "I can't help with that", not as a filter block. Confusing those two failure modes is the classic interview trap: one is an API-level block you detect from metadata, the other is ordinary generated text. ## Practical guidance Pick thresholds per feature, not per company. A moderation-review console reading user reports needs permissive settings simply to display what it is reviewing; a children's tutoring product wants the strict end. Log the returned ratings alongside your own request id so you can measure the false-refusal rate before you loosen anything, and keep the loosening scoped to the single category that is actually firing. Document the choice: "we set `HARM_CATEGORY_DANGEROUS_CONTENT` to `BLOCK_ONLY_HIGH` because clinical dosage questions were being blocked at MEDIUM" is a defensible sentence in a design review; "we set everything to `BLOCK_NONE`" is not.
- What happens to the categories you leave out of the safetySettings array?They keep the model's default threshold for that category. The array is a set of overrides, not a complete specification — omitting harassment while overriding dangerous content leaves harassment exactly as the service ships it. Because defaults can differ between model families, teams that care about reproducibility list every category explicitly rather than relying on what the default happens to be today.
- If you set every category to BLOCK_NONE, will the model answer anything?No. Two things still stop it. A non-configurable filter layer covers core harms such as child sexual abuse material regardless of your settings, and the model's own training still leads it to decline requests it judges harmful — that arrives as an ordinary successful response containing a refusal sentence, with a normal stop finish reason and no block metadata at all.
- Do safety settings apply to the prompt as well as the output?Yes, to both. A prompt that trips a threshold is rejected before generation, so you receive zero candidates and a block reason in promptFeedback. A generated candidate that trips one stops with a safety finish reason. Handling only the second case is a common bug: the code indexes candidates[0] and throws an IndexError on exactly the requests that were blocked at the input.
saying these in an interview costs you the question
- Thinks BLOCK_NONE makes the model answer anything
- Reads thresholds as severity levels rather than classifier confidence
- Believes safety settings are an account-wide console toggle
- Assumes filters only inspect model output, never the prompt
- Confuses a model's polite refusal text with an API safety block