Your Gemini service keeps hitting safety blocks on legitimate user content — how do you diagnose it?
answer
- Measure the rate before touching settings
- Which single category actually fires
- Input side or output side
- System instruction before threshold change
- One category, one surface, recorded
basics
~20 sInstrument first: log the block reason, the finish reason and every category/probability pair per request. That shows which single category fires and where. Then fix in order — prompt and system instruction, then that one category's threshold, never a blanket loosening.
solid answer
~50 sStart with evidence, not settings. Log for every call the `promptFeedback.blockReason`, the candidate `finishReason`, and the full `safetyRatings` list with categories and probabilities, keyed to your own request id — then measure how many blocks per thousand requests, and which category dominates. Almost always it is one category (frequently dangerous content or sexually explicit) firing on domain vocabulary: clinical, security-research, legal or moderation text. Next check *where* it fires: an input block means the user's text tripped it, an output block means your prompt steered the model into it. Fix in escalating order — reframe the system instruction so the model answers clinically rather than operationally, strip or summarise pasted content that carries the trigger, and only then raise that one category from `BLOCK_MEDIUM_AND_ABOVE` to `BLOCK_ONLY_HIGH`. Keep the other categories untouched, and give users an explicit refusal message rather than an empty box.
go deeper
Know that blocks are visible in the response metadata and must be logged, and that the first question to ask is which harm category fired rather than assuming the whole filter is too strict.
Explain how to separate input blocks from output blocks and what each implies, and why prompt or system-instruction changes are tried before any threshold change.
Show the full loop: instrument, measure a base rate per surface and category, fix in escalating order, re-measure on the same traffic slice, and handle refusals in the UI instead of returning an empty box.
Own the governance: who may loosen a category, what evidence and sign-off is required, how the decision is documented for audit, and how block-rate regression is caught when a model version changes.
## Establish the base rate before you change anything The failure mode of this investigation is guessing. Someone reports "the assistant went silent on my ticket", a threshold gets loosened globally, and nobody can say afterwards whether it helped. So the first move is instrumentation: for every call, record the block reason if present, the candidate finish reason, and the complete safety ratings array — category and probability for each entry — alongside your request id, the feature that made the call, and enough of the prompt to reproduce it (hashed or truncated if it is user data). Now you have a rate (blocks per thousand calls), a breakdown by category, and a breakdown by surface. That data almost always collapses the problem. It is rarely "safety is too aggressive"; it is "`HARM_CATEGORY_DANGEROUS_CONTENT` fires at MEDIUM on 4% of calls from the clinical-notes feature and nowhere else". ## Input block or output block? The two block paths point at different owners. If `promptFeedback.blockReason` is set, the *user's* text tripped the classifier before the model ran — the trigger is in what you sent, so look at pasted documents, retrieved chunks in a RAG prompt, or a conversation history that has accumulated a flagged turn. If instead a candidate comes back with a safety finish reason, the model generated the problematic content, which usually means your instructions steered it there: a prompt that says "give step-by-step instructions" over a hazardous topic will earn a block even when the user's question was innocuous. RAG pipelines deserve special suspicion. You may be injecting retrieved passages that carry the flagged vocabulary and blaming the user's question. Log the assembled prompt's provenance, not just the user turn. ## Fix in escalating order **1. System instruction and prompt framing.** Tell the model what register to answer in — "respond with clinical, non-operational information suitable for a licensed practitioner" moves output away from the shape that trips dangerous-content filters. This is the cheapest fix and the one that survives model upgrades. **2. Input hygiene.** If a pasted transcript is what fires the filter, summarise or excerpt it before sending, or chunk it so a single flagged passage does not poison the whole request. For moderation-review products, consider whether the flagged content even needs to reach the generative model or can be handled by a classifier. **3. Threshold change — one category.** Raise the specific category that your logs implicate, typically from `BLOCK_MEDIUM_AND_ABOVE` to `BLOCK_ONLY_HIGH`. Leave the others at their defaults. Setting everything to the loosest value is the move that fails a review: it is unmeasured, unjustified, and removes protections unrelated to your actual problem. **4. Route or fail gracefully.** Some requests genuinely should be refused. Give the user a clear message naming that the request could not be answered, offer a rephrase, and keep a support path. An empty response box teaches users the product is broken. ## What not to do Do not retry the blocked request unchanged — the classifier is stable enough that you will get the same block, paying quota and latency for nothing. Do not silently swallow blocks into empty strings; that is how blocked outputs end up persisted as legitimate empty summaries in a database. Do not conclude "safety got stricter" without evidence — but *do* re-baseline after a model version change, because thresholds are interpreted by a classifier that ships with the model, and a family upgrade can shift the rate in either direction. Pin the model version in production for this reason, and roll new versions with the block rate on the dashboard. Also separate the true refusals. A polite "I can't help with that" in the response body is not a filter block at all — finish reason STOP, tokens billed, no block metadata. If your users' complaint is really about those, no threshold change will help; that is prompt and model-selection work. ## Close the loop After a change, compare the block rate on the same traffic slice, and sample the requests that now get through to confirm the outputs are ones you are happy to ship. Write the decision down: which category, from what to what, on which surface, with the measured before/after and who approved it. That record is what makes the loosening defensible when someone asks six months later why this service runs a permissive dangerous-content threshold.
- Your logs show blocks only on the RAG-backed endpoint. What does that suggest?That the trigger is in the retrieved passages, not the user's question. Log the assembled prompt with chunk provenance and check which document ids appear in blocked requests. Fixes are usually retrieval-side: filter or summarise the offending source, cap how much raw text you inject, or chunk so one flagged passage does not carry the whole request into a block.
- After a model version upgrade the block rate doubled. What is your first move?Roll back to the pinned previous version if the surface is user-facing, then compare like-for-like on a replayed traffic sample to confirm the model change caused it. The classifier ships with the model family, so the same thresholds can behave differently across versions. Re-baseline the rate and, if the new behaviour is acceptable, adjust the implicated category deliberately rather than reverting wholesale.
- Users complain the assistant refuses, but your logs show no blocks at all. What is happening?Those are model-generated refusals, not filter blocks: finish reason STOP, output tokens billed, refusal text in the content. No safety setting will change them. The levers are prompt framing, a system instruction that establishes the legitimate professional context, sometimes a different model tier, and — if you need visibility — classifying the response text so refusals show up in your metrics at all.
saying these in an interview costs you the question
- Loosens every category at once instead of the implicated one
- Retries blocked requests unchanged hoping for a different result
- Has no per-category logging, so cannot say what fired
- Blames the safety filter for model-generated refusal text
- Swallows blocked responses as empty strings into storage