skip to content

Guardrails & Safety

The controls around model input and output: classifiers on both sides, prompt-injection defenses, moderation APIs, PII redaction, and jailbreak mitigation. Interviewers probe this because a model wired to tools and private data fails in ways a chatbot cannot.

on this pageshow

questions

19

Why do content moderation classifiers return per-category scores, not one flag?

level: juniorimportance: must knowfreq 55%

answer

  1. one flag, one action — too blunt
  2. categories exist to differ in response
  3. cost of a miss varies by harm
  4. self-harm needs routing, not blocking
  5. thresholds live per category

basics

~20 s

Different harms need different responses. Per-category scores let a system hard-block one category, route self-harm signals to a crisis flow, and only warn on mild insults. A single unsafe flag forces one action and one cutoff for every kind of harm.

solid answer

~50 s

A harm taxonomy is the list of categories a moderation system scores against — things like hate, harassment, sexual content, violence, and self-harm. Classifiers return a score per category rather than one boolean because the product's response differs by category. In a competitive shooter's party chat, a message scoring high on self-harm should page a human within minutes and surface crisis resources; a message scoring high on trash talk about the match should probably do nothing at all. One flag collapses those into a single action. Per-category scores also let you set a different threshold per category, which matters because the cost of a false positive is wildly uneven: wrongly muting a player for banter is annoying, wrongly missing a credible threat is a safety incident. The design test for any category is simple — if two harms trigger the same action at the same threshold, they do not need to be separate categories.

go deeper

for a junior

Be able to name a few harm categories and say plainly that different categories trigger different actions, so a single unsafe flag is not enough.

for a middle

Explain the response test for whether a category should exist, and give a concrete pair of harms where the correct actions differ.

for a senior

Show how per-category thresholds are set from the cost of each error type, and how you handle overlap precedence and dead categories in a live taxonomy.

for a principal

Own the taxonomy as a product artifact: who approves changes, how new categories are introduced without invalidating historical metrics, and how enforcement policy is documented and defended externally.

## What a harm taxonomy is A harm taxonomy is the enumerated set of categories your moderation system recognises: hate, harassment, sexual content, violence, self-harm, and so on, often with subcategories. A classifier scores a piece of text against each category independently and returns a score per category, not a single verdict. The taxonomy is the contract between the classifier and the product: it defines what you can even talk about detecting. ## Why one flag is not enough A single `unsafe: true` flag forces three things you almost never want: 1. **One action for every harm.** Blocking is right for some categories and actively wrong for others. Suppressing a message that scores high on self-harm hides a user who may need help. 2. **One threshold for every harm.** Categories have wildly different base rates and wildly different error costs. A cutoff calibrated for mild profanity is far too loose for credible threats of violence. 3. **One review path.** Operationally you need to know *why* something was flagged in order to route it — to an automated mute, to a general review queue, or to an urgent queue with a tight response time. A boolean throws away all the information needed to make those decisions. ## Designing categories: the response test The useful design rule is that **a category earns its existence when it maps to a distinct response**. Not when it feels conceptually different, and not because a vendor's default list happens to contain it. Apply the test: if category A and category B always produce the same action at the same threshold, merge them. If one broad category is producing two different actions depending on what a reviewer sees inside it, split it. ## A worked example: party chat in a competitive shooter Consider voice-to-text party chat in a competitive online shooter. A flat "harassment" category performs badly here, because the same words carry opposite meanings. "You're trash, get off my team" between friends who queue together every night is banter. The identical string aimed at a random matchmade stranger is abuse. And "this map is trash, whoever designed it should be fired" is neither — it is trash talk about the match. Splitting harassment into **targeted-at-a-player** and **trash-talk-about-the-match** passes the response test, because the responses genuinely differ: the first can lead to a chat mute or a report escalation, while the second is left alone at any score. That split also lets you set the targeted subcategory's threshold low (catch more, tolerate some false positives) and effectively disable enforcement on the other, which a merged category cannot express. Note what the split does *not* solve: relationship context. The classifier sees text, not the fact that these two players are friends. Systems handle that with signals outside the classifier — party membership, prior reports, whether the recipient blocked the sender — combined with the score. Naming that limitation is what separates a thoughtful answer from a naive one. ## Per-category thresholds Once you have categories, each gets its own cutoff. Rough shape in practice: - **Categories where a miss is catastrophic** (child safety, credible threats): low threshold, accept false positives, human review on everything flagged. - **Categories where a false positive damages the product** (mild profanity, edgy humour in an adult context): high threshold, or no automated enforcement at all. - **Categories that need routing rather than blocking** (self-harm): moderate threshold, action is escalation and resource surfacing, not suppression. ## Where taxonomies go wrong - **Categories that never fire.** Dead categories add classifier cost and reviewer confusion. Measure per-category volume and prune. - **Overlapping categories with no precedence rule.** One message scoring high on three categories needs a defined winner, or your action is nondeterministic. - **Adopting a default taxonomy unchanged.** Off-the-shelf category lists are a starting point; the categories that matter to a shooter's chat, a health forum, and a legal-research tool are not the same set. - **Assuming categories are mutually exclusive.** They are not — scores are independent and a message can be high on several at once. ## What interviewers listen for They want the response test articulated, at least one concrete example where two harms demand different actions, and awareness that thresholds are per category rather than global. A candidate who says "we call the moderation check and block if it says unsafe" has described a system that will both over-block and under-protect.

  • When should you merge two harm categories rather than keep them separate?
    When they always produce the same action at the same threshold. If nothing downstream branches on the distinction, the split adds classifier surface, reviewer ambiguity and dashboard noise for no behavioural difference. The reverse also holds: split a category the moment reviewers start applying two different outcomes to items inside it.
  • A message scores high on three categories at once. How should the system decide what to do?
    With an explicit precedence rule, defined in advance, usually ordered by severity of the required response. Self-harm routing outranks a chat mute, and a hard-block category outranks a warning. Without a stated rule the action depends on evaluation order in code, which makes enforcement nondeterministic and impossible to audit.
  • Why can't a text classifier alone distinguish banter between friends from abuse toward a stranger?
    Because the distinguishing information is not in the text. Identical words carry opposite intent depending on the relationship between sender and recipient. Production systems combine the classifier score with out-of-band signals — party membership, mutual history, prior reports, whether the recipient blocked the sender — and reserve enforcement for the combination, not the score alone.

saying these in an interview costs you the question

  • Treats moderation as a single unsafe boolean
  • Uses one global threshold across all harm categories
  • Assumes blocking is the right response to every category
  • Adopts a default category list without checking it fits the product
  • Assumes a message belongs to exactly one category

context

open as a page

Which output channels let an LLM app leak data without running code?

level: middleimportance: must knowfreq 62%

basics

~20 s

Any path where model output causes a fetch is an egress channel: rendered image and link URLs, outbound tool or webhook calls, redirect targets, even hostname lookups. Controls only work once every such sink is enumerated.

open as a page

Why bind the tenant filter on RAG retrieval server-side, not in the prompt?

level: middleimportance: must knowfreq 58%

basics

~20 s

Anything the model can influence can be widened by text that reaches it. Derive the tenant scope from the authenticated session and bind it at the query layer, so the model contributes only a search string and never the filter that decides which corpus is readable.

open as a page

Why is the instruction/data split inside an LLM prompt not a real trust boundary?

level: middleimportance: must knowfreq 72%

basics

~20 s

A model receives one flat token sequence. System text, user text and retrieved text carry no enforced privilege difference — only a learned tendency to prefer operator wording. Prompt-level separation shifts odds; the enforceable boundary has to live outside the model.

open as a page

In an LLM app, why screen model output when the user input already passed moderation?

level: middleimportance: must knowfreq 62%

basics

~20 s

Input screening only sees what the user typed. A model can still emit harmful text from a benign prompt, from retrieved documents, or from another user's content pulled into context, so the output is a separate surface that needs its own classifier.

open as a page

Why tokenize client names before an LLM call instead of just deleting them?

level: middleimportance: must knowfreq 55%

basics

~20 s

Tokenization swaps each identifier for a stable surrogate such as CLIENT_0447, so the model can still track who did what and the application can restore the real name afterwards. Deleting the name destroys the references the answer depends on.

open as a page

In LLM APIs, what does a training opt-out still leave retained on the provider side?

level: middleimportance: must knowfreq 60%

basics

~20 s

A training opt-out only stops your text improving future models. Providers typically still hold prompts and responses for a bounded abuse-monitoring window, and features such as prompt caching, batch jobs and fine-tuning keep their own copies with their own lifetimes.

open as a page

How do you stop model-authored URLs and images from leaking data at render?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Allowlist the destinations instead of inspecting payloads: strip or rewrite every model-authored image source through a same-origin proxy that only fetches approved hosts, restrict link targets to an approved list, never render model-authored raw HTML, and add a browser-enforced content policy as a backstop.

open as a page

What must a human approval prompt show before an agent's irreversible action?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Show the resolved call the system will actually make — tool, target and every argument verbatim — plus where the data that triggered it came from. A model-authored summary of the action is attacker-influenceable text, so approving on the summary approves nothing.

open as a page

How do you pick the score threshold for a content-moderation classifier?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Not from a default. Score a labelled sample, plot precision against recall per category, and choose the operating point where the cost of a false positive and the cost of a missed harm balance for that category — then check the resulting flag volume fits your review capacity.

open as a page

How do you keep API keys and credentials out of LLM prompt logs and traces?

level: juniorimportance: should knowfreq 40%

basics

~20 s

Keep secrets out of the prompt in the first place by injecting them at the tool layer, then redact at the write path: match known credential formats before a trace record is stored, replace matches with a hash or placeholder, and restrict who can read the trace store.

open as a page

In prompt-injection defense, what is spotlighting and why randomize delimiters?

level: middleimportance: should knowfreq 52%

basics

~20 s

Spotlighting makes the untrusted region of a prompt unmistakable — by delimiting it, marking every token inside it, or encoding it. A per-request random delimiter matters because a fixed tag can be reproduced inside the content itself to fake the end of the data block.

open as a page

How do you stop an agent's HTTP tool from becoming an exfiltration path?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Remove the model's control over the destination. Have tools take an identifier and build the URL server-side, run the tool runtime with default-deny network egress and an allowlist of approved hostnames, refuse cross-host redirects, and cap outbound body size and rate.

open as a page

How does an isolated extraction pass keep untrusted text out of a tool-calling model?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A first model with no tools reads the untrusted document and emits only a fixed, typed record. The privileged model that can call tools sees that record, never the free text — so an injected imperative has no channel into the component holding the capabilities.

open as a page

If your moderation classifier times out, should the request fail open or fail closed?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It depends on the context, and it must be written down in advance. Fail closed where unchecked content is unacceptable, fail open where a classifier outage taking down the product is worse than a rare miss — and log every bypassed request for retrospective review either way.

open as a page

Where must a deletion request reach in an LLM app beyond the app database?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Every store that received a copy of the text: trace and log stores, eval and golden datasets mined from production traces, agent memory, vector-index payloads, provider-side retention and caches, and any fine-tuning set. Data already trained into weights cannot be deleted at all.

open as a page

Which prompt-injection controls should sit outside the model rather than in the prompt?

level: principalimportance: should knowfreq 44%

basics

~20 s

Anything that must hold after the model is persuaded belongs in code: approval before side effects, the set of capabilities the agent can invoke at all, and isolation of untrusted text from the privileged component. Prompt-level measures are a first filter, never load-bearing.

open as a page

How do you size and operate the human review queue behind a moderation classifier?

level: principalimportance: should knowfreq 38%

basics

~20 s

Derive daily escalation volume from traffic and the chosen thresholds, staff to that number with headroom for spikes, tier by urgency with a response-time target on the most time-critical categories, and sample allowed traffic too — flagged-only review can never reveal what you missed.

open as a page

How do you decide where LLM inference physically runs for regulated client data?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat processing location as an architecture constraint set per data class, not a vendor checkbox. Classify what may leave the jurisdiction, check which providers offer in-region inference and which subprocessors are involved, and accept the capability gap that regional or self-hosted options impose.

open as a page