skip to content

LLM Safety & Security

You will learn how to ship LLM features that survive hostile input: layered injection defenses, guardrails, moderation, PII handling, and safely scoped tool-using agents. Interviewers probe this to check you can defend a production AI product, not just build one.

on this pageshow

explore

questions

page 1 of 2

Why is pasting a full patient record into a chat assistant a leak even without model training?

level: juniorimportance: must knowfreq 70%

answer

  1. prompts are not ephemeral
  2. count the copies you make
  3. history is resent every turn
  4. logs, traces, support tooling, caches
  5. send fields, not documents

basics

~20 s

Text placed in a prompt becomes data your own system holds: it sits in conversation history and is resent every turn, and it lands in request logs, traces and support tooling. Training use is a separate, narrower question.

solid answer

~50 s

"The provider does not train on it" answers only one of several exposure paths, and not the common one. Once text is in a prompt it is in the conversation history, so it is re-sent with every subsequent turn of that session; it is in whatever request/response logging, tracing and error capture the application does; it is visible to anyone who can open that thread — support staff, an admin console, a shared link — and it may be picked up later when someone samples production traffic. A whole discharge summary also carries far more identifiers than the task needs. The control is minimization at the input boundary: pull only the fields the task requires, redact before the call, and design the interface so the narrow path is the easy one rather than leaving a free-text box that invites a full paste.

go deeper

for a junior

Be ready to name the places a prompt gets copied: conversation history, logs and traces, support views, caches. Say plainly that not training on the data is a different question from not storing it.

for a middle

Explain minimization concretely — assemble prompts from the specific fields a task needs instead of accepting a pasted document, and keep full prompt bodies out of default logging.

for a senior

Show that you would go looking for the copies: which store holds threads, what the tracing layer captures, what support tooling can view, what sampling feeds evaluation. Then argue for the interface change that removes the paste path.

for a principal

Own the tradeoff between debuggability and exposure. Full prompt capture makes incidents solvable and makes every log store sensitive; argue for a default of metadata-only capture with deliberate, time-boxed exceptions.

## The claim being tested The reassurance people reach for is "the model does not learn from my input." Whether or not that is true for a given deployment, it addresses one exposure path — the model's weights — and leaves untouched every copy your own system makes of the same text. Most real LLM privacy incidents are of the second kind. Nothing exotic happens; the sensitive text simply ends up in more places than the person pasting it imagined. ## The copies your system makes **Conversation history.** A chat feature is stateless underneath: to answer turn five, the application re-sends turns one through four. A discharge summary pasted at the start is transmitted again on every later message and sits in whatever store holds the thread. It also keeps consuming context, which is why teams later add summarisation or truncation — and a summary of a record is still a record. **Logs and traces.** Application logs, an observability/tracing layer, an error report that captures the request body, a queue message, a retry buffer. These are usually built to be broadly readable by engineers, precisely because they exist for debugging. A prompt written into a log line inherits that audience. **Operational surfaces.** Admin consoles, support tooling that lets a human view a user's session to reproduce a complaint, exported conversation transcripts, screenshots pasted into a ticket, a shared thread link. Each is a legitimate feature and each widens the readership of the pasted text. **Downstream reuse.** Sampled production traffic is a normal source of evaluation examples and debugging fixtures. Text that entered as a prompt can end up in a dataset that outlives the session. **Caches.** Repeated prefixes are frequently cached to save cost and latency. A cached prefix is another stored copy with its own lifetime. None of these are misconfigurations; they are the default shape of a production application. Privacy work here is about noticing that a prompt is not ephemeral just because a conversation feels ephemeral. ## Why the whole record is the wrong unit A discharge summary bundles identity (name, address, record number, dates of birth and admission) with clinical content (diagnoses, medications, notes). Almost no task needs all of it. "Draft a follow-up appointment message" needs a diagnosis category, a discharge date and a clinician name; it does not need an address or a full medication history. Pasting the whole document imports every identifier into every copy listed above, and it also makes the text re-identifiable: identifiers such as record numbers are resolvable by anyone inside the organisation with access to the chart system. ## Minimization in practice *Fetch narrowly.* If the application can pull structured fields from the source system, pull the three fields the prompt template needs rather than the document. A template with named slots is both safer and more reliable than a free paste. *Redact before the call.* Run detection and replace identifiers with placeholders. This is lossy and imperfect, so treat it as reducing blast radius, not as a guarantee. *Shorten the retention of the copies you cannot avoid.* Do not log full prompts by default; log a request identifier and enough metadata to debug. If prompt capture is needed for a specific investigation, make it deliberate and time-boxed. *Design the easy path.* This is the part teams miss. If the only affordance is a large text box, users will paste whole documents, because that is the fastest way to get a good answer. If the interface offers "summarise this visit" with the record selected by identifier and assembled server-side, the sensitive text never passes through the user's clipboard or the chat log at all. Behaviour follows the interface; a policy asking clinicians to paste less will lose to a UI that rewards pasting more. *Give users a way to correct mistakes.* Someone will paste something they should not have. A visible delete on the conversation, that actually removes the stored copies, is worth more than a warning banner. ## How to talk about it The strong answer separates the two questions cleanly. Question one: what does the model provider do with the data — a contractual and configuration matter. Question two: what does *your* application do with it — a design matter entirely within your control, and the one that produces most incidents. A candidate who only answers question one has not thought about their own system.

  • The team says they will just scrub the logs. Is that enough?
    It helps, but it only covers one copy. Conversation history, caches, support views, exported transcripts and sampled traces are separate stores with separate lifetimes. Scrubbing logs while the thread store keeps the full text moves the problem rather than solving it. Minimization at the input boundary shrinks every copy at once, which is why it comes first.
  • If the assistant summarises the record before storing it, is the privacy problem solved?
    No. A summary of a patient record is still patient data, and summarisation happens after the full text has already been sent, logged and cached. It can reduce what persists in history going forward, which is worth something, but the exposure at the moment of the call is unchanged. Treat summarisation as a context-budget technique, not a privacy control.
  • Why does the interface design matter more than a written policy here?
    Because the fastest path wins. A free-text box makes pasting an entire document the lowest-effort way to get a useful answer, so people do it regardless of training. Assembling the prompt server-side from selected structured fields removes the temptation entirely and gives you one place to enforce minimization and redaction.

saying these in an interview costs you the question

  • Says training opt-out means the data is not stored anywhere
  • Thinks a prompt is discarded once the model replies
  • Forgets conversation history is resent on every turn
  • Treats logs and traces as private because only engineers read them
  • Assumes summarising the record removes the privacy concern

context

open as a page

Why doesn't a strict JSON schema on an LLM response make its content safe?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A schema constrains shape, not meaning. It guarantees the reply parses and carries the required fields and types. It says nothing about whether a field holds an invented price, a competitor's name, an unsafe instruction, or text that is dangerous where you paste it.

open as a page

Why should an autonomous agent hold its own credentials instead of the operator's?

level: middleimportance: must knowfreq 62%

basics

~20 s

An agent running on a human's session inherits every permission that human has, and its actions are logged as theirs. Giving the agent its own principal lets you scope it narrowly, revoke it alone, and attribute what it actually did.

open as a page

In an LLM feature, why moderate both user input and model output?

level: middleimportance: must knowfreq 68%

basics

~20 s

Input screening rejects harmful requests before you pay for a generation and before they enter the context. Output screening catches harm the model produces anyway, including harm that arrived through retrieved documents or tool results the input check never saw.

open as a page

Why do off-the-shelf PII detectors miss hospital MRNs, and what does over-redaction cost?

level: middleimportance: must knowfreq 62%

basics

~20 s

Shipped recognizers cover identifiers standardised nationally or industry-wide; a facility-assigned medical record number has no fixed format, so nothing matches it. Misses leak silently, while blanket masking strips the dosages, dates and lab values the answer depended on.

open as a page

Where do you place input, output and tool-call guardrails around an LLM agent?

level: middleimportance: must knowfreq 72%

basics

~20 s

Three placements, each catching a different failure. Input rails screen the incoming turn before or alongside generation. Output rails screen the finished response before the user sees it. Tool rails wrap each function call, checking arguments before the side effect and treating the returned result as fresh untrusted input.

open as a page

What is the lethal trifecta, and how do you break it in an email assistant?

level: middleimportance: must knowfreq 72%

basics

~20 s

The lethal trifecta is one agent session holding three things at once: access to private data, exposure to untrusted content, and an outbound channel. Any two are survivable; all three let injected text steal data. Remove one capability.

open as a page

Why does a fixed red-team prompt suite overstate an LLM feature's real safety?

level: middleimportance: must knowfreq 62%

basics

~20 s

A fixed suite only measures attacks someone already wrote down. A real attacker sees your defense and adapts, so passing the suite proves the known attacks are patched — not that the system resists new ones.

open as a page

In LLM security, how do direct and indirect prompt injection differ as risk categories?

level: middleimportance: must knowfreq 78%

basics

~20 s

Direct injection arrives in the user's own turn, so the user is the adversary. Indirect injection hides in content the application feeds the model later — retrieved documents, tool results, memory — so a third party attacks through your data pipeline.

open as a page

How do you scope a deploy agent's tools so a hijacked run stays contained?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Replace generic tools with narrow, capability-shaped ones, bind the dangerous parameters server-side rather than letting the model choose them, split read from write, and enforce every limit at the API that executes the call — never in the prompt.

open as a page

In a multi-facility RAG corpus, where must the tenant filter be enforced, and why?

level: seniorimportance: must knowfreq 55%

basics

~20 s

At query time in the retrieval layer, as a pre-filter derived from the caller's verified identity — or as physically separate indexes. Anything above that, including instructions to the model, is advice rather than a boundary, because a retrieved chunk is already disclosed.

open as a page

Why is a prompt-injection detection classifier not a security boundary?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Detection is probabilistic and only has to be wrong once, while an attacker can retry and tune against it. Adaptive attacks against a dozen published defenses succeeded over 90% of the time. Filters reduce noise; a deterministic architectural limit is the boundary.

open as a page

How do you report attack-success rate so it doesn't overstate LLM safety?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Attack-success rate is meaningless without its denominator and the attacker's budget. Report success per attempt, not per prompt, state how many attempts and whether the attacker adapted, and publish the curve of success against attempts rather than one number.

open as a page

Where do the trust boundaries sit in a retrieval- and tool-using LLM assistant?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Not at the network edge. Everything that lands in the context window is one untrusted zone, whatever its source. The boundaries that matter sit between that context and any privileged action, and between each writer and each data source the system ingests.

open as a page

What must a sandbox provide before an agent may run model-written code?

level: middleimportance: should knowfreq 48%

basics

~20 s

Isolation strong enough to assume the code is hostile: a hardened runtime or microVM rather than a shared process, no host mounts, no ambient cloud credentials or metadata access, default-deny network egress with a narrow allowlist, resource caps, and a fresh instance destroyed after each session.

open as a page

In multimodal moderation, how do you catch harm that lives only in the image?

level: middleimportance: should knowfreq 38%

basics

~20 s

Score the image, not just the words around it. Text-only classifiers pass a clean caption over an unsafe picture, so the pipeline needs image-capable moderation — either a joint text-plus-image call or a dedicated image classifier — plus sampled frames for video rather than one thumbnail.

open as a page

Which internal content is unsafe to place in a system prompt, and what should hold it instead?

level: middleimportance: should knowfreq 48%

basics

~20 s

Treat anything in the context window as disclosable to that session's user: no credentials, no other users' data, no internal policy you would not publish. Secrets belong in the tool-execution layer, access decisions in code, sensitive policy behind an authorized retrieval call.

open as a page

How does the Plan-Then-Execute pattern contain prompt injection, and what does it miss?

level: middleimportance: should knowfreq 44%

basics

~20 s

Plan-Then-Execute fixes the sequence of tool calls before any untrusted content is read, so injected text cannot add or reorder steps. It stops control-flow hijacking. It does not stop untrusted data from poisoning the arguments and content of the steps already planned.

open as a page

What do model-level safety probes miss that agent-level red-team suites catch?

level: middleimportance: should knowfreq 45%

basics

~20 s

Model-level probes such as garak or PyRIT attack a bare model endpoint and measure what it says. Agent-level suites such as AgentDojo or AgentHarm attack the assembled system and measure what it does — which tools fired and whether a harmful task was completed end to end.

open as a page

Why does OWASP publish a separate Agentic Top 10 alongside the GenAI LLM Top 10?

level: middleimportance: should knowfreq 52%

basics

~20 s

The LLM Top 10 models a prompt-in, text-out application. The Agentic Top 10 covers what appears only when a system plans, keeps memory across sessions, calls tools and talks to other agents — goal hijack, tool misuse, memory poisoning, rogue agents.

open as a page

How do you defend an agent against a poisoned tool description from a third-party server?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Treat tool metadata as untrusted code that lands in the model's context: pin servers to a reviewed version or digest, require a diff review before any description change takes effect, fail closed on unexpected changes, and keep capability enforcement at the resource server so no description can grant access.

open as a page

When does a policy-conditioned moderation classifier beat a fixed hazard taxonomy?

level: seniorimportance: should knowfreq 45%

basics

~20 s

When your policy is specific to your product and changes often. A policy-conditioned classifier reads your written policy at inference time, so a rule change is a text edit rather than a labelling-and-retraining cycle. Fixed taxonomies stay cheaper and faster for stable, universal harm categories.

open as a page

When is a declarative guardrail framework worth it over hand-rolled validators?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A framework pays when policy churns, when non-engineers must read or edit it, and when you need many off-the-shelf checks composed the same way. Hand-rolled validators win when you have three rules, tight latency budgets, and no appetite for a second runtime and its DSL.

open as a page

How do you design an LLM refusal so it isn't a dead end for the user?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Treat a blocked turn as a product state, not an error. Name the category of what you cannot do, offer the sanctioned path for it, and carry the user's context into that path. A refusal that ends the conversation converts a policy win into an abandoned session.

open as a page

How do CaMeL-style information-flow controls go beyond the Dual-LLM pattern?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Dual-LLM keeps untrusted text away from the privileged model by passing symbolic references, but nothing checks what the privileged program then does with those values. CaMeL and FIDES attach provenance and permission labels to every value and enforce policy at the point of action.

open as a page

How would you gate a release on an automated red-team suite in CI?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Run the red-team job against every release candidate, block on any new success in high-severity categories rather than on an aggregate score, keep every past incident as a permanent case, and run each case several times because any single success counts as a failure.

open as a page

How should credentials flow from a user through an agent to a third-party tool server?

level: principalimportance: should knowfreq 34%

basics

~20 s

Each hop gets its own token. A tool server must never replay a token minted for something else at a downstream API, and every service must reject tokens whose audience does not name it — otherwise the server becomes a confused deputy acting with authority it was never granted.

open as a page

How do you design human escalation and appeals into a moderation pipeline?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat human review as a component of the system, not a safety net bolted on. Every automated decision records the policy clause and model version behind it, affected users are told which rule they broke, appeals go to a reviewer with that evidence, and reversals feed back into the policy.

open as a page

How do you resolve conflicts when several guardrails judge one LLM response?

level: principalimportance: should knowfreq 30%

basics

~20 s

Give rails distinct verbs and a fixed precedence. Blocks outrank redactions, which outrank rewrites; a block short-circuits. Decide per rail whether it fails open or closed when it errors, run mutating rails in a defined order, and cap any repair attempt so rails cannot fight each other indefinitely.

open as a page

Under the Agents Rule of Two, how would you restructure a triage agent that reads public issues while holding private-repo access?

level: principalimportance: should knowfreq 36%

basics

~20 s

The Rule of Two allows an agent session at most two of: untrusted input, sensitive access, and the ability to change state or communicate externally. A triage agent holding all three should be split into a public-only reader that emits a structured result and a privileged session that never reads issue text.

open as a page

showing 1–30 of 33