skip to content

Prompt Injection Defense

Detection and instruction hardening are nudges, not boundaries. Interviewers expect containment: the lethal trifecta, the Agents Rule of Two, and patterns like Plan-Then-Execute and Dual-LLM.

on this pageshow

questions

5

What is the lethal trifecta, and how do you break it in an email assistant?

level: middleimportance: must knowfreq 72%

answer

  1. three capabilities, not one bug
  2. count them per session, not per product
  3. private data plus attacker-controlled text
  4. the third leg is a way out
  5. any two of the three is survivable

basics

~20 s

The lethal trifecta is one agent session holding three things at once: access to private data, exposure to untrusted content, and an outbound channel. Any two are survivable; all three let injected text steal data. Remove one capability.

solid answer

~40 s

The lethal trifecta, named by Simon Willison, is the observation that data theft by prompt injection needs three properties in the same agent session: access to private data, exposure to attacker-controlled content, and some way to send data out. Drop any one leg and an injection may still cause nonsense, but it cannot exfiltrate. An assistant with inbox read, send-mail and web-fetch has all three, so a meeting invite whose body carries instructions can make the agent mail your calendar to an outsider. I use it as a design review gate: for each session, list the three legs honestly — outbound channels include image URLs, webhooks, search queries and anything that touches the network — and if all three are present, the architecture changes, not the system prompt.

go deeper

for a junior

Be able to name the three legs — private data, untrusted content, outbound channel — and say that all three in one agent is the dangerous shape. Give one concrete example, such as an assistant that reads email and can also send it.

for a middle

Explain why any two legs are survivable and the third makes it theft, and enumerate non-obvious channels like rendered image URLs and search queries. Show that you would fix the architecture by cutting a capability rather than by editing the system prompt.

for a senior

Apply the test per session over a real system, including retrieval indexes and third-party tool results as sources of private data and untrusted content. Be ready to propose a concrete cut and to argue why it holds deterministically even when the model is successfully fooled.

for a principal

Own the trifecta as an intake gate every AI feature passes before it ships, and be honest about its limits: it scores exfiltration only, so pair it with a constraint that also counts irreversible state change. Decide which capability the organisation gives up by default.

## What the trifecta names The lethal trifecta is a design test, not an attack. It says that for an injected instruction to actually *steal* something, three properties have to coexist in the same agent session: 1. **Access to private data** — mailboxes, tickets, source code, customer records, anything the attacker does not already have. 2. **Exposure to untrusted content** — any text the agent reads that an outsider can influence: an email body, a web page, a PDF, a calendar invite, a code comment, a tool result. 3. **An ability to communicate externally** — any path by which bytes leave the trust boundary. The framing matters because it moves the conversation away from "is the model gullible?" (it is, and that will not change soon) to "what can a gullible model reach?" ## Why any two legs are survivable An agent with private data and untrusted content but no outbound path can be tricked into producing wrong answers, which is a correctness problem, not a breach. An agent with untrusted content and an outbound path but no private data can be made to send garbage, which is abuse, not theft. An agent with private data and an outbound path but no untrusted content has no attacker in the loop at all. Only the full set turns "the model believed a stranger" into "the stranger got your data." ## The worked example Take an email-and-calendar assistant with three tools: read the inbox, send mail, and fetch a URL for context. That is the textbook trifecta. An attacker sends a meeting invite whose description contains text addressed to the agent rather than the human. When the user asks "what's on my calendar tomorrow?", the agent reads the invite, follows its instructions, and mails the week's meetings to an external address — or, more quietly, fetches an attacker-controlled URL with the data encoded in the query string. The user clicked nothing suspicious; the untrusted content arrived on its own. That zero-click shape is exactly what the EchoLeak class of findings in enterprise assistants demonstrated: ordinary received content is enough to start the chain. ## Finding the legs honestly Most failed reviews are failures of enumeration, not of judgement. - **Private data hides in retrieval.** A RAG index built from internal wikis is private data even though no tool is named "read secrets." - **Untrusted content hides in tool results.** Anything an agent reads back from the outside world is attacker-influenceable, including search snippets and error messages from third-party APIs. - **Outbound channels hide everywhere.** Rendering a markdown image whose URL the model chose is an outbound channel. So is a web search, a DNS lookup, a support-ticket comment, a git push, a webhook, and a "harmless" link the user is likely to click. If you cannot draw the three sets for a feature, you cannot claim it is contained. ## Using it as a gate Run the test per **session**, not per product. A product may legitimately do all three things as long as no single session holds all three. Common cuts: - **Cut the outbound leg** for the session that reads untrusted content — let it produce only a structured result that a separate, non-network step consumes. - **Cut the private-data leg** — do the untrusted reading in a session whose credentials see nothing sensitive. - **Cut the untrusted leg** — pin the privileged session to content you control, and have a quarantined step summarise the rest into a constrained shape. The important property of all three cuts is that they are deterministic. They hold whether or not the model was fooled, which is the whole point: model-layer mitigations are probabilistic and an attacker who can retry will eventually find the phrasing that gets through. ## What the trifecta does not cover The trifecta is scoped to **exfiltration**. An agent with untrusted input and write access but no private data is outside it, and can still delete a production table or post a defamatory message. That is why the broader pick-two constraints treat "changes state" as a risk leg in its own right, and why the trifecta is a first-pass filter rather than a complete threat model. It is also silent about how you *implement* the cut — which credentials, which sandbox, which egress policy — those are separate engineering problems. What it gives you is a five-minute review question that reliably catches the architectures that cannot be saved by prompting.

  • Your agent has no send-mail tool at all — why might the outbound leg still be present?
    Because exfiltration channels are rarely named as such. A rendered markdown image, a fetched URL, a web-search query, a DNS resolution, an outbound webhook, or a link the model writes into the UI for a human to click all move bytes to a party the attacker controls. Enumerate every path by which model-chosen content reaches the network or a human's browser, not just the tools whose names sound like sending.
  • If two legs are genuinely required by the product, how do you keep the feature?
    Split the work across sessions so no single session holds all three. Do the untrusted reading in a session with no private-data credentials and no network egress, and have it emit a constrained, typed result. A second session that never sees the raw untrusted text consumes that result and holds the privileges. The cut is architectural, so it holds even when the model is fooled.
  • Does the trifecta apply to a purely conversational assistant with no tools?
    Only partly. With no tools, a user pasting hostile text can still extract whatever is in the context window and read it themselves — but they are the trust boundary, so nothing crossed it. The trifecta becomes urgent the moment the agent reads content on the user's behalf or acts on the user's behalf, because then the attacker and the victim are different people.

saying these in an interview costs you the question

  • Says a stronger system prompt or a warning closes the trifecta
  • Counts only explicitly named send tools as outbound channels
  • Treats internal wiki or RAG content as trusted because it is internal
  • Assumes the user must click something for injection to fire
  • Applies the test to the whole product rather than to each session

context

open as a page

Why is a prompt-injection detection classifier not a security boundary?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Detection is probabilistic and only has to be wrong once, while an attacker can retry and tune against it. Adaptive attacks against a dozen published defenses succeeded over 90% of the time. Filters reduce noise; a deterministic architectural limit is the boundary.

open as a page

How does the Plan-Then-Execute pattern contain prompt injection, and what does it miss?

level: middleimportance: should knowfreq 44%

basics

~20 s

Plan-Then-Execute fixes the sequence of tool calls before any untrusted content is read, so injected text cannot add or reorder steps. It stops control-flow hijacking. It does not stop untrusted data from poisoning the arguments and content of the steps already planned.

open as a page

How do CaMeL-style information-flow controls go beyond the Dual-LLM pattern?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Dual-LLM keeps untrusted text away from the privileged model by passing symbolic references, but nothing checks what the privileged program then does with those values. CaMeL and FIDES attach provenance and permission labels to every value and enforce policy at the point of action.

open as a page

Under the Agents Rule of Two, how would you restructure a triage agent that reads public issues while holding private-repo access?

level: principalimportance: should knowfreq 36%

basics

~20 s

The Rule of Two allows an agent session at most two of: untrusted input, sensitive access, and the ability to change state or communicate externally. A triage agent holding all three should be split into a public-only reader that emits a structured result and a privileged session that never reads issue text.

open as a page