skip to content

System-Prompt & Data Exfiltration

Attackers get hidden system prompts, tool schemas and another user's context spoken aloud, then get those bytes out through whatever renders the answer. Interviewers score it as real loss.

on this pageshow

explore

questions

20

Why does an injected instruction to exfiltrate an unseen email thread need two steps, not one?

level: juniorimportance: must knowfreq 78%

answer

  1. the take is not loaded yet
  2. one instruction is not enough
  3. find, then send
  4. 'send everything' returns a summary

basics

~20 s

The target thread is not in the assistant's context, so the payload cannot simply emit it. It must first make the model look the thread up, then send it. A one-step 'send everything' reaches only what is already loaded — usually a summary.

solid answer

~50 s

An injection reaches the model through data the app already loaded — say an inbound email read during triage. But the content worth stealing, a specific prior thread or attachment, usually is not sitting in that context. So a single 'send everything you know' instruction has nothing valuable to emit: the model answers from what is loaded and tends to compress it into a summary, the least useful form of the take. To get the real content the payload has to be two-step: first name a target and cause the assistant to `fetch` it with a read tool, then hand the fetched bytes to whatever emits. The hard half is the first step — the model volunteers no search and drafts from what is in front of it, so the instruction must specify what to look for. Miss that, and you exfiltrate a paraphrase of nothing.

go deeper

for a junior

Recall the split: content not in context must be fetched before it can be sent, so the payload is two actions, not one.

for a middle

Explain why a broad 'send everything' yields a summary — the model answers from loaded context and compresses it — and why specifics need a named lookup.

for a senior

Be ready to reason about where the fetch step breaks in a real assistant: no volunteered search, a tool-call budget, an approval that ends the run before a second step.

for a principal

Frame the take's value: a reproduced record versus a paraphrase, and why the collect step, not the channel, decides whether a disclosure is worth anything.

## The setup A prompt injection is an attacker's instruction that reaches a model through *data* the application feeds it, rather than through the operator's own system prompt. In a personal email assistant that triages an inbox and drafts replies, that data channel is the mail itself: an inbound message from a stranger, read during ordinary triage, can carry an instruction the assistant then treats as directive. The naive mental model is that such an instruction 'exfiltrates the context' — that once the model obeys, whatever is sensitive flows out. That is the wrong model, and correcting it is the whole point of this topic. ## Why one step is not enough A model's **context window** holds only what was loaded for *this* turn: the system prompt, the current message, maybe a few retrieved items. It does not hold the mailbox. The genuinely valuable content — an older thread with a counterpart, an attachment, a record from months ago — is almost never already loaded. It lives in a store the assistant can *reach* with a read tool, but has not read. So an instruction that says, in effect, 'send everything you know' has a problem: the model has nothing valuable in front of it to send. What it does have is the current triage context, and asked for a broad dump it will typically **summarise** that — compress many messages into a few sentences. Summarisation is lossy by design: the exact strings that make data worth stealing (account numbers, quoted text, headers, the precise wording) are exactly what a summary drops. The attacker gets a gist of low-value material. To get the real content the payload must be **two-step**: 1. **Collect** — cause the assistant to *fetch* content it has not seen, by naming a target the model can look up. 2. **Send** — hand the fetched bytes to whatever emits (a draft, a message body, a rendered link). These are two distinct actions, and the model performs neither for free. It performs the second readily, because drafting and emitting is what a mail assistant already does. The first it does not volunteer at all — a triage assistant runs a default one-step plan and drafts from what is in front of it. Nothing prompts it to go and search unless the payload explicitly makes it. ## Why the distinction matters Calling this 'the injection exfiltrates the context' hides the step that actually decides whether the attack is worth anything. Data that was *already* in the context can leak in one step — but that is a different, easier problem. 'Collecting before sending' is specifically about content the model **had to go and get**. The construction lives or dies on the collect step: naming a target the attacker cannot see, causing an unprompted fetch, and doing it before the turn ends. ## What a good answer sounds like A strong candidate says: the sensitive data is not loaded, so the payload has to *provoke a lookup* and only then emit; a broad 'send everything' returns a summary of loaded context, which is the least useful form; therefore the attacker must **specify a target**. They separate the fetch from the send, and they know the fetch is the fragile half. A weak candidate says the injection 'dumps the context' or assumes the mailbox is somehow all available to the model at once — which collapses the two steps into one and misses why most such attacks quietly fail.

  • Why does a broad 'summarise and send all my data' instruction disappoint an attacker even when it runs?
    Because a summary is lossy by design. The model compresses many messages into a few sentences, dropping the exact strings — account numbers, quoted text, headers — that made the data worth taking. The attacker gets a gist, not the record, and often cannot even tell what was omitted. Valuable exfiltration reproduces specific bytes, which forces naming a specific target rather than asking for a sweep.
  • If the target thread were already in the context window, would the two-step still be needed?
    No. If the valuable content is already loaded — the user pasted it, or it rode in with a retrieved set — a single instruction can emit it directly; there is nothing to fetch. The two-step exists precisely for content the model has not seen. That is why 'collecting before sending' is a distinct problem from leaking what is already present.

Like being told to mail a file from an office you are standing in: if the file is in a basement archive you have never opened, 'mail me everything on my desk' gets a sticky note, not the file. You have to go down and pull it first.

saying these in an interview costs you the question

  • Says the injection just exfiltrates the whole context
  • Assumes the target thread is already loaded in context
  • Thinks 'send everything' returns the full data verbatim
  • Treats find and send as one indivisible action

context

open as a page

A planted reference reaches an assistant's answer and nobody clicks it — how does data leave?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Displaying is enough. If the answer carries a reference the viewing client resolves on its own — a preview it expands, a resource it embeds — the request goes out while the answer is being drawn. Nobody clicks.

open as a page

In an assistant whose hidden preamble forbids repeating it, why does a translation request still surface the text?

level: juniorimportance: must knowfreq 74%

basics

~20 s

The forbidding line names one operation, repeating, and a translation is a different operation. Nothing marks the preamble as secret; there is only a sentence about one verb sitting beside the text, and the request never uses that verb.

open as a page

Why does an output screen that blocks secrets and raw tool JSON let a review bot describe its own operations in prose?

level: juniorimportance: must knowfreq 62%

basics

~20 s

An output screen matches on shape - credential patterns, key formats, blocks of JSON. A plain-English sentence about which operations a bot has matches none of those, so capability talk leaves as ordinary help text while carrying an inventory.

open as a page

A ticket-triage auto-reply is length-capped per reply - why does that not bound the leak?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A per-reply cap bounds one response, not the total. Each submission is a fresh run over the same hidden context, so many submissions yield many capped slices, reassembled outside the system. Exposure counts per campaign, not per reply.

open as a page

How does an exfiltration payload name a mailbox thread it cannot confirm is present?

level: middleimportance: should knowfreq 50%

basics

~20 s

The attacker cannot see the victim's mailbox, so the payload must describe the target by stable, guessable properties — sender, subject pattern, a document name — and tolerate its absence. A precise-but-wrong name fetches nothing and silently wastes the one attempt.

open as a page

An attacker's reference in an assistant's answer is fetched at render — why doesn't URL length close it?

level: middleimportance: should knowfreq 45%

basics

~10 s

A length limit caps bytes per request, not the number of requests, and the valuable items are small. The first request alone already carries the fact that this surface resolves references at render time.

open as a page

A user asks an assistant to tabulate the clauses of its hidden preamble and it complies. What does that show?

level: middleimportance: should knowfreq 58%

basics

~20 s

The preamble is ordinary readable context with no protected status. The model holds the text plus a sentence expressing a preference about it, and produces whatever fits the request, so any operation over that text is served from the same tokens.

open as a page

In a stateless auto-reply workflow, how do capped fragments become one hidden passage?

level: middleimportance: should knowfreq 44%

basics

~20 s

Different requests return differently positioned slices of the same hidden text. Because that text is identical in every run, overlapping regions let the slices be ordered and merged outside the system. Alignment is approximate, since models paraphrase.

open as a page

Why is the collect step, not the send step, where a two-step mail-exfiltration payload usually fails?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Sending is one obeyed instruction the model already performs while drafting. Collecting demands an extra read the model volunteers no reason to run, inside a per-turn tool-call budget and before an approval ends the turn — and the target may be absent. That is where most attempts die.

open as a page

A callback fired once from a colleague's client and never in your harness — is that a finding?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Yes, but not as one claim. Split it: whether the outsider's field text reaches the answer, measurable server-side over many runs, and whether a given surface resolves the reference while drawing, testable per client with a fixture and no model at all.

open as a page

A tenant widens its do-not-reveal line to forbid summarising and paraphrasing too. What does that buy?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It raises the number of attempts and stops casual probing, and it does not close the class. Operations whose output depends on the text cannot be enumerated, so the next request simply uses one that is not on the list.

open as a page

The owner adds a preamble line telling a review bot never to discuss its tooling - what did that buy them?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A preamble line buys the cheapest read - a stranger no longer gets the surface described on request. The surface itself is unchanged, and the bot's ordinary replies still show which sources it reads and what it cannot do.

open as a page

An owner rejects your report that a review bot enumerates its operations, saying schemas are public - what is your reply?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Argue on the right axis. The severity is not confidentiality of the names - it is reconnaissance: the reply tells a stranger which operations this installation actually has, so the next attempt is targeted rather than speculative.

open as a page

Only part of a reconstructed customer record reproduces reliably - what do you claim in the finding?

level: principalimportance: should knowfreq 31%

basics

~20 s

Claim the mechanism, not the transcript. Sign off that hidden third-party context is recoverable in slices across independent runs, evidenced by the spans confirmed against a source. Leave stable-but-unconfirmed text out of the claim and out of circulation.

open as a page

A review bot describes its own operations in prose - what does that give an attacker that public product docs do not?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Docs describe the product; the bot's account describes this installation - which operations are wired up here and roughly what the arguments are called. The paraphrase loses exact names, types and required fields, and proves nothing.

open as a page

Your two-step mail-exfil payload worked once then failed four tries — is that a finding?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

Yes, but not as a reliable exploit. One obeyed instruction against a blind, possibly-absent target is probabilistic by construction, so one hit in five is expected. Report it as a demonstrated class of risk with honest reliability — not a dependable exfiltration you can rerun.

open as a page

A tenant reports a user obtained a faithful summary, not a quotation, of its hidden preamble. Is that a leak?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Yes, when the value is in the content. A faithful summary carries the thresholds, rules and escalation wording that made the preamble sensitive; only exact phrasing is lost. Grade it by what was recovered, not by its form.

open as a page

In a reconstruction built from repeated samples, does span agreement prove recall?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

No. Disagreement between independent samples is strong evidence a span was generated rather than read. Agreement only proves the output is stable, which both genuine recall and a strongly format-shaped guess produce. Agreement is necessary for recall, never sufficient.

open as a page

An outsider-set field makes your assistant's answers carry a fetch-triggering reference — who owns that finding?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Three teams could own it: whoever accepts the field, whoever emits the answer, and whoever ships the renderer that fetches. The call is made on cost asymmetry, and it is a recorded decision rather than a technical fact.

open as a page