Why does an injected instruction to exfiltrate an unseen email thread need two steps, not one?
answer
- the take is not loaded yet
- one instruction is not enough
- find, then send
- 'send everything' returns a summary
basics
~20 sThe target thread is not in the assistant's context, so the payload cannot simply emit it. It must first make the model look the thread up, then send it. A one-step 'send everything' reaches only what is already loaded — usually a summary.
solid answer
~50 sAn injection reaches the model through data the app already loaded — say an inbound email read during triage. But the content worth stealing, a specific prior thread or attachment, usually is not sitting in that context. So a single 'send everything you know' instruction has nothing valuable to emit: the model answers from what is loaded and tends to compress it into a summary, the least useful form of the take. To get the real content the payload has to be two-step: first name a target and cause the assistant to `fetch` it with a read tool, then hand the fetched bytes to whatever emits. The hard half is the first step — the model volunteers no search and drafts from what is in front of it, so the instruction must specify what to look for. Miss that, and you exfiltrate a paraphrase of nothing.
go deeper
Recall the split: content not in context must be fetched before it can be sent, so the payload is two actions, not one.
Explain why a broad 'send everything' yields a summary — the model answers from loaded context and compresses it — and why specifics need a named lookup.
Be ready to reason about where the fetch step breaks in a real assistant: no volunteered search, a tool-call budget, an approval that ends the run before a second step.
Frame the take's value: a reproduced record versus a paraphrase, and why the collect step, not the channel, decides whether a disclosure is worth anything.
## The setup A prompt injection is an attacker's instruction that reaches a model through *data* the application feeds it, rather than through the operator's own system prompt. In a personal email assistant that triages an inbox and drafts replies, that data channel is the mail itself: an inbound message from a stranger, read during ordinary triage, can carry an instruction the assistant then treats as directive. The naive mental model is that such an instruction 'exfiltrates the context' — that once the model obeys, whatever is sensitive flows out. That is the wrong model, and correcting it is the whole point of this topic. ## Why one step is not enough A model's **context window** holds only what was loaded for *this* turn: the system prompt, the current message, maybe a few retrieved items. It does not hold the mailbox. The genuinely valuable content — an older thread with a counterpart, an attachment, a record from months ago — is almost never already loaded. It lives in a store the assistant can *reach* with a read tool, but has not read. So an instruction that says, in effect, 'send everything you know' has a problem: the model has nothing valuable in front of it to send. What it does have is the current triage context, and asked for a broad dump it will typically **summarise** that — compress many messages into a few sentences. Summarisation is lossy by design: the exact strings that make data worth stealing (account numbers, quoted text, headers, the precise wording) are exactly what a summary drops. The attacker gets a gist of low-value material. To get the real content the payload must be **two-step**: 1. **Collect** — cause the assistant to *fetch* content it has not seen, by naming a target the model can look up. 2. **Send** — hand the fetched bytes to whatever emits (a draft, a message body, a rendered link). These are two distinct actions, and the model performs neither for free. It performs the second readily, because drafting and emitting is what a mail assistant already does. The first it does not volunteer at all — a triage assistant runs a default one-step plan and drafts from what is in front of it. Nothing prompts it to go and search unless the payload explicitly makes it. ## Why the distinction matters Calling this 'the injection exfiltrates the context' hides the step that actually decides whether the attack is worth anything. Data that was *already* in the context can leak in one step — but that is a different, easier problem. 'Collecting before sending' is specifically about content the model **had to go and get**. The construction lives or dies on the collect step: naming a target the attacker cannot see, causing an unprompted fetch, and doing it before the turn ends. ## What a good answer sounds like A strong candidate says: the sensitive data is not loaded, so the payload has to *provoke a lookup* and only then emit; a broad 'send everything' returns a summary of loaded context, which is the least useful form; therefore the attacker must **specify a target**. They separate the fetch from the send, and they know the fetch is the fragile half. A weak candidate says the injection 'dumps the context' or assumes the mailbox is somehow all available to the model at once — which collapses the two steps into one and misses why most such attacks quietly fail.
- Why does a broad 'summarise and send all my data' instruction disappoint an attacker even when it runs?Because a summary is lossy by design. The model compresses many messages into a few sentences, dropping the exact strings — account numbers, quoted text, headers — that made the data worth taking. The attacker gets a gist, not the record, and often cannot even tell what was omitted. Valuable exfiltration reproduces specific bytes, which forces naming a specific target rather than asking for a sweep.
- If the target thread were already in the context window, would the two-step still be needed?No. If the valuable content is already loaded — the user pasted it, or it rode in with a retrieved set — a single instruction can emit it directly; there is nothing to fetch. The two-step exists precisely for content the model has not seen. That is why 'collecting before sending' is a distinct problem from leaking what is already present.
Like being told to mail a file from an office you are standing in: if the file is in a basement archive you have never opened, 'mail me everything on my desk' gets a sticky note, not the file. You have to go down and pull it first.
saying these in an interview costs you the question
- Says the injection just exfiltrates the whole context
- Assumes the target thread is already loaded in context
- Thinks 'send everything' returns the full data verbatim
- Treats find and send as one indivisible action