Why is the collect step, not the send step, where a two-step mail-exfiltration payload usually fails?
answer
- sending is the easy half
- the model volunteers no lookup
- one tool call, then approval ends it
- no target, no take
basics
~20 sSending is one obeyed instruction the model already performs while drafting. Collecting demands an extra read the model volunteers no reason to run, inside a per-turn tool-call budget and before an approval ends the turn — and the target may be absent. That is where most attempts die.
solid answer
~50 sThe send half is cheap: the assistant is already drafting a message, so emitting content into it is within its normal behaviour. The collect half is where the construction is fragile, for several compounding reasons. A triage assistant runs a default one-step plan — it drafts from what is in front of it and volunteers no lookup — so the payload has to actively provoke a search the model had no reason to do. That search costs a tool call against a per-turn budget, and it must complete before the send-approval step ends the run; a plan that needs two tool calls where one is allowed simply never reaches the send. Then there is the blind target: even a provoked, budgeted search returns nothing if the named thread is not in the mailbox. So an attacker pricing this counts most attempts as silent losses at the collect step — not at the channel, which is the part people focus on.
go deeper
Recall that the two halves differ: sending rides on work the assistant already does; collecting needs an extra, unprompted action.
Explain why the collect step is provoked, not free — the default plan is one-step and volunteers no lookup — and how a tool-call budget limits chaining.
Diagnose a real assistant: identify whether the budget, the approval deadline, or an absent target is dropping the fetch, and price attempts as mostly-silent losses there.
Argue that reliability is set by the collect step, so a red-team estimate that models only the channel overstates the method's dependability.
## Two halves, very different odds A collect-before-send payload has two actions: fetch content the model has not seen, then emit it. People new to this fixate on the *send* — the egress channel, the rendered link, the draft that leaves. But in practice the send is the cheap half, and the collect is where the construction breaks. Understanding why is a senior-level diagnosis. ## Why sending is cheap A mail assistant's whole job is to draft and, on approval, send messages. Emitting content into an outgoing draft is squarely inside its normal behaviour; the model does not need to be coerced into producing output. If the bytes are in hand, getting them into something that leaves is not the bottleneck. ## Why collecting is fragile — four compounding reasons **1. No volunteered search.** A triage assistant runs a *default one-step plan*: read the current message, draft from what is in front of it, stop. It has no reason to go looking for other threads and does not do so unless something makes it. The payload therefore has to actively provoke a lookup the model would never have performed — and provoking an unprompted action is harder and less reliable than riding one the assistant was already going to take. **2. A per-turn tool-call budget.** Assistants commonly cap how many tool calls a turn may make before they draft and stop. If the cap is one call, a payload that needs a *search* and *then* a *send* has already overrun: the search consumes the call, and the turn ends before anything is emitted — or the send fires with nothing collected. The tighter the budget, the more the collect has to happen inside an action the assistant was already taking, which is a much harder payload to build. **3. The approval deadline.** Many mail assistants stop at a human approval boundary before a message actually goes out. That boundary matters to the attacker not because they must defeat it, but because it **ends the run**. If the fetch did not finish before the assistant reached draft-and-confirm, there is nothing collected to send — approved or not. The deadline truncates the plan. **4. Blind target absence.** Even a provoked, budgeted, in-time search returns nothing if the named thread is not in the mailbox. Blind targeting means many attempts fetch empty simply because the guessed descriptor did not match anything real. ## The consequence for pricing An attacker who has built this counts most attempts as **silent losses at the collect step**. The send channel — the thing defenders and juniors alike tend to scrutinise — is rarely the reason an attempt failed. It failed because the fetch was never provoked, or overran the budget, or missed the deadline, or found nothing. A red-team estimate that models only the channel will *overstate* the method's dependability, because it is measuring the easy half. ## What a good answer sounds like A strong candidate inverts the naive focus: they say the send is normal behaviour and cheap, then enumerate why the collect is fragile — unprompted, budgeted, time-boxed, and dependent on a target that may not exist. They can say *which* of those is likely dropping the fetch in a given assistant, and they price attempts accordingly. A weaker candidate keeps pointing at the egress channel, or assumes the model will search on its own, and cannot explain why real attempts mostly fail before anything is ever emitted.
- How does a per-turn tool-call budget interact with a two-step payload?It caps how much the payload can chain. If the assistant allows one tool call before it drafts and stops for approval, a payload that needs a search and a send has already overrun: the search consumes the call and the run ends before emitting, or the send happens with nothing collected. The tighter the budget, the more the collect has to happen inside an action the assistant was already going to take.
- Why does the send-approval step matter to the attacker even though approval is about output?Because approval ends the run. The attacker does not need to defeat the approval to have already failed — if the collect step did not finish before the assistant reached the draft-and-confirm boundary, there is nothing collected to send, approved or not. The deadline truncates the plan. Even a perfect descriptor loses if the fetch could not complete inside the turn.
saying these in an interview costs you the question
- Thinks the egress channel is the main failure point
- Assumes the model will search on its own initiative
- Ignores the per-turn tool-call budget entirely
- Treats the approval click as the only obstacle