skip to content

Your two-step mail-exfil payload worked once then failed four tries — is that a finding?

level: seniorimportance: nice to knowfreq 35%

answer

  1. probabilistic by construction
  2. one shot per victim
  3. one success is not a rerun
  4. report the pathway and the denominator

basics

~20 s

Yes, but not as a reliable exploit. One obeyed instruction against a blind, possibly-absent target is probabilistic by construction, so one hit in five is expected. Report it as a demonstrated class of risk with honest reliability — not a dependable exfiltration you can rerun.

solid answer

~50 s

It is a finding, but you have to characterise it honestly. This construction is probabilistic by design: one instruction, one shot per victim, a target you named blind that may or may not exist, and a collect step that can be truncated by a budget or a deadline. So a low hit rate is not evidence the finding is invalid — it is the expected shape of the method. What you must not do is round one success up to a repeatable exploit; and what you must not do is dismiss it because it failed four times. Report the mechanism — that an injected instruction can cause the assistant to fetch and emit content the user never surfaced — the conditions under which it fires, and the measured reliability with its denominator. A triager should understand the risk is real and the exploitation is opportunistic, because a single unflagged success can still remove real data. The value is the demonstrated pathway, not the win rate.

go deeper

for a junior

Know that these attacks are chancy: a low success rate does not by itself mean there is no vulnerability.

for a middle

Explain why the method is probabilistic — blind target, one shot, a truncatable collect step — so variance is expected rather than disqualifying.

for a senior

Characterise the finding: report the demonstrated pathway, the firing conditions and the measured reliability with its denominator, without over- or under-claiming.

for a principal

Own the judgement on what a mostly-silent method is worth to a programme, and why a single unflagged success justifies treating an opportunistic pathway as real risk.

## The question behind the question An interviewer who asks this is testing whether you can characterise a **probabilistic finding** honestly — neither inflating one success into a reliable exploit nor discarding a real pathway because it mostly failed. On this class of attack that judgement is the whole job, because the method is chancy by construction. ## Why this method is probabilistic by construction Several independent sources of variance stack up: - **One obeyed instruction.** The payload is a single injected instruction; there is no interactive loop to adjust. - **One shot per victim.** Realistically you get one meaningful attempt before the mail is read, deleted, or the behaviour is noticed. You cannot grind. - **Blind targeting.** The target is named without visibility; it may not exist in this mailbox at all, so some attempts fetch nothing for reasons unrelated to the payload's quality. - **A truncatable collect step.** A per-turn tool-call budget or an approval deadline can end the run before the fetch completes, independent of anything else. Given all that, **one success in five attempts is an expected shape**, not an anomaly. A deterministic bug either reproduces or it does not; this does not behave that way, and treating it as if it should is the mistake. ## The two failure modes of the triager, and of you There are two ways to report this badly: 1. **Overclaim.** Rounding a single win up to 'a reliable exfiltration exploit' is dishonest and erodes trust; the next reviewer who cannot reproduce it discounts everything you file. The method is opportunistic and the report should say so. 2. **Underclaim / dismiss.** Throwing the finding out because it failed four times hides a real pathway. A single unflagged success can still remove real data from a real mailbox. Variance is not the same as absence of risk. ## How to report it honestly Separate the **mechanism** from the **hit rate**: - **Mechanism / pathway.** State plainly what was demonstrated: an injected instruction caused the assistant to fetch content the user never surfaced and emit it. That pathway either exists or it does not, and here it does. - **Firing conditions.** Name what has to be true for it to succeed: the target present, the descriptor matching, the fetch completing inside the budget and before approval. - **Reliability with its denominator.** Give '1 of 5', not a bare '20%' floating free — the denominator and the conditions let a reader weigh it. A fraction with no context misleads in either direction. The triager should come away understanding that the *risk is real* and the *exploitation is opportunistic*: not something an attacker fires with confidence, but something that, when it lands unnoticed, is a genuine loss. ## Why you cannot just rerun for confidence With ordinary code you raise confidence by re-running until the behaviour is stable. Here each run is a fresh blind attempt against a possibly-different mailbox state, and re-running does not converge the way a deterministic exploit does. You also may not *get* many honest reruns per victim. So confidence comes from understanding the conditions that make it fire — the mechanism — not from a pass/fail count you would trust for deterministic software. ## What a good answer sounds like A strong candidate says 'yes, it is a finding' immediately, then explains *why the low rate is expected* and *how to report it without over- or under-claiming*: mechanism, conditions, reliability with a denominator. A weaker candidate either dismisses it ('it only worked once, not real') or overstates it ('it works, ship the exploit'), both of which misread a probabilistic finding.

  • How do you keep a reviewer from dismissing a one-in-five result as noise?
    Separate the mechanism from the hit rate. Show the pathway is real — an injected instruction demonstrably caused a fetch-and-emit of unsurfaced content — and that the low rate comes from named conditions: blind targeting, target absence, a truncating budget. Give the denominator and the conditions rather than a bare fraction. A single success that removes real data is a real loss; weigh the pathway's existence, not the variance.
  • Why can you not just rerun the payload to raise confidence, as with a deterministic bug?
    Because each run against a victim is a fresh blind attempt with a possibly-different mailbox state, and re-running the same instruction does not converge on a stable outcome the way a deterministic exploit does. You may also get only one realistic shot per victim before the attempt is noticed. Confidence comes from understanding the conditions that make it fire, not from a pass/fail rerun count you would trust for ordinary code.

saying these in an interview costs you the question

  • Rounds one success up to a reliable, rerunnable exploit
  • Dismisses the finding because it mostly failed
  • Reports a bare hit rate with no conditions or denominator
  • Treats a probabilistic finding like a deterministic bug

context