skip to content

When refused, absent, and empty-but-successful tool calls all surface as the same bland sentence, how do you build a reliable read/write oracle?

level: seniorimportance: should knowfreq 40%

answer

  1. the sentence carries almost no signal
  2. decline means withheld, not nothing ran
  3. move the signal off-process
  4. an observable the attacker controls
  5. confidence is a repeat rate

basics

~20 s

Stop reading the sentence and design probes where the three cases diverge somewhere observable — an out-of-band signal, timing, or a later-visible effect. A bland reply proves the answer was withheld, not that nothing ran, so you separate the events by side effects, not by wording.

solid answer

~60 s

The failure mode here is silence: a refused call, an absent capability, and a call that ran and returned nothing can all produce the identical bland sentence, so the reply text carries almost no signal. Reading it harder does not help — a decline proves the answer was withheld, not that no tool executed. A reliable oracle comes from arranging for the three cases to differ *outside* the reply. For the network write, an off-process observable — a page authored to be reached that later shows an access — distinguishes 'ran' from 'did not run' regardless of what the assistant says. Timing and ordering can separate a call that executed from one that never dispatched. Because the model is probabilistic, one observation is weak; you repeat until the difference is stable, and you report the oracle's confidence in terms of that repeat rate, not a single success. The core discipline is respecting the direction of each claim — the sentence tells you what was declined, the side effect tells you what happened.

go deeper

for a junior

Understand that a refusal or empty reply does not mean nothing happened; the assistant can act and still say little.

for a middle

Explain why the reply text is weak signal when refused, absent, and silent-success collapse onto one sentence, and why side effects carry more.

for a senior

Design an oracle that separates the cases off-process — an observable the attacker controls — and treat reliability as a repeat rate under a probabilistic model.

for a principal

Be ready to say what such an oracle's confidence is worth in a report, and how to price a finding whose signal was mostly inferred rather than observed.

### Why silence is the hard case When an attacker probes an assistant they cannot inspect, the ideal is an *oracle*: a probe whose result cleanly answers 'does a read reach private context here?' or 'does a write reach the network here?'. The obstacle this question names is that the assistant's surface behaviour collapses several very different internal events onto one output. A capability that is **absent**, a capability that is **present but refused**, and a capability that **ran and returned nothing** can each produce the same flat sentence — 'I wasn't able to do that.' The reply is therefore nearly information-free, and an attacker who scores it as 'no' is often wrong. ### The direction-of-claim discipline The first move is to be strict about what each observation proves. A bland decline proves the *answer was withheld* — not that no tool executed, not that the capability is absent, not that nothing left the process. A response that returns no content proves the *reply* was empty, not that the underlying call did nothing. Conflating 'the assistant said no' with 'nothing happened' is exactly the misread that makes a silent success look like a failure. The whole method follows from refusing that conflation. ### Build the oracle outside the reply If the reply cannot distinguish the cases, arrange for something *else* to. For the write half, the reliable instrument is an off-process observable the attacker controls: a page authored to be arrived at — reached by a browsing agent from a search result rather than handed to it — whose later access is visible to the attacker independent of anything the assistant says. Now 'the call ran' and 'the call never ran' diverge on the attacker's own side, and the bland sentence stops mattering. Timing and ordering are a weaker second instrument: a probe that dispatches a tool and one that does not may differ in latency or in the sequence of what the assistant mentions. The principle is to move the distinguishing signal from the reply — where three cases collapse — to a channel where they separate. ### Separating the three cases A careful probe can pull the cases apart: - **Absent vs present-but-refused.** A capability that is absent cannot produce the off-process effect under any phrasing; one that is present but declining may still produce a partial effect, or its decline may vary in specificity across probes. Repeated, differently-shaped requests expose the difference between 'never possible' and 'possible but withheld.' - **Refused vs silent success.** This is the dangerous pair, because a silent success is a working capability that looks like a dead end. Only an off-process observable settles it: if the effect appears, the call ran regardless of the sentence. ### Probabilism and cost The model is not deterministic, so a single observation is weak evidence. A tell that appears once may not repeat; a genuine capability may decline on one phrasing and comply on another. A reliable oracle is therefore a *repeat rate*, not a one-off — you probe until the difference between the cases is stable, and you price the oracle honestly: how many probes it took, and how much of the conclusion is observation of an off-process effect versus inference from the wording of a reply. An oracle that fires once in five attempts is not the same instrument as one that fires four in five, and a report that hides that distinction overstates what was established. ### What the oracle is and is not Even a clean oracle establishes only presence and reachability: that a read touched private context, or that a write left the process, in this session. It does not itself complete a transfer, and it does not prove the two compose — that still requires showing both halves live together. The value delivered is a confident yes/no plus its cost, which is precisely what a red-team report on this construction is expected to carry.

  • Why is an off-process observable more trustworthy than the assistant's own reply for confirming a write?
    Because the reply is authored by the same system whose behaviour is in question and collapses refused, absent, and silent-success onto one sentence, whereas an observable the attacker controls — a later-visible access to a page authored to be reached — records the fact of the call independent of anything the assistant says. The assistant can stay silent about an action it took; the side effect cannot be un-happened. Trust the channel the target does not narrate.
  • How should the oracle's reliability appear in a red-team report?
    As a rate and a cost, not a claim of certainty. State how many probes established the fact, what fraction of attempts reproduced it, and how much of the conclusion rested on an observed off-process effect versus inference from reply wording. A pair confirmed four times in five with a controlled observable is a far stronger finding than one seen once, and collapsing that difference into a bare 'confirmed' misrepresents what the probing actually bought.

saying these in an interview costs you the question

  • Reads a bland decline as proof no tool executed
  • Treats an empty reply as an empty side effect
  • Trusts the assistant's own wording to confirm a write
  • Reports a single observation as a reliable finding
  • Ignores that the model is probabilistic across probes

context