skip to content

From an assistant's replies alone, how would an attacker infer which of its tools reads private data and which reaches the network?

level: middleimportance: should knowfreq 50%

answer

  1. you can't see the config
  2. infer capabilities from behaviour
  3. steer toward a task per capability
  4. content only a private read explains
  5. confirm both live in one session

basics

~20 s

By differential probing: steer the assistant toward tasks that would each need a private-context read or a network write, then read the replies for tells — content that could only come from a private source proves the read; an observable outbound effect proves the write. Behaviour reveals the set the attacker cannot see.

solid answer

~50 s

The attacker cannot see the configuration, so they infer it from behaviour. They ask the assistant for things that would require each capability in turn: a task only answerable from private context, and a task whose only completion is an outbound request. A reply that reflects private content that was never supplied in the conversation is evidence the read reached private context in that session; an effect visible off-process — a page authored to be reached that later shows a visit, or any out-of-band signal — is evidence the write left the process. The target is not either capability in isolation but their coexistence in **one** session, so the probes must establish both halves without the assistant resetting context between them. Crucially, they are reading tells, not configuration: what the assistant volunteers, how specifically it phrases a decline, and content it could only have obtained one way. None of this completes an actual transfer; it only establishes that the pair is present and usable.

go deeper

for a junior

Know that an attacker cannot see the configuration and must infer capabilities from how the assistant behaves, not from any settings page.

for a middle

Explain differential probing: design tasks that each need one capability, read the reply against a baseline, and use content only a private read explains as evidence.

for a senior

Discuss confirming both halves within a single session and using an out-of-band observable for the write, plus how a probabilistic model forces repeated probes.

for a principal

Be ready to weigh how much of such an inference is observation versus guesswork, and what that uncertainty does to the confidence you can report.

### The problem the attacker is solving The attacker wants to know two facts about a deployment they cannot inspect: does this assistant have a capability that **reads** something the operator treats as private, and does it have a capability that **writes** to somewhere off the process — and are both live in the **same session**? Configuration is invisible from outside, so the only instrument is the assistant's own behaviour. This is inference from responses, and it is a discovery task, not an exploitation one: the goal here is knowledge, not a completed transfer. ### Differential probing The method is to design requests that *would* exercise a given capability and watch how the reply differs from a baseline that would not. Two families of tell matter: - **For the private read.** Ask for something that can only be answered from private context — a detail that was never placed in the conversation and is not general knowledge. If the reply contains that detail, the read reached private context in this session. If it declines or produces only generic content, either the capability is absent, or it is present but declined to use it — an ambiguity that has to be resolved separately. - **For the network write.** Ask for a task whose only faithful completion involves an outbound action, and arrange for that action to be observable off-process. In the browsing-agent setting, the classic observable is a page authored to be *arrived at* — reached by the agent from a search result rather than handed to it — whose later access tells the attacker a fetch occurred. The attacker never touches the request path; they learn of the visit from what comes back to them later. Comparing the probed reply against the baseline is what turns a bland sentence into a bit of signal — hence *differential*. ### Reading tells, not configuration The useful tells are behavioural: content the assistant volunteers that it could only have from a private source; the *specificity* of a decline (a refusal that names a constraint reveals more than a generic one); ordering and latency that suggest a tool ran; and any downstream effect the attacker can observe independently. The attacker is assembling a picture of the tool set from the shape of responses, the way one maps a service from its error messages rather than its source. ### The one-session requirement The payoff is specifically the *pair in one session*. A read capability proven in session A and a write proven in session B do not compose if the deployment never puts them together — so the probing has to establish both halves without the context resetting in between. This is what separates this discovery from merely cataloguing an assistant's tools: the attacker is confirming a co-location, not a feature list. ### The direction of the evidence Every inference here has a strict direction, and getting it backwards is the common error. An assistant answering from private content proves the read reached private context *in that session* — not that a store was breached and not, on its own, that a write is also present. An observed fetch proves a network write occurred — not which tool issued it or what data it could carry. A decline proves the answer was withheld, not that the capability is absent. Because the model is probabilistic, a tell that appears once may not appear again; a single confirming reply is weaker evidence than a repeatable difference, which is why probing tends to cost multiple attempts per fact. ### Where it stops The method depends on responses carrying tells. A deployment whose replies are uniformly bland — the same sentence whether a call ran, was refused, or is absent — starves the attacker of signal, and separating those cases becomes the hard part (the subject of the silence question on this topic). And if the read and write never share a session, no amount of probing composes them, because the thing being discovered does not exist.

  • How does the attacker learn a network write happened when the write itself returns nothing useful?
    Through an out-of-band observable rather than the reply. In the browsing-agent case, a page authored to be reached from a search result later shows an access, so the attacker learns of the fetch from their own side, not from the assistant. The reply may say nothing; the evidence lives off-process. This is why the write half is often confirmed by watching for a downstream effect the attacker controls, not by parsing the assistant's sentence.
  • Why is it not enough to confirm the read and the write in separate sessions?
    Because the vulnerability is the co-location, not the two features. If the deployment reads private context in one session and reaches the network in another, an attacker who can only act within a single session has no path from one to the other. The probing therefore has to show both capabilities live together, without the context resetting between them, or it has proven only that the assistant owns two tools — not that they compose.

saying these in an interview costs you the question

  • Assumes the attacker can read the tool configuration
  • Confirms each capability but ignores the one-session requirement
  • Treats a private-sounding reply as proof a store was breached
  • Reads a bland decline as proof a capability is absent
  • Believes a single confirming reply is reliable from a probabilistic model

context