Why can an attacker exploit a read and a write that each passed an AI assistant's per-capability review?
answer
- one tool fine, two together not
- the risk is in neither part
- review sees items, not the set
- the pair is born at deployment
- assembled by someone reviewing nothing
basics
~20 sBecause the danger lives in the pair, not either tool. A per-capability review judges each on its own and never sees the combination, and the combination is assembled later, when a deployment wires a private-context read and a network-reaching write into the same session.
solid answer
~50 sA single read that reaches private context is fine; a single write that reaches the network is fine; the two in one session are what let attacker-controlled text pull private data toward an outbound call, with the model as the courier. A review that signs off tools one at a time approves each on its own merits and has nowhere to see the set — because the set does not exist at review time. It is created when whoever configures the deployment enables both capabilities for the same assistant, and that person is wiring features together, not adjudicating risk. So the exploitable combination is born after every review has passed, and no individual approval was wrong. The attacker does not defeat a control; they notice that the pair coexists. The composition is the vulnerability, and it belongs to the deployment, not to either tool's author.
go deeper
Be ready to state that the risk is a property of the combination — a read that reaches private context plus a write that reaches the network in one session — not of either tool alone.
Explain why a per-capability review is structurally blind: its unit is the single tool, but the vulnerability's unit is the set, and the set is not visible at review time.
Show that the exploitable pair is created at deployment by whoever enables both capabilities, after every review has passed, and that this makes it a discovery problem rather than a bypass.
Be ready to argue who owns a risk that no capability author created, and why 'review each item harder' cannot address a property of the configured set.
### The claim, precisely An AI assistant becomes dangerous less because it holds any one bad tool and more because it holds a *combination*: one capability that reads content the operator treats as private (working files, another user's records, internal context) and another that moves bytes off the process (a fetch, an outbound request, a message send). Either alone is unremarkable. Together, in the same session, text that an attacker gets in front of the model can steer the private read toward the outbound write — the model carries the data out on the attacker's behalf. ### Why a per-capability review cannot see it A review that signs off tools one at a time asks, of each: *is this capability, on its own merits, acceptable?* A private-file reader in isolation is acceptable — reading files is its whole point. A web fetch in isolation is acceptable — fetching pages is its whole point. Neither judgement is wrong. The mismatch is one of *units*: the review's unit is the capability, and the vulnerability's unit is the **set**. Nothing in a per-capability process has the job of asking which capabilities coexist, because coexistence is not a property of any single capability — it is a property of the configuration that turns them both on. ### The set is assembled after review, by someone reviewing nothing This is the part candidates miss. The exploitable pair does not exist at review time. It comes into being later, when whoever stands up the deployment enables both capabilities for the same assistant. That act — placing a private reader and a network writer into one session — is configuration, not review. The person doing it is typically turning on features a customer asked for, not weighing an attack. So the vulnerable combination is *born* at the moment every review has already passed, in the hands of someone who is not looking for it. This is exactly why 'just review each capability more carefully' is the wrong answer: no amount of care on the individual item surfaces a risk that only exists once two approved items sit together. ### What follows for the attacker Because the pair is a property of the deployment, the attacker's task is not a *bypass* — there is no single gate that says no. It is a *discovery*: notice that a read reaching private context and a write reaching the network are live in the same session. That is solved from the outside by watching behaviour, since the attacker cannot see the configuration. The capabilities were each honestly approved, so nothing flags when they are used; the attacker is reading which tools respond, not tripping an alarm. ### Where the construction runs out The composition is the whole vulnerability, so it exists only while the composition does. If a given deployment never co-locates a private read and a network write in one session — separate assistants, separate sessions, no shared context — then there is no pair to find and the attacker's discovery has nothing to land on. The individual capabilities are then as safe as their reviews said. Note the direction of every claim here: a capability passing review proves only that it was judged acceptable alone; it says nothing about the set it will later join. And an assistant that answers a probe using private content proves the read reached private context in that session — not that the store was breached, and not, by itself, that a write is also present. Establishing the *pair* takes evidence about both halves, which is the subject of the harder questions on this topic. ### Why interviewers ask it Excessive agency and composed tool paths are the load-bearing idea behind agent security. An interviewer wants to hear that you locate the risk in the configured set and name the deployment/composition owner as the party who created it — not that you'd re-review the parts that were never individually wrong.
- If every capability author did their job correctly, who owns the resulting vulnerability?The party who assembled the deployment — the configuration or composition owner who enabled a private-context read and a network-reaching write for the same assistant. The risk is a property of that chosen set, not of any capability, so it cannot be assigned to a tool author who was reviewed and approved in isolation. Naming that owner is often the hardest part of filing the finding, because the pair belongs to whoever wired it, and wiring was not treated as a review step.
- Does adding a third capability change how you think about the risk?Yes — the exposure grows with the number of dangerous pairings, not linearly with the number of tools. Each new capability can form a read-plus-write pair with several existing ones, so the set of exploitable combinations expands combinatorially. This is another reason per-capability review scales badly: the review effort grows with tools while the risk grows with pairs, and no single-item review ever looks at a pair.
saying these in an interview costs you the question
- Says reviewing each capability closely enough would catch it
- Treats it as one tool being vulnerable rather than the pair
- Blames the reader tool's author for the combination
- Believes approving the parts approves the whole
- Calls it a control bypass rather than a discovery