skip to content

Agent, Tool-Use & MCP Attacks

You will learn how attacks escalate once a model can call tools, browse, or drive other agents, including confused-deputy actions and poisoned MCP tool descriptions. Interviewers focus here because agentic systems convert a text trick into real side effects, which is where LLM security is heading.

on this pageshow

explore

questions

28

Why can an attacker exploit a read and a write that each passed an AI assistant's per-capability review?

level: juniorimportance: must knowfreq 68%

answer

  1. one tool fine, two together not
  2. the risk is in neither part
  3. review sees items, not the set
  4. the pair is born at deployment
  5. assembled by someone reviewing nothing

basics

~20 s

Because the danger lives in the pair, not either tool. A per-capability review judges each on its own and never sees the combination, and the combination is assembled later, when a deployment wires a private-context read and a network-reaching write into the same session.

solid answer

~50 s

A single read that reaches private context is fine; a single write that reaches the network is fine; the two in one session are what let attacker-controlled text pull private data toward an outbound call, with the model as the courier. A review that signs off tools one at a time approves each on its own merits and has nowhere to see the set — because the set does not exist at review time. It is created when whoever configures the deployment enables both capabilities for the same assistant, and that person is wiring features together, not adjudicating risk. So the exploitable combination is born after every review has passed, and no individual approval was wrong. The attacker does not defeat a control; they notice that the pair coexists. The composition is the vulnerability, and it belongs to the deployment, not to either tool's author.

go deeper

for a junior

Be ready to state that the risk is a property of the combination — a read that reaches private context plus a write that reaches the network in one session — not of either tool alone.

for a middle

Explain why a per-capability review is structurally blind: its unit is the single tool, but the vulnerability's unit is the set, and the set is not visible at review time.

for a senior

Show that the exploitable pair is created at deployment by whoever enables both capabilities, after every review has passed, and that this makes it a discovery problem rather than a bypass.

for a principal

Be ready to argue who owns a risk that no capability author created, and why 'review each item harder' cannot address a property of the configured set.

### The claim, precisely An AI assistant becomes dangerous less because it holds any one bad tool and more because it holds a *combination*: one capability that reads content the operator treats as private (working files, another user's records, internal context) and another that moves bytes off the process (a fetch, an outbound request, a message send). Either alone is unremarkable. Together, in the same session, text that an attacker gets in front of the model can steer the private read toward the outbound write — the model carries the data out on the attacker's behalf. ### Why a per-capability review cannot see it A review that signs off tools one at a time asks, of each: *is this capability, on its own merits, acceptable?* A private-file reader in isolation is acceptable — reading files is its whole point. A web fetch in isolation is acceptable — fetching pages is its whole point. Neither judgement is wrong. The mismatch is one of *units*: the review's unit is the capability, and the vulnerability's unit is the **set**. Nothing in a per-capability process has the job of asking which capabilities coexist, because coexistence is not a property of any single capability — it is a property of the configuration that turns them both on. ### The set is assembled after review, by someone reviewing nothing This is the part candidates miss. The exploitable pair does not exist at review time. It comes into being later, when whoever stands up the deployment enables both capabilities for the same assistant. That act — placing a private reader and a network writer into one session — is configuration, not review. The person doing it is typically turning on features a customer asked for, not weighing an attack. So the vulnerable combination is *born* at the moment every review has already passed, in the hands of someone who is not looking for it. This is exactly why 'just review each capability more carefully' is the wrong answer: no amount of care on the individual item surfaces a risk that only exists once two approved items sit together. ### What follows for the attacker Because the pair is a property of the deployment, the attacker's task is not a *bypass* — there is no single gate that says no. It is a *discovery*: notice that a read reaching private context and a write reaching the network are live in the same session. That is solved from the outside by watching behaviour, since the attacker cannot see the configuration. The capabilities were each honestly approved, so nothing flags when they are used; the attacker is reading which tools respond, not tripping an alarm. ### Where the construction runs out The composition is the whole vulnerability, so it exists only while the composition does. If a given deployment never co-locates a private read and a network write in one session — separate assistants, separate sessions, no shared context — then there is no pair to find and the attacker's discovery has nothing to land on. The individual capabilities are then as safe as their reviews said. Note the direction of every claim here: a capability passing review proves only that it was judged acceptable alone; it says nothing about the set it will later join. And an assistant that answers a probe using private content proves the read reached private context in that session — not that the store was breached, and not, by itself, that a write is also present. Establishing the *pair* takes evidence about both halves, which is the subject of the harder questions on this topic. ### Why interviewers ask it Excessive agency and composed tool paths are the load-bearing idea behind agent security. An interviewer wants to hear that you locate the risk in the configured set and name the deployment/composition owner as the party who created it — not that you'd re-review the parts that were never individually wrong.

  • If every capability author did their job correctly, who owns the resulting vulnerability?
    The party who assembled the deployment — the configuration or composition owner who enabled a private-context read and a network-reaching write for the same assistant. The risk is a property of that chosen set, not of any capability, so it cannot be assigned to a tool author who was reviewed and approved in isolation. Naming that owner is often the hardest part of filing the finding, because the pair belongs to whoever wired it, and wiring was not treated as a review step.
  • Does adding a third capability change how you think about the risk?
    Yes — the exposure grows with the number of dangerous pairings, not linearly with the number of tools. Each new capability can form a read-plus-write pair with several existing ones, so the set of exploitable combinations expands combinatorially. This is another reason per-capability review scales badly: the review effort grows with tools while the risk grows with pairs, and no single-item review ever looks at a pair.

saying these in an interview costs you the question

  • Says reviewing each capability closely enough would catch it
  • Treats it as one tool being vulnerable rather than the pair
  • Blames the reader tool's author for the combination
  • Believes approving the parts approves the whole
  • Calls it a control bypass rather than a discovery

context

open as a page

Why can someone uploading one file to a summarising service outspend its per-request token cap?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A per-request token cap bounds one model call, not one submission. An accepted upload fans out into extraction, a call per unit of content, an aggregation pass and retries, so dozens of individually compliant calls are billed to the operator.

open as a page

Why doesn't ending the session clear an injected fact an email assistant wrote to long-term memory?

level: juniorimportance: must knowfreq 72%

basics

~10 s

Session teardown discards the conversation context, not the durable store. Anything the assistant extracted and saved during that session outlives it by design, so a planted claim is read back into later, unrelated conversations.

open as a page

For an agent's tool calls, what is the difference between an attacker causing a new call and supplying an existing call's arguments?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Causing a new call adds an operation nobody asked for. Supplying arguments leaves the operation exactly as planned and changes only what it acts on. The second is much quieter, because the call itself still looks routine.

open as a page

Why does per-step verification against an agent's recorded plan confirm a goal-substitution attack?

level: middleimportance: must knowfreq 66%

basics

~20 s

Because the check compares each action against the recorded objective, and the recorded objective is what was edited. It tests consistency, not authenticity, so after the edit every action genuinely serves the goal on file and the verifier signs each one off.

open as a page

Why does an attacker with one obeyed instruction often skip an assistant's most privileged operation?

level: middleimportance: must knowfreq 70%

basics

~20 s

A target is worth its reach multiplied by the chance the call passes unremarked. The most privileged operation is the one its owners already watched. The productive target is the widest operation that still looks routine.

open as a page

Two components each passed review alone; one now launders text into the other. How do you scope the reachable effect?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Scope it as the union, not the intersection: whatever the reading component reaches is now reachable by content the emitting component read. Each review measured one component's blast radius; neither measured the pair, and internal provenance measured nothing at all.

open as a page

Why is an internal handoff's notes field still untrusted when our own component wrote it from attacker-supplied text?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Provenance is not integrity. The notes member is free text one component composed out of attacker-supplied input, so the attacker's wording crossed the hop intact. What changed at the boundary was the label on the text, not the text.

open as a page

In an LLM agent, how does editing the standing objective differ from getting one tool call obeyed?

level: juniorimportance: should knowfreq 58%

basics

~20 s

One obeyed call ends when that step ends. An edited objective persists: the agent re-reads it at the top of every later step and plans from it, so it keeps pursuing the substituted goal on its own initiative with no further injected text.

open as a page

Why does scoping a tool-using assistant's granted operations to its job still leave an attacker a choice of target?

level: juniorimportance: should knowfreq 62%

basics

~20 s

Scoping removes the operations the job never needs, not the ones it does. A real job spans many capabilities, so what survives is still a menu, and an attacker holding one obeyed instruction picks a target from that menu.

open as a page

From an assistant's replies alone, how would an attacker infer which of its tools reads private data and which reaches the network?

level: middleimportance: should knowfreq 50%

basics

~20 s

By differential probing: steer the assistant toward tasks that would each need a private-context read or a network write, then read the replies for tells — content that could only come from a private source proves the read; an observable outbound effect proves the write. Behaviour reveals the set the attacker cannot see.

open as a page

In a chain of components, what property makes one hop the right place for an attacker to launder text through?

level: middleimportance: should knowfreq 44%

basics

~20 s

The hop worth laundering through is the one whose output the next component reads as free text rather than a typed record, sitting behind the single arrival check, and read by a component with capabilities the emitter lacks.

open as a page

In a file-summarising pipeline, why does an attacker pick structure over wording?

level: middleimportance: should knowfreq 45%

basics

~20 s

Structure decides the unit count, wording does not. Part, sheet and embedded-object counts and nesting depth set how many calls one submission becomes, and a structural file carries no directive span for a content screen to score.

open as a page

What does a stored memory entry's provenance prove when an extractor wrote it during the user's own session?

level: middleimportance: should knowfreq 54%

basics

~20 s

It proves which component wrote the entry and in which session, and nothing else. Both stamps are truthful for a claim lifted from a stranger's email, so provenance cannot separate attacker-seeded facts from ones the user actually stated.

open as a page

In an invoice-processing agent, how does a supplier-controlled document field reach a tool call's argument slot?

level: middleimportance: should knowfreq 55%

basics

~20 s

By being read as data the whole way. An extraction step lifts the field from the ingested document, a mapping step places it in the argument the operation takes, and no stage asks who wrote it.

open as a page

When refused, absent, and empty-but-successful tool calls all surface as the same bland sentence, how do you build a reliable read/write oracle?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Stop reading the sentence and design probes where the three cases diverge somewhere observable — an out-of-band signal, timing, or a later-visible effect. A bland reply proves the answer was withheld, not that nothing ran, so you separate the events by side effects, not by wording.

open as a page

A public upload reached no private data yet drained a shared model quota - is that a finding?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Yes, when the asymmetry holds: an unauthenticated submission costing the operator far more than the submitter, repeatable, and drawing down capacity other tenants share. No data reached proves the read path was not exercised, not that nothing was spent.

open as a page

In a long-running LLM agent, what must an obeyed span reach to change the standing objective rather than one step?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It must reach the record the loop re-reads, through a step that writes into it — usually one consolidating the standing goal. The agent performs that write, so the span must read as task detail, not an instruction.

open as a page

What does an attacker give up by aiming at an assistant's end-of-turn memory extractor instead of the live answer?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Control and feedback. Whether the claim is kept, how it is reworded and when it is read back are decided by components the sender never sees, with nothing echoed back. What is bought is persistence with no expiry.

open as a page

An obeyed instruction reaches a coding assistant inside failed CI output. How does that change target choice?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It moves the bar. Whether a call looks routine depends on the task the assistant appears to be doing, and an instruction arriving in a tool's return value lands mid-run, when repository writes already fit the story.

open as a page

An agent's call log shows an approved operation with valid arguments — how do you triage a report that it should not have run?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Stop asking whether the call was allowed and ask who chose each value. The log settles which operation ran and that it validated; it carries no provenance, so compare each argument against the record meant to determine it.

open as a page

You confirmed a read-plus-write path in an assistant where every capability passed review — whom do you file it against and how do you defend its severity?

level: principalimportance: should knowfreq 35%

basics

~20 s

File it against the party who assembled the deployment — the configuration owner who put a private-context read and a network write in one session — not any tool author. Defend severity as the composed capability: private data reaching an outbound call, established with a stated probe count and reproduction rate.

open as a page

Why does tool-call argument validation that checks type, length and format pass an attacker-written value?

level: middleimportance: nice to knowfreq 34%

basics

~10 s

Because it judges the string, and the objection is about the string's origin. Type, length and format are properties a value has on its own, so a correctly formed identifier passes whoever wrote it.

open as a page

The audit trail names the upstream component as initiator of a downstream action. What does that record prove?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

It proves which call ran, under which authority, at whose request. It does not record who chose the argument values or where the prose that motivated them came from, so a trusted identity is credited with the action.

open as a page

Does upload fan-out against a summarising service stay worth red-teaming as inference gets cheaper?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Exposure is unit price times a multiplier the ingest design chooses, and only the first term falls. Cheaper models and caching shrink the price; neither touches the multiplier, and a bound per submission relocates it to submission count.

open as a page

A goal-substitution finding reproduces once in five runs and the step trace is green — how do you adjudicate it?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Adjudicate the structural property, not the outcome. That the objective is run-writable and the step check reads it reproduces every time, while the end-to-end result is one in five. Report the rate with its trials, and file it against the architecture owner.

open as a page

An assistant's memory store holds thousands of extracted claims with no source text - what do you tell the owner about trusting it?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

No claim in an auto-extracted memory store can be attributed, so the choice is what to accept, not what to verify: keep it and carry claims of unknown authorship, discard it and lose personalisation, or re-derive what surviving data supports.

open as a page

A red-team run made only granted, ordinary-looking calls. Do you file that as a bug or a design limit?

level: principalimportance: nice to knowfreq 33%

basics

~10 s

Neither label fits until you name the claim the run falsified. Nothing malfunctioned, so there is no component to file against; what broke is the belief that a roster granted once bounds a run.

open as a page