skip to content

For an agent's tool calls, what is the difference between an attacker causing a new call and supplying an existing call's arguments?

level: juniorimportance: must knowfreq 72%

answer

  1. same operation, different target
  2. nothing new appears in the plan
  3. the call was always going to happen
  4. only the field values changed

basics

~20 s

Causing a new call adds an operation nobody asked for. Supplying arguments leaves the operation exactly as planned and changes only what it acts on. The second is much quieter, because the call itself still looks routine.

solid answer

~40 s

Two different moves sit behind the same phrase 'the agent did something bad'. In the first, the attacker gets an operation to happen that was never part of the workflow — a send, a delete, a fetch that had no business running. In the second, the workflow runs exactly as designed: an invoice-processing agent files the disbursement it was always going to file today, for a real amount — but the beneficiary field carries a value the attacker wrote into the ingested document, not one the requester chose. Same operation, same schedule, same actor, same shape. The first move must survive somebody noticing an operation that should not be there; the second only has to survive checks that ask *which* operation ran and whether its fields are well formed.

go deeper

for a junior

Be ready to state the two moves in one sentence each and say which one changes the plan. Naming the second move at all already puts you ahead of candidates who only know 'the agent did something bad'.

for a middle

Explain why the second move survives the checks the first one trips: operation kind, rate and sequence all stay normal. Be able to point at which layer each move is visible to.

for a senior

Show you can reason about cost. Say what each move requires of the attacker, why authoring a document field is cheaper than changing an agent's task, and where the argument move runs out of room.

for a principal

Frame it as which expectation the workflow encoded: if an outside party's document determines what a payment lands on, that is a design position somebody took, whether or not anyone wrote it down.

## Two moves that look the same in a summary and are not the same at all An agent that can act does so through tool calls: a named operation plus a set of argument values. Anything an attacker does to that machinery lands in one of two places — the *choice of operation*, or the *values that operation runs on*. Interview answers collapse these together constantly, and the distinction is the whole subject of this leaf. ### Move one: cause a call that was not going to happen Here the attacker gets the agent to perform an operation outside the plan. A queue that only ever files disbursement requests suddenly issues a credential read, or a message send, or a bulk export. The effect is new. It is also, relatively speaking, loud: the operation is one nobody expected on this workflow, so anything that looks at the *shape of the day* — which operations this agent performs, how often, in what order — has something anomalous to look at. Rate, sequence and operation-kind are all cheap to watch, and an unexpected kind stands out even in a log nobody reads carefully. ### Move two: supply the values on a call that was always going to run Here nothing about the plan changes. Consider an autonomous procurement and expense workflow: invoices arrive in a queue, an agent matches each one to a purchase order and files a disbursement request. Today's run was always going to file a disbursement for this invoice. The attacker's contribution is not the call — it is one field inside it. A remittance line, a supplier reference, an account string on the document itself carries the value the attacker chose, and the pipeline forwards that value into the argument slot the operation takes. The resulting record is an approved operation, of an approved kind, on a real invoice, for an amount inside policy, with every field correctly formatted. The only thing wrong with it is *whose choice* one value represents — and choice is not a property any of the usual checks can see. ### Why the second move is quieter Three layers commonly stand between an agent and the world, and none of them is built to notice this: | Layer | What it examines | What it misses here | | --- | --- | --- | | Argument validation | Type, length, enum membership, format | A well-formed value is well formed regardless of who wrote it | | The call log | Which operation ran, when, with which values | Nothing in the record states where any value came from | | A reviewer's expectation | Does this agent do things like this? | It does — this is exactly the thing it does | Each of those is a real control and each is doing its job. They are simply aimed at the operation, and the attacker never touched the operation. ### What each move costs the attacker Causing a new call generally requires getting the agent to accept a change of task — the harder thing, and the thing most likely to be caught by whatever notices an operation out of character. Supplying an argument requires only the ability to author a field on something the pipeline ingests, plus a working guess about which document field lands in which argument slot. In a workflow that accepts documents from outside parties by design, that first requirement is not an obstacle at all; it is the product feature. ### Where the second move stops working It stops where the argument is not a free choice. If the operation derives that field from internal state — the purchase order's own record of the payee, a value already fixed by the request that opened the workflow — there is no slot for the attacker to fill, and nothing the document says can change it. The class exists exactly to the extent that a value the outside world authored is allowed to determine what the operation lands on. ### What this is not It is not an escape from the argument into some other syntax: no quoting boundary is crossed, no command or query is smuggled, and the value never stops being a value. It is also not the model inventing an identifier or drifting out of an enumeration — that is an accuracy problem with a different owner. Here the value is in-specification by construction; that is the point of it. And the direction of the claim matters: a well-formed argument reaching a call proves the pipeline forwarded it, not that the model was hijacked.

  • Does the agent have to be talked into a different task for the second move to work?
    No, and that is what makes it cheap. The task stays the one the operator wanted: match this invoice, file this disbursement. The attacker contributes a value the workflow was going to read anyway. Anything watching for a change of plan sees no change of plan, because there is not one.
  • A teammate says the call was legitimate, so there is nothing to investigate. What does 'legitimate' establish?
    It establishes that the operation was in policy and in character for this agent. It says nothing about who chose the values it ran on. Those are separate claims, and the second one is the one under dispute. Treating in-policy as equivalent to intended is how this class stays open.
  • Which of the two moves would you expect an operations team to notice first?
    The new call, almost always. An operation of an unexpected kind breaks the daily shape of the workflow, and kind, rate and sequence are the things people actually watch. An expected operation with correctly formatted fields breaks nothing anyone is looking at.

One attacker adds a line to the cheque run. The other leaves the cheque run untouched and changes the payee on a cheque that was already being printed.

saying these in an interview costs you the question

  • Treats any tool-call abuse as one undifferentiated category
  • Says the agent must be hijacked into a new plan for harm to occur
  • Assumes an in-policy operation implies an intended target
  • Calls this command injection because untrusted text reached a call
  • Believes a complete call log makes the difference visible

context