skip to content

Why does scoping a tool-using assistant's granted operations to its job still leave an attacker a choice of target?

level: juniorimportance: should knowfreq 62%

answer

  1. the job needs more than one operation
  2. scoping trims the menu, not the choice
  3. granted once at onboarding, never re-read
  4. scoped to the job, not to this run
  5. a roster read as a target list

basics

~20 s

Scoping removes the operations the job never needs, not the ones it does. A real job spans many capabilities, so what survives is still a menu, and an attacker holding one obeyed instruction picks a target from that menu.

solid answer

~50 s

An obeyed instruction on its own changes nothing outside the transcript; it becomes an effect only when it is aimed at one of the operations the assistant is actually allowed to call. Least-privilege scoping shrinks that roster by removing what the assistant's job never needs — but a job like "help maintain a repository" legitimately needs to open branches, edit files, move labels, rerun builds and cut releases, so the surviving set is a dozen or more operations rather than one. Two further facts widen it in practice: the roster is usually granted in a single onboarding decision and never re-read, and it is scoped to the job as a whole rather than to the run in progress. So the attacker's question is never "can I get a new permission" — it is "which of the ones already granted do I aim at". That selection is the whole move.

go deeper

for a junior

Be ready to say plainly that an obeyed instruction only becomes an effect through an operation the assistant was already allowed to call, and that a scoped roster is usually still a dozen operations rather than one.

for a middle

Explain why scoping cannot cut the set to one: permissions are granted per job, not per run, and a useful job spans many operations. Note that a grant records necessity, not safety.

for a senior

Show you have looked at a real roster. Expect to describe how grants accumulate in a single onboarding decision, are never re-read, and stay live during runs that have nothing to do with why they were requested.

for a principal

Own the framing that the roster is an inventory somebody must periodically re-read and re-justify, and that whoever granted it, not whoever wrote the assistant, is the person whose decision determines what an obeyed instruction can ever be worth.

## What the question is really about A tool-using assistant is a language model wired to a set of callable operations: read a file, open a branch, edit a workflow definition, rerun a build, move a label, cut a release. Text the model produces can request one of these, and an orchestrator executes it. That wiring is the only path from words to consequences. Everything else the model emits is text in a transcript. An attacker who has succeeded in getting a span of untrusted content treated as an instruction — arriving in retrieved content, in a document, or in the text a tool handed back — now holds exactly one thing: an instruction the assistant is willing to follow. That is worth nothing by itself. It has to be pointed at something, and the only somethings that exist are the operations already on the roster. ## What least-privilege scoping does and does not remove Scoping means granting the assistant only the operations its job requires, rather than everything the underlying credential could reach. It is real and it matters: it is the difference between an assistant that can touch one repository and one that can touch an organisation's whole estate. What it does not do is reduce the roster to a single operation, because jobs are not single-operation. A coding assistant that is supposed to be useful on a failing build has to be able to read the build's output, read and edit files, open a branch, push a change and re-run the job. An assistant that is supposed to help with releases needs the operations that make releases. Each grant is individually justified. The set is still plural. Three practical facts make the surviving set wider than anyone remembers: - **It is granted once.** Rosters are typically assembled during onboarding, when someone is trying to make the assistant work, and then not re-read. Forty registered operations is an ordinary number and nobody has an accurate mental list of them. - **It is scoped to the job, not the run.** The permission that makes the assistant useful during a release is present during every other run too. Scoping has no notion of "needed right now". - **A grant is a statement about intent, not about consequence.** That an operation appears on the roster proves someone judged it necessary for the assistant's work. It does not prove the operation is small, reversible, or unlikely to matter when it fires for a reason nobody intended. ## The direction of the claim Be careful about what each observation supports. That a call was inside the grant proves the orchestrator was permitted to make it — it says nothing about who chose to make it or why. That an operation survived scoping proves it is needed by the job, not that it is harmless. Candidates routinely collapse those two, and it is the collapse that makes this whole class of finding surprising to the people who own the roster. ## Why this is the first thing to understand about agentic abuse Every later question in this area assumes it. Which capability the instruction is aimed at, what the arguments were, whether two operations compose into a path — all of that presupposes a plural roster to choose from. If scoping really did reduce the set to one, there would be no selection problem and this material would not exist. It does not, so the selection problem is the material. The practical framing to carry into an interview: an assistant's granted operation list is, from the attacker's side, a target list. It was written by people asking "what does this assistant need". It is read by someone asking "which of these gets me the effect I want, and which one will nobody look at twice". Those two readings of the same list produce different orderings, and that divergence is where the interesting answers live.

  • If the instruction is obeyed but no operation on the roster reaches the effect the attacker wants, what has actually happened?
    Nothing outside the transcript. The model can produce text describing the effect, and that text may still matter if something downstream reads it, but the assistant has no path to the effect itself. This is why the roster, not the instruction, sets the ceiling on what an obeyed instruction is worth in a given deployment.
  • Does an operation appearing in the assistant's action log prove someone intended that call?
    No. The log proves which operation ran with which arguments under the assistant's credential. It carries no evidence about what caused the model to request it — an operator's request and an obeyed span in retrieved text produce identical entries. Provenance of the intent is exactly what these records do not hold.

A building pass scoped to one department still opens every door that department uses. Narrowing it to the job is not the same as narrowing it to the errand someone is on today.

saying these in an interview costs you the question

  • Claims least-privilege scoping leaves the attacker no target choice
  • Treats an operation as harmless because it was deliberately granted
  • Assumes the attacker must obtain a new permission rather than use an existing one
  • Thinks an obeyed instruction produces effects without a callable operation
  • Confuses scoping to a job with scoping to the current run

context