skip to content

Why does an attacker with one obeyed instruction often skip an assistant's most privileged operation?

level: middleimportance: must knowfreq 70%

answer

  1. reach times unremarkability, not power
  2. the powerful one is the watched one
  3. the instruction is spent once
  4. widest thing that still looks boring
  5. consequence ordering versus ordinariness ordering

basics

~20 s

A target is worth its reach multiplied by the chance the call passes unremarked. The most privileged operation is the one its owners already watched. The productive target is the widest operation that still looks routine.

solid answer

~50 s

An obeyed instruction is spent once, so the attacker is choosing where to spend it, and the ranking is not by power. Two properties matter together: how far the operation's effect reaches, and how unremarkable the call looks to whatever reads the run afterwards. The most privileged operation scores badly on the second, precisely because it is the one somebody thought about — it is the one behind an approval, the one a reviewer's eye stops on in the end-of-run summary. A less impressive operation that mutates shared state the whole team pulls from can score better on both axes at once: it reaches everyone, and it reads on the summary line like the dozen other calls the run made. This is why the wrong answer in an interview — "aim at the most powerful thing it can do" — is wrong: a privileged call that triggers a look is worth less than an unwatched one that quietly reaches the same shared state.

go deeper

for a junior

Know that an assistant only causes effects through operations it was granted, and that the flashiest permission is not automatically the one an attacker cares about.

for a middle

Explain the two axes that decide a target's value: how far one call's effect reaches, and how likely the call is to pass without anyone stopping on it. Say why they pull against each other.

for a senior

Show judgment about real rosters: which operations reach state a whole team consumes, which sit behind an approval, and why an operation's name is a poor predictor of its reach.

for a principal

Own the observation that attention is allocated by an ordering somebody chose, and be ready to argue what follows when that ordering and the ordering by consequence disagree across forty granted operations.

## The wrong instinct, stated plainly Asked which capability an attacker aims an obeyed instruction at, most competent engineers answer "the most privileged one". It is the natural transfer from classic access control, where the goal of an escalation chain is the highest privilege you can reach. In an agentic deployment it is the wrong ordering, and understanding why is most of this topic. ## Two axes, not one The attacker is choosing a target on the assistant's granted roster. The value of each candidate is roughly the product of two things: - **Reach** — how far the effect of one call travels. A change to something one person reads has small reach. A mutation to state a whole team pulls from, or that a pipeline consumes on every run, has large reach. Irreversibility multiplies it further: an effect that can be undone in a minute costs the defender less than one that cannot be recalled once others have consumed it. - **Remarkability** — how much attention the call attracts on its way through, and afterwards. This covers the approval click in front of certain operations, and it covers what happens after the run: the action summary somebody skims, the triage that classifies each listed call as ordinary or not. Ranking only by reach produces the wrong answer because the two axes are anti-correlated by construction. Operations get approvals and get watched *because* someone judged their consequences large. The privileged operation is expensive to reach precisely in proportion to how obviously valuable it is. ## Why an obeyed instruction is a budget This matters because the instruction is a one-shot resource in the run where it lands. If the chosen call stops at an approval, the operator sees a request they did not ask for, and the whole construction has announced itself while achieving nothing. The attacker has traded a working path into the context for a notification. That is a bad trade, and it is why the ranking weighs "does this go through unremarked" as heavily as "how much does this do". ## What "unremarkable" actually means It is not a property of an operation in the abstract. It is a property of how the call reads to a human doing a fast triage of a list. A reviewer scanning an end-of-run summary is asking, per line, "is this the sort of thing this assistant does". A build rerun and a tag cut occupy one line each and read alike at that speed. The reviewer's ordering is by *ordinariness*; the risk register's ordering, when there is one, is by *consequence*. Those two orderings disagree, and the productive target sits exactly in the gap: high on consequence, low on apparent ordinariness. ## The shape of the answer an interviewer wants Something close to: *the useful target is the widest-reaching operation that is still boring.* Then the reasoning: reach and irreversibility set the payoff; the review path sets the price; the most privileged operation is the one with the worst price-to-payoff ratio because attention was allocated to it deliberately. Candidates who can also name what makes the ordering diverge — that a fast human triage sorts by how usual a call looks, not by what it costs to undo — are demonstrating that they have watched a real review happen rather than reasoning from first principles. ## Care with the claims A few directions that are easy to get backwards: - That a call completed without an approval prompt proves it was not behind one, not that it was harmless. - That a reviewer marked a line ordinary proves it resembled the assistant's usual work, not that its effect was small. - That an operation is unimpressive by name proves nothing about its reach. Names are chosen by the people who built the integration, and the correspondence between an operation's name and the size of its effect is weak. That last one cuts both ways, and it is the reason this selection is a skill rather than a lookup. The attacker's judgment is about the roster as deployed here, not about a general ranking of operation types.

  • If reach and unremarkability conflict on a roster, which does the attacker usually favour?
    Unremarkability, up to a point. A slightly narrower effect that lands is worth more than a wider one that gets stopped, because being stopped also burns the path into the context. The attacker only reaches for the loud operation when nothing quiet on the roster touches shared state at all, which is a fact about how the roster was scoped.
  • Does an operation completing without an approval prompt tell you it was low-consequence?
    No. It tells you the operation was not among those someone chose to put a prompt in front of. That set was drawn up by people predicting which calls would be regretted, and predictions made during onboarding do not track consequence evenly across forty operations. The gap between the two is the attacker's working space.
  • Why is the effect being shared and irreversible more valuable than it being large?
    Because a shared effect is consumed by people and pipelines that never see the run that produced it, and an irreversible one keeps its value after the run is noticed. Size measured in the moment fades; a mutation others have already pulled from does not, which is what makes the payoff outlive the discovery of the construction.

A courier stealing from a warehouse does not head for the safe. The safe is where the camera points. The pallet that ships out on every truck moves more value and never looks unusual on the manifest.

saying these in an interview costs you the question

  • Ranks targets by privilege alone and stops there
  • Says the attacker aims at whatever the assistant is most trusted to do
  • Ignores that an approval click makes a call announce itself
  • Assumes an operation's name predicts the size of its effect
  • Treats a call that drew no prompt as evidence it was harmless

context