skip to content

What must a human approval prompt show before an agent's irreversible action?

level: seniorimportance: must knowfreq 58%

answer

  1. the summary is attacker-influenced
  2. render the call, not the prose
  3. show resolved arguments and target
  4. provenance: what triggered this
  5. fatigue turns gates into clicks

basics

~20 s

Show the resolved call the system will actually make — tool, target and every argument verbatim — plus where the data that triggered it came from. A model-authored summary of the action is attacker-influenceable text, so approving on the summary approves nothing.

solid answer

~50 s

Human approval is the control that still holds after the model has been persuaded, but only if what the human sees is the action itself rather than the model's description of it. Under injection the prose is precisely what the attacker controls: an agent can narrate "sending the shortlist to the hiring manager" while the resolved arguments name a different recipient. So the approval surface must render the concrete call — tool name, target identifier, full argument values, diffs for edits — from the serialized request the harness will execute, not from generated text. It should also show provenance: which document or tool result led here, so a reviewer can spot that an action originated in an uploaded file rather than the user's request. Finally, gate only side-effecting and irreversible operations. Prompting on every read produces approval fatigue, and a fatigued reviewer clicks through the one that mattered.

go deeper

for a junior

Know that actions with real-world effects — sending, deleting, paying, changing status — should stop for a person, and that the person needs to see what will actually happen, not a summary.

for a middle

Explain that the approval surface must be rendered by the surrounding code from the resolved call, because model-generated prose is influenceable by whatever the agent just read. Name what belongs on the screen: tool, target, arguments, reversibility.

for a senior

Demonstrate the tradeoff you have lived: gate narrowly enough that reviewers still read, add provenance so an action originating in an uploaded document is visible, and instrument approval rate against review time to detect fatigue.

for a principal

Own the policy question of which capabilities may exist behind an approval at all versus which must not be wired to an agent, and be explicit that approval bounds side effects but not information flow.

## Why approval is the layer that survives Every prompt-level mitigation shares one dependency: the model has to cooperate. A human approval gate does not. It is enforced in code before the side effect fires, so it holds whether or not the model was persuaded by something it read. That makes it the most valuable single control in this taxonomy — and it makes the integrity of the approval *surface* the thing worth interrogating, because a gate whose display is attacker-influenced fails silently. ## The failure that makes this a senior question Picture a recruiting screener that reads candidate PDFs and can advance a candidate to interview. A résumé carries an invisible line instructing the reader to advance this candidate. The agent complies and asks for confirmation, and the confirmation reads: "Advance the strongest candidate to the interview stage?" The recruiter approves. The gate was present, the human was in the loop, and nothing was actually reviewed — because the sentence the recruiter read was generated by the same model the attacker had already influenced. The fix is structural. The approval UI must be rendered by the harness from the serialized call it is about to execute, never from model prose. Concretely: - **The tool and target.** Which capability, and which specific record, file, recipient or endpoint. Identifiers resolved to human-readable names, but the identifier shown too. - **Every argument, verbatim.** Not a paraphrase, not a truncation that hides a suffix. For edits and deletions, a diff or the exact rows affected. - **Provenance.** Which document, retrieved chunk or tool result the agent was reading when it proposed this. An action traceable to an uploaded PDF rather than to the user's own request is the single most useful signal a reviewer gets. - **Blast radius.** Whether the action is reversible, and how many entities it touches. "Advance 1 candidate" and "advance 340 candidates" must not look alike. - **Nothing rendered as markup.** The approval surface is itself an output sink; argument values must be escaped so a crafted value cannot restyle or hide part of the dialog. ## Where the gate is placed At the side-effect boundary, in the harness, not in the model's judgment. Two placements are wrong in ways that recur. Asking the model to decide whether an action is risky and to request approval accordingly puts the gate back inside the compromised component. The classification must be a property of the tool, declared in code. Gating everything — reads included — is the other failure. Approval fatigue is not a UX quibble; it is the mechanism by which the control degrades to a click-through. Gate writes, external sends, deletions, spends and state transitions; let reads run. Grouping a batch into one approval is fine only if the group is shown item by item. ## What approval does not solve It does not stop an agent from reading something it should not, and it does not stop information leaving through channels that are not gated. A model that has been persuaded can encode data into an argument of an action a human will plausibly approve. So the gate is one layer: it constrains side effects, while keeping untrusted free text away from the deciding component and restricting what the agent can invoke at all constrain the rest. It also does not survive a poor human factor. If the reviewer cannot tell what "good" looks like — because the argument is an opaque identifier, or the diff is 200 lines — the gate is decorative. Design the surface so the decision is answerable in a few seconds. ## Operating it Record the approval: who approved, the exact serialized call, the provenance, and the time. That record is what makes an incident reconstructable and what tells you whether reviewers are approving in one second or ten. Rising approval rates with falling review time is the measurable signature of fatigue, and it is a reason to narrow what gets gated rather than to add more gates. ## Interview framing Lead with the sentence that shows you understand the threat: under injection the model's description of the action is attacker-controlled, so the approval must render the resolved call from the harness. Then add provenance, blast radius and the fatigue tradeoff. Candidates who answer only "add a human in the loop" have named the layer without knowing what makes it work.

  • How do you keep approval gates from degrading into reflexive clicking?
    Gate narrowly — writes, external sends, deletions, spends, state transitions — and let reads through. Make each decision answerable in seconds: resolved names, a diff rather than a blob, blast radius stated. Then instrument it: track approval rate against time-to-decide. Near-100% approval at one second per item means the control has degraded, and the response is to narrow the gate, not to add another.
  • An agent proposes 340 record updates in one step. Do you ask for one approval or 340?
    One approval, but the surface must show the scope honestly: the count, the affected entities enumerated or sampled, and the diff shape. The danger of batching is that a single innocuous-looking line stands in for the whole set. If the batch mixes action types or targets, split by type so each approval covers one homogeneous, comprehensible group.
  • Could an attacker exfiltrate data through an action a human approves?
    Yes — that is the limit of this control. Data can be encoded into arguments of an operation the reviewer finds plausible, such as free text in a message body. Approval bounds side effects, not information flow, which is why it is layered with keeping untrusted text out of the privileged component and constraining what the agent can invoke at all.

saying these in an interview costs you the question

  • Approving on the model's natural-language description of the action
  • Letting the model decide which actions need approval
  • Gating every step including read-only lookups
  • Showing truncated or paraphrased arguments in the dialog
  • Treating a human in the loop as sufficient on its own

context