skip to content

What must an AI agent record when a human approves a destructive action?

level: seniorimportance: should knowfreq 44%

answer

  1. approval of what, exactly?
  2. who, what, when, and against which plan
  3. bind execution to the literal arguments
  4. hash the approved payload
  5. identity from the IdP, not the model

basics

~20 s

Record who approved — from an authenticated identity, never from model output — what exactly they saw and accepted as literal tool arguments, when, against which plan version, and the decision itself. Then execute those stored arguments so the approval binds to a specific action.

solid answer

~50 s

An approval record has to answer "who agreed to what, when, and on what basis" well enough that someone reading it a year later can reconstruct the decision. That means: approver identity taken from the authenticated session, not a name the model wrote; the exact tool name and literal argument payload rendered to the human; a hash of the plan or payload so you can prove nothing changed between approval and execution; timestamp and expiry; the policy rule that fired; the decision plus any edits the approver made; and a correlation id linking the record to the run's trace. The binding matters more than the logging. If the agent re-prompts the model after approval, the executed call can differ from the approved one — a $500 refund becomes a $50,000 one. Execute the stored payload, re-check policy, and for high-value actions require two distinct approvers who are not the requester.

code

json · 19 lines
json
{
  "approval_id": "apr_01J9",
  "run_id": "run_7c2f",
  "policy_rule": "treasury.wire.over_50k",
  "policy_version": "2026-03-11",
  "proposed_call": {
    "tool": "wire_transfer",
    "args": {"beneficiary": "ACC-88421", "amount_usd": 75000, "currency": "USD"}
  },
  "payload_sha256": "3f9a...c17d",
  "decision": "approved_with_edit",
  "edited_args": {"beneficiary": "ACC-88421", "amount_usd": 50000, "currency": "USD"},
  "approvers": [
    {"subject": "u_4412", "idp": "okta", "role": "treasury_approver", "at": "2026-03-14T09:31:08Z"},
    {"subject": "u_9087", "idp": "okta", "role": "treasury_approver", "at": "2026-03-14T09:44:52Z"}
  ],
  "expires_at": "2026-03-14T10:01:08Z",
  "executed_payload_sha256": "a01b...66e2"
}

go deeper

for a junior

Know that approving an agent's action means recording who said yes and to exactly which action, and that the human should see the real values being used, not a summary.

for a middle

Explain the field list and why the executed call must come from the stored payload rather than a fresh model turn, since re-generation lets the action drift away from what was approved.

for a senior

Show the enforcement thinking: hash comparison at execution, identity from the IdP, entitlement checks on the approver, re-validated policy on resume, and tamper-evident append-only storage the agent cannot write to.

for a principal

Own the governance shape — separation of duties and four-eyes thresholds, retention and redaction for payloads carrying personal data, and making approval events queryable alongside traces so audits are a query rather than an investigation.

## The record is a control, not a log line Teams usually treat approval logging as compliance paperwork bolted on after the gate works. That gets it backwards. The record is what makes the approval *mean* something: it pins a specific human to a specific action at a specific moment, and it is the artifact you execute against. A row that says `approved: true, at: 14:32` proves nothing — approved what, by whom, and is the thing that ran the thing they saw? ## Fields that earn their place - **Approver identity, from the identity provider.** The authenticated session of the person who clicked. Never a name supplied by the model, parsed from free text, or inherited from the agent's service account. If the agent can influence who the log says approved, the log is fiction. - **The literal payload.** Tool name and the exact arguments, stored verbatim — not the model's prose summary. A screen reading "issue the customer refund" gives the human nothing to verify; a screen reading `issue_refund(order=A-8812, amount_usd=500.00)` does. - **A hash of what was approved.** Hash the argument payload, and for plan-level approvals the plan document too. Comparing the hash at execution time turns "we think it was the same call" into a check. - **Timestamp and expiry.** When the request was raised, when the decision landed, when the approval stops being valid. - **The policy that fired.** Which rule demanded the gate, and its version. Without it you cannot tell later whether a gap was a policy bug or an operator decision. - **Decision and edits.** Approve, reject, or approve-with-modification, plus the modified payload and any free-text rationale. Rationale is often the only place the *why* survives. - **Correlation.** Run id, trace id, conversation id, tenant. The record must join to the agent's trajectory, or investigating an incident means guessing. ## Binding: approval must attach to a payload, not a vibe The defect that turns a working gate into theatre is re-generation. The flow that fails looks like: model proposes a call, the human approves, the harness tells the model "the human approved, go ahead", and the model emits a *new* tool call. Nothing constrains that new call to match the old one. Sampling variance, a re-read of a poisoned observation, or an ambiguous instruction is enough for the amount, the account or the environment to shift. The human approved a $500 refund; a $50,000 one executes; the record says "approved". The fix is structural. The pending call is stored at gate time. On resume, the executor takes the stored payload, verifies the hash, re-evaluates policy against it, checks preconditions, and executes it directly. The model is not consulted about the approved action at all — it is only resumed with the *result*. If the approver edited the arguments, the edited payload is what gets stored, hashed and re-checked; edits are a first-class path, not a bypass. ## Four eyes, and who may approve For high-consequence actions, identity is not enough — you need separation of duties. The approver must not be the person who initiated the run, and above a threshold you want two distinct approvers. A treasury agent releasing a wire over $50,000 under a four-eyes rule needs two authenticated humans whose sessions are independent, and the record must show both. This is also where entitlement checks belong. Being logged in is not the same as being *allowed* to approve production DDL or a chargeback reversal. Check the approver's role against the action class, and record the entitlement that permitted it, so a later privilege change does not erase the reasoning. ## Making the record trustworthy A few properties separate an audit trail from a text file: append-only or otherwise tamper-evident storage; the agent's own credentials having no write access to it; retention aligned with the regulatory shape of the action, which for payments is usually years; and redaction discipline, because argument payloads carry account numbers and personal data. Emitting the record as structured events on the same tracing backbone as the agent's spans is what makes it queryable — you want to ask "show every gated production DDL last quarter, its approver, and whether it executed" and get an answer in one query. ## What interviewers listen for The distinguishing answer connects the record to enforcement rather than reciting fields. Say that you store the payload and execute *that*, that identity comes from the IdP, that the human is shown the real arguments and not a paraphrase, and that a rejected or expired approval leaves a record too — the absence of an execution is itself evidence. Candidates who describe only "we log approvals to a table" have built the paperwork without the control.

  • Why store a hash of the approved payload rather than just a boolean decision?
    Because the hash is checkable at execution time. A boolean tells you a human said yes; a hash lets the executor prove that the payload it is about to run is byte-identical to the one rendered on the approval screen. It also survives investigation: months later you can recompute the hash from the stored arguments and show that nothing was edited in between. For plan-level approvals, hashing the plan document catches silent rewrites of steps the approver never saw.
  • The approval UI shows the model's summary of the call instead of the arguments. What goes wrong?
    The human approves a description, not an action, and the model authored the description. A summary can omit the environment, round the amount, or simply be wrong, and prompt injection can shape it deliberately. Render the literal payload — tool name and resolved arguments — as the primary artifact, with a generated summary as an optional aid beside it. If a payload is too large to read, that is a signal the gate is at the wrong granularity, not a reason to paraphrase.
  • Who counts as the approver when the request is confirmed through an automated workflow tool?
    Whoever's authenticated session made the decision, propagated end to end. If a workflow or chat integration sits in the path, it must carry the user's verified identity through rather than acting under its own bot credential, or every approval collapses to "the integration approved it". Record the identity, the channel it came through, and the entitlement that permitted this action class — and if the path cannot prove a human identity, treat the approval as invalid.
  • Should a rejected approval be recorded as carefully as an accepted one?
    Yes, and it is often more useful. Rejections are the signal that the gate is doing work: they show which proposals a human refused, which is the raw material for tightening policy, fixing prompts, or removing a capability. Record the same fields plus the rejection reason, and keep them joined to the trajectory so you can see what the agent did next. A gate with zero recorded rejections over months is a gate nobody is really reading.

Launching a nuclear weapon needs two officers turning distinct keys, and the order they authenticate is the exact order that executes — not a paraphrase someone retypes afterwards.

saying these in an interview costs you the question

  • Logs only a timestamp and approved-yes with no payload
  • Lets the model regenerate the tool call after approval
  • Takes the approver's identity from the model's output
  • Shows a natural-language summary instead of the real arguments
  • Lets the requester approve their own high-value action

context