skip to content

How do you set approval timeouts for an AI agent without creating rubber stamps?

level: principalimportance: should knowfreq 38%

answer

  1. pending requests cannot wait forever
  2. decide what expiry means
  3. default-deny on irreversible actions
  4. watch approval rate and dwell time
  5. a near-100% yes rate is a smell

basics

~20 s

Give every pending approval an expiry sized to the action's real urgency, and default to deny on expiry for anything irreversible. Route to a rota rather than one person, and gate few enough actions — with good enough evidence — that approvers still read them.

solid answer

~50 s

Two failure modes sit on opposite ends. Requests that never expire silently stall work and leave decisions applied against a stale world. Requests that auto-approve on timeout convert a safety control into a delay. My default is deny-on-expiry for irreversible actions, with the run recorded as blocked so someone owns the retry; auto-continue on expiry is defensible only for reversible, low-blast-radius steps where blocking costs more than the action. The harder problem is human. Approval fatigue is a genuine security failure — the same dynamic as MFA push fatigue — and it arrives whenever volume rises or the evidence shown is too thin to judge. Watch the approval rate and time-to-decision: near-100% yes with a few seconds of median dwell time means the gate is decorative. Fix it by narrowing what is gated, batching similar requests, improving the evidence rendered, and routing to an on-call rota with escalation instead of one named owner.

go deeper

for a junior

Know that an approval request needs a deadline, and that when nobody answers in time the safe outcome for a dangerous action is to not do it.

for a middle

Explain expiry semantics per action class, escalation to a secondary approver before expiry, and why silent auto-approval on timeout defeats the point of the gate.

for a senior

Show that you instrument the gate itself — approval rate, dwell time, rejection rate, queue depth — and use those numbers to re-scope policy, improve the evidence shown, and staff a rota with break-glass and escalation.

for a principal

Own the capacity tradeoff across a fleet: how much genuine human adjudication exists, whether to buy safety with capability controls instead of gates, and when pre-authorized envelopes beat per-call approval for keeping throughput without giving up control.

## Timeouts are a policy decision, not a config default Every pending approval is a run holding open state and, usually, a user waiting. The window you give a human should come from the action, not from a framework default. A checkout-flow correction that is worthless after the customer leaves might get 90 seconds. A production index change that will be applied in tonight's maintenance window can wait eight hours. A treasury wire under a four-eyes rule might get 30 minutes during trading hours and roll to the next business day outside them. Sizing the window is where domain knowledge shows up. ## Deny by default when it matters The question interviewers actually want answered is what happens at expiry. Three options, and the choice is per action class: - **Deny and close.** The correct default for irreversible or externally visible actions. Nobody answered, so nobody assented; a fresh request must be raised, with preconditions revalidated, because the world moved. Record the expiry — an approval that timed out is a signal about staffing or scope, and you want it in the data. - **Continue without the action.** Sometimes the agent can complete a degraded but useful result: finish the analysis, skip the send, hand back a draft. Better than a dead end when the gated step was optional. - **Proceed on expiry.** Only for reversible, contained steps where blocking is the larger harm — and even then it needs an explicit policy entry and a louder audit trail, never an inherited default. Silent auto-approval on timeout is how gates become theatre. Escalation belongs alongside expiry: nudge the primary approver, escalate to a secondary after a fraction of the window, then expire. And a break-glass path — an authorized human forcing an action through with mandatory justification and heightened logging — is worth having, because the alternative is people disabling the gate wholesale during an incident. ## Fatigue is the real failure mode A gate is only a control while a human is genuinely deciding. Push enough requests at an approver and they will start clicking approve, exactly as people accept MFA prompts they did not initiate. The mechanism is well understood: high volume, low base rate of genuine problems, and evidence too thin to evaluate, so approving becomes the rational default. The measurable symptoms are the approval rate and the decision dwell time. If a gate is approved 99.8% of the time with a median decision of four seconds, one of two things is true: the policy is scoped too broadly and is catching routine traffic, or nobody is reading. Both are fixable, and both are invisible unless you instrument them. Rejection rate near zero over months is the same signal. ## Designing for attention rather than volume Attention is the scarce resource, so spend it deliberately: - **Narrow the policy.** Move the threshold until the gate fires on the tail that genuinely needs judgement. Under a tiering scheme where refunds at or under $200 run automatically, only the exceptional cases reach a person. - **Improve the evidence.** Show the resolved arguments, the diff or blast-radius estimate ("this DDL touches a table serving 4,100 requests per minute"), the agent's stated reason, and what it already checked. A reviewer who can judge in ten seconds *because the evidence is good* is not fatigued; one who clicks in ten seconds because there is nothing to read is. - **Batch and group.** Fifty similar low-risk requests reviewed as one set with per-item opt-out beats fifty notifications. - **Make rejection cheap and useful.** If saying no means the run dies and the requester starts over, people say yes. Let a rejection carry feedback or an edited payload so the loop can continue sensibly. - **Rotate the load.** An on-call rota with defined hours, coverage across timezones, and explicit handoff. Routing to one named owner guarantees that their week off becomes a queue of expiries. ## The throughput conversation At fleet scale this becomes a capacity question that a lead owns: how many gated actions per day can the organization actually adjudicate, and what is the latency budget for a human decision inside a product flow? If the answer is that approvals would need three full-time reviewers, the design is wrong — either narrow the gates, or invest in capability controls (scoped credentials, sandboxes, per-tenant limits) so fewer actions need a human at all, or accept a slower product. Pre-authorized envelopes are the middle path: a human approves a bounded budget or scope once — this quarter's refund allowance, this maintenance window's table list — and the agent operates inside it without per-call gates, with the envelope itself audited and expiring. The honest framing to give an interviewer is that there is no universally right timeout or approval rate. There is a per-action-class tradeoff between the cost of a wrong action and the cost of delay, plus a hard constraint on how much genuine human attention exists — and a design that ignores the second one degrades into rubber-stamping regardless of how good the first looks on paper.

  • Which metric most cleanly tells you a gate has become a rubber stamp?
    Approval rate paired with median time-to-decision. Near-total approval with dwell times of a few seconds means the reviewer cannot be evaluating anything, and a rejection rate flat at zero over months confirms it. Segment by approver too — one person approving everything in under two seconds while colleagues take minutes is an individual signal, not a policy one. Treat both as SLOs on the gate itself, reviewed like any other production metric.
  • Is defaulting to allow on timeout ever the right call?
    Only for reversible, contained actions where the cost of stalling exceeds the cost of the action — an internal draft, a sandbox change, a retry of an idempotent step. It must be an explicit per-action-class policy entry with its own audit marker, never an inherited default, and never for anything that moves money, leaves the system, or destroys data. If you find yourself wanting allow-on-timeout for an irreversible action, the real answer is that the action should not have been gated by a human at all.
  • How do you keep an agent useful when approvers are asleep?
    Tier so the overnight-common cases are automatic, run a follow-the-sun rota for what remains, and use pre-authorized envelopes: a human approves a bounded scope in advance — a spending allowance, a named set of tables for a maintenance window — and the agent works inside it without per-call gates. The envelope itself is audited, expiring and revocable. Anything outside it queues to the morning rather than proceeding, and the queue depth is the metric that tells you whether the envelope was drawn correctly.

saying these in an interview costs you the question

  • Lets pending approvals sit forever with no expiry
  • Auto-approves irreversible actions when the timeout elapses
  • Routes every request to one named person with no rota
  • Treats a near-100% approval rate as proof gates work
  • Adds gates without budgeting the human attention to service them

context