skip to content

Agent Loops and Control Flow

The control flow around the model call: observe, think, act, repeat — plus the stopping conditions, token budgets, and human checkpoints that keep the loop from running forever. This is the operational half of agent design, and interviewers probe it because runaway loops are expensive and irreversible actions are worse.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

questions

13

Which AI-agent tool calls need human approval, and how do you tier the rest?

level: middleimportance: must knowfreq 72%

answer

  1. not every tool carries the same risk
  2. think blast radius
  3. reversible versus irreversible, and how wide
  4. policy keys on arguments, not names
  5. unclassified tools default to human

basics

~20 s

Gate by blast radius, not by tool name. Anything irreversible, externally visible, or above a value threshold — money movement, production schema changes, outbound messages, deletions — needs a human. Reads and cheaply reversible writes run automatically.

solid answer

~50 s

I tier tools by reversibility and blast radius rather than gating everything. Tier 0 is read-only: run automatically, log it. Tier 1 is reversible, contained writes — a draft, a ticket, a sandbox file — automatic with an undo path. Tier 2 is irreversible or externally visible: money movement, production DDL, emails to customers, deletions. Those pause for a human. Tier 3 is not exposed to the agent at all. The policy has to key on **arguments**, not just the tool name, because the same tool spans tiers: refunds at or under $200 run automatically, chargeback reversals always need a person; a migration against staging is tier 1, the same statement against production Postgres is tier 2. Enforcement lives in the harness that executes the call, never in the system prompt — prompt-level rules are advisory and fall over under prompt injection. New, unclassified tools default to deny.

code

yaml · 16 lines
yaml
default: human            # unclassified tools are gated until reviewed
tools:
  search_docs:
    tier: auto
  create_ticket:
    tier: auto
  issue_refund:
    tier: auto_if
    condition: "args.amount_usd <= 200"
  reverse_chargeback:
    tier: human
  run_sql:
    tier: auto_if
    condition: "args.env != 'production' and args.kind == 'select'"
  rotate_root_credentials:
    tier: forbidden

go deeper

for a junior

Know that an agent should not be able to move money, delete data, or message customers without a person confirming, and that reads are usually fine to run automatically.

for a middle

Be ready to lay out an explicit tier table and explain that the policy is evaluated over the tool arguments, not just the tool name, with unclassified tools defaulting to approval.

for a senior

Show that enforcement lives in the harness and sits on top of scoped credentials and sandboxes, and that you would pick thresholds from real traffic so the gate fires on the tail rather than on every ticket.

for a principal

Own the tradeoff between safety and throughput across a fleet: which actions should be removed from agents entirely rather than made approvable, who owns the policy table, and how you measure whether each gate still changes decisions.

## Why the gate exists at all An agent loop is a model choosing actions from a tool list. The model is probabilistic, its inputs include untrusted text it fetched, and it has no concept of "I cannot take this back." A human checkpoint is the cheapest safety mechanism available: it costs one round trip of latency and buys a veto on exactly the actions you cannot undo. The engineering question is never *whether* to have gates — it is which actions sit behind one, because a gate on everything is the same as no gate at all once approvers stop reading. ## Classify by blast radius, not by tool The useful axes are: - **Reversibility.** Can you undo it with a compensating action you already have, and would you notice in time? A created draft is reversible. A sent email is not. A deleted index can be rebuilt, but the rebuild may take hours on a hot production table, so treat "reversible but expensive" as irreversible. - **Externality.** Does the effect leave your system? Messages to customers, payments, filings, posts, calls to third-party APIs that mutate their state. - **Magnitude.** Value moved, rows touched, users affected. - **Data sensitivity.** A read is not automatically safe: reading a table of customer records into a context window that later reaches an external tool is an exfiltration path, not a read. A practical four-tier ladder: | Tier | Examples | Policy | |---|---|---| | 0 read-only | search, lookup, query non-sensitive data | auto, logged | | 1 reversible write | create draft, open ticket, write to a scratch file, sandbox change | auto, undoable, logged | | 2 gated | wire transfer, production DDL, delete, send-to-customer, refund above threshold | human approval required | | 3 forbidden | rotate root credentials, disable audit logging, mass delete | not in the tool list | Tier 3 matters more than people expect. Some actions should not be *approvable* by whoever happens to be on call; removing them from the agent's surface entirely is stronger than any checkpoint. ## The policy is a function of arguments Gating on tool identity alone is the most common design error. One `issue_refund` tool covers a $12 shipping credit and a $40,000 enterprise credit; one `run_sql` tool covers `SELECT` on staging and `DROP INDEX` on the production primary. Write the policy as a predicate over the resolved call: tool name plus arguments plus environment plus the caller's own entitlements. Typical predicates are amount thresholds, environment allowlists, statement-type checks (any DDL is gated, DML under a row-count estimate is not), and recipient allowlists for outbound channels. Thresholds also let you preserve throughput. A support agent that auto-approves refunds at or under $200 handles the overwhelming majority of tickets end to end, and the small tail that reaches a human is genuinely worth a human's minute. That is the whole design goal: push the gate to where it changes decisions. ## Enforce in the harness, default to deny A rule written in the system prompt — "always ask before sending money" — is a suggestion to a probabilistic system that is also reading attacker-controlled web pages and tool output. It is worth writing, because it improves behaviour and makes the agent's proposals better shaped, but it is not a control. The control is the tool router: the harness resolves the call, evaluates the policy, and either executes or suspends. The model cannot argue with code it does not run. Default-deny closes the other common hole. When an engineer registers a new tool and nobody classified it, the safe behaviour is that it requires approval until someone reviews it — not that it inherits the auto tier. Unclassified-means-gated turns a missing review into slow throughput instead of a silent incident. ## Layering with least privilege Approval gates are the last layer, not the only one. Beneath them sit scoped credentials (the agent's database role simply cannot `DROP`), network egress limits, per-tenant data scoping, and sandboxes. A gate protects against a bad decision; a credential boundary protects against a compromised loop that never reaches your gate code. Interviewers like the candidate who says the checkpoint is a *decision* control layered on top of *capability* controls, because it shows they know a human click is not a security perimeter. ## What good looks like in review A reviewable policy is a small, declarative table, versioned in the repo, with a test per tier; every gated execution carries an approval record; and there is a metric for how often each gate fires and how often the answer was no. If a gate has never once been answered "no", it is either mis-scoped or nobody is reading it — both are findings.

  • Beyond a binary approve or reject, what other intervention should a checkpoint offer?
    Let the human edit before they accept. Editing the tool arguments — trimming a refund to $180, narrowing a SQL predicate — or editing the proposed plan turns a dead end into a corrected action, and it captures the reviewer's intent instead of forcing a reject-and-reprompt cycle. Also offer approve-and-remember for a class of call, and free-text feedback that flows back into the loop as an observation. The cost is a bigger, better-validated approval surface: an edited call still has to be re-checked against policy before it runs.
  • Why isn't a system-prompt rule like "always ask before deleting" a sufficient gate?
    Because the model both proposes the action and decides whether the rule applies, and it is reading tool output and web content an attacker may control. A prompt rule is advisory: injection, goal drift, or an ambiguous edge case will eventually route around it, silently and without a log. The enforceable version lives in the tool router, which evaluates policy over the resolved call. Keep the prompt rule too — it makes proposals better shaped — but do not count it as a control.
  • Reads are usually auto-approved. When is a read worth gating?
    When the read itself is the sensitive act: pulling a full customer PII table, opening another tenant's records, or fetching credentials. Once that data is in the context window, any later outbound tool can carry it out, so a read plus an egress channel is an exfiltration path. Gate reads by scope rather than by tool: broad or cross-tenant queries need approval, targeted ones do not, and per-tenant scoping should make most of the dangerous shapes unrepresentable.

A pharmacy lets any technician count out ibuprofen but requires a second signature for controlled substances — the countersignature is placed where a mistake cannot be walked back, not on every transaction.

saying these in an interview costs you the question

  • Says a system-prompt instruction to ask first is enough
  • Gates by tool name only, ignoring amounts or target environment
  • Treats every write as equally risky and gates all of them
  • Assumes reads are always safe regardless of what they pull
  • Adds new, unclassified tools as auto-approved by default

context

open as a page

Why does agent self-correction improve results with a test suite but often not without one?

level: middleimportance: must knowfreq 72%

basics

~20 s

Self-correction needs a signal the model did not produce. Failing tests, compiler errors and schema violations are outside evidence of a defect. Pure self-judgment re-samples the same model that wrote the output, so it rarely catches what it already missed.

open as a page

What stop conditions end an agent loop besides the model returning no tool call?

level: middleimportance: must knowfreq 62%

basics

~20 s

Agent loops end naturally when the model returns a final answer with no tool call, or calls an explicit done tool. They end forcibly on a max-iteration cap, a wall-clock deadline, a token or cost budget, or an unrecoverable error.

open as a page

How do you pause an AI agent for human approval that may take days?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Persist the loop's state to a durable checkpoint, emit the approval request, and let the process exit. When the human answers — hours or days later, on a different machine — load the checkpoint, inject the decision, and resume from the interrupt point. Never block a thread.

open as a page

Why can asking an LLM "are you sure?" turn a correct answer into a wrong one?

level: juniorimportance: should knowfreq 56%

basics

~20 s

A challenge like "are you sure?" supplies doubt, not evidence. Models tend to go along with implied disagreement, so the second pass often swaps a correct but unusual detail — an exact date, an odd spelling — for a more ordinary-sounding one, lowering accuracy.

open as a page

Why split an agent's generator and critic into two separate prompts?

level: middleimportance: should knowfreq 46%

basics

~20 s

Separation stops the critic from inheriting the generator's framing. A fresh prompt with its own system message, the task, the draft, and explicit failure criteria — but not the generator's reasoning — reviews the artifact on its merits instead of defending the path that produced it.

open as a page

What is self-consistency sampling, and when does it beat an iterative critique pass?

level: middleimportance: should knowfreq 45%

basics

~20 s

Self-consistency samples the same prompt several times at non-zero temperature and returns the answer the samples agree on most often. It beats iterative critique when the task has one comparable answer, because agreement is a real signal that needs no critic and the samples run in parallel.

open as a page

What must an AI agent record when a human approves a destructive action?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Record who approved — from an authenticated identity, never from model output — what exactly they saw and accepted as literal tool arguments, when, against which plan version, and the decision itself. Then execute those stored arguments so the approval binds to a specific action.

open as a page

How many reflection rounds should an agent run before it stops revising?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Most of the gain arrives in the first one or two rounds, so cap revision at two or three and stop earlier on signal: the verifier passes, or a round fails to reduce the failure count. Every extra round costs a full generation and its latency.

open as a page

How does an advisory token budget an agent paces against differ from an enforced cap?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An advisory budget is told to the model so it can pace itself and wrap up gracefully; an enforced cap is applied by the harness and cuts the run off wherever it happens to be. Advisory budgets shape behaviour, enforced caps bound spend.

open as a page

How do you detect an agent loop that keeps repeating steps without making progress?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Hash each step — tool name, arguments and the observation it returned — and count repeats. When the same hash recurs a few times, or a window of recent hashes cycles, the agent is stuck. Trip a no-progress stop instead of waiting for the iteration cap.

open as a page

How do you set approval timeouts for an AI agent without creating rubber stamps?

level: principalimportance: should knowfreq 38%

basics

~20 s

Give every pending approval an expiry sized to the action's real urgency, and default to deny on expiry for anything irreversible. Route to a rota rather than one person, and gate few enough actions — with good enough evidence — that approvers still read them.

open as a page

When an agent exhausts its budget mid-task, should it return partial results or continue?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

Return the partial result with an explicit exhausted marker whenever partial work is usable — a bug list missing the last level beats nothing at all. Continue only when the task is checkpointable and someone has knowingly authorised the extra spend.

open as a page