How does an advisory token budget an agent paces against differ from an enforced cap?
answer
- one informs, one binds
- the model can read a soft budget
- a fuse does not negotiate
- graceful landing vs hard truncation
- soft number inside a hard one
basics
~20 sAn advisory budget is told to the model so it can pace itself and wrap up gracefully; an enforced cap is applied by the harness and cuts the run off wherever it happens to be. Advisory budgets shape behaviour, enforced caps bound spend.
solid answer
~60 sAn **advisory** budget — providers now expose this as a token-denominated task budget, alongside coarse effort levels — is passed into the request and shown to the model. The model sees roughly how much it has to spend, paces its exploration, and lands on an answer before running out. It changes *behaviour*: with 40,000 tokens the agent reads three files; with 400,000 it reads thirty. What it cannot do is bind, because the model estimates its own consumption and can be wrong. An **enforced** cap — the per-response output-token limit, an iteration cap, a wall-clock deadline, a session cost ceiling — is checked by the harness and fires regardless of what the model intended, truncating mid-sentence if it must. You need both layers: advisory alone leaves you exposed to a model that misjudges, enforced alone gives you an abrupt truncation with no summary of what was learned. In practice the advisory number sits comfortably below the enforced one, so the hard stop is a fuse that should rarely blow.
go deeper
Know that an agent run has a token budget at all, and that a limit you tell the model is different from a limit enforced by the code around it.
Explain the mechanism: an advisory budget is in the prompt so the model paces itself, an enforced cap is checked by the harness and truncates regardless. Say why the soft number sits below the hard one.
Describe the layered arrangement you would ship — advisory budget, enforced token and iteration caps, session cost ceiling — plus the telemetry that tells you which layer is doing the stopping, and what an oversized tool payload does to the model's own pacing.
Own the cost model: effort tiers as a product control, per-tenant currency caps that survive a model swap, and the argument that token budgets bound model spend but not the blast radius of the tools the agent called.
## Two different jobs Budgeting an agent run has two distinct goals that people routinely conflate. The first is **behavioural**: you want the agent to choose an appropriate depth of work — a quick lookup should not trigger a forty-step investigation. The second is **protective**: no matter what the agent chooses, spend must not exceed a bound you can afford. Advisory budgets do the first job. Enforced caps do the second. Neither substitutes for the other. ## Advisory budgets: information the model acts on An advisory budget is a number placed in the model's context — as of mid-2026, providers expose this directly as a token-denominated task budget, and as coarse effort settings that scale how much the model explores. The mechanism is simply that the model can read it. A well-behaved model given a task budget of roughly 40,000 tokens will plan a shallower investigation, skip speculative side-quests, and start assembling its answer as it approaches the limit rather than after it. The property that matters is **graceful landing**: the run ends with a coherent result rather than in the middle of reading a file. The same idea appears as a product-facing control. Exposing "quick" (about five iterations) and "thorough" (about forty) as a request parameter lets the caller buy the depth they actually need, which is usually a better cost lever than tuning a single global cap. It also gives you a clean story for latency: quick mode is a few seconds, thorough mode is a minute, and the user chose. The weakness is that an advisory budget is a *request*. The model estimates its own consumption, and estimates are imperfect. A single tool that returns a 60,000-token log blows the plan the model made three steps earlier. Nothing about an advisory budget stops the run. ## Enforced caps: mechanisms that do not negotiate Enforced limits are checked outside the model. The per-response output-token limit truncates generation. Iteration caps stop the loop. Wall-clock deadlines abort. Session cost ceilings — say two dollars per user session — halt the run whatever it was doing. These are the fuse in the circuit: they do not care what the agent intended, and that is precisely their value. When an enforced cap fires, the output is by definition ungraceful; you may get half a sentence or an unfinished tool argument, and your harness has to handle that. ## Layering them The standard arrangement is a soft number well inside a hard one: - advisory task budget: the depth you actually want, say 40,000 tokens - enforced token/iteration/clock caps: comfortably above it, sized so they fire only on pathology - session cost cap: the outermost ring, denominated in currency so a model swap cannot quietly multiply spend Instrument the gap. If the advisory budget is regularly overshot and the enforced cap is doing the stopping, your soft number is not being respected — either the task genuinely needs more room, or a tool is returning payloads the model could not anticipate. A healthy fleet ends the overwhelming majority of runs by completing inside the advisory budget, with hard caps firing on a small tail. ## Cost is denominated in more than tokens One subtlety: the resource you care about is rarely a single number. A run can be cheap in tokens and expensive in wall clock (waiting on slow tools), or cheap in both and expensive in *risk* (it wrote to production twice). Token budgets bound the model's spend; they say nothing about the side effects of the tools it called. Separate guards — approval gates, rate limits on write tools — cover that, and a candidate who conflates "I capped tokens" with "the agent is safe" has missed the distinction. ## Why interviewers ask Because it separates people who have run agents in production from people who have read about them. The advisory-versus-enforced split is the concrete form of a general operations principle: give the system information so it behaves well, and put a mechanism underneath so that when it behaves badly anyway, the damage is bounded. Reasoning about that layering, and about what a graceful landing looks like at exhaustion, is exactly the judgement the question is probing.
- What goes wrong if you ship only the enforced cap?The run dies wherever it happens to be — mid-sentence, mid tool argument, with no summary of what was learned. The caller gets nothing useful and pays for everything consumed. An advisory budget lets the model see the ceiling approaching and spend its last tokens writing down findings instead of opening another file.
- How would you expose depth as a product control rather than a hidden constant?Offer named effort levels the caller picks per request — a quick tier of roughly five iterations for lookups, a thorough tier of roughly forty for investigations — and map each tier to an advisory budget plus its own enforced caps. Callers buy the latency and cost they need, and you get a clean per-tier cost model instead of one global compromise.
- A single tool returns a 60,000-token payload and blows the plan. What do you change?Bound the tool, not just the loop: truncate or paginate the result at the tool boundary, return a reference the agent can query rather than the whole payload, and report the real size so the model knows what it skipped. Budgets the model paces against only work if no single observation can dwarf them.
saying these in an interview costs you the question
- Believes telling the model a budget guarantees it is respected
- Ships a hard cap only, then wonders why output is truncated mid-sentence
- Sets the advisory and enforced numbers equal
- Caps tokens and calls the agent safe, ignoring tool side effects
- Uses one global budget for every task class