You are about to run a multi-step agentic harmful-task suite against a metered hosted endpoint. How do you estimate the model and tool calls it will consume before you start, and what does capping the steps per task do to the results you get out?
answer
- tasks x steps x turns x repeats
- paper estimate is always low
- pilot ten tasks, measure tokens
- context grows, late turns cost more
- spike at the cap = you measured your budget
basics
~20 sMultiply tasks by required steps by model turns per step, add retries and repeats, then pilot ten tasks and measure the real per-task call count before extrapolating. A step cap bounds the bill but truncates long chains, so attempts that hit it look like non-completions unless you record the cap as their stop reason.
solid answer
~50 sEstimate top-down, then correct with a pilot. The paper estimate is `tasks x mean required steps x turns per step x repeats`, plus a retry allowance, plus the tool calls, which are cheap if the environment is mocked locally and are not if any tool reaches a paid service. The paper number is always low, because agents take exploratory turns that are not on the required path — reasoning turns, wrong tool choices, re-reads. Run ten representative tasks, measure the actual turns and tokens per attempt, and scale from that. The control you have is the per-task step or turn cap. It bounds worst-case spend on a runaway loop, and it biases the measurement: every attempt that would have finished at cap+1 is recorded as unfinished. Check the depth distribution for a spike at the cap. If there is one, the cap, not the agent, is producing your headline number, and you should either raise it for the affected tasks or report the truncated fraction beside the result.
go deeper
Knows a run costs calls per step per task and that a cap is needed so an agent cannot loop forever.
Builds the multiplicative estimate, includes repeats and retries, and knows the cap turns some genuine completions into recorded non-completions.
Pilots before committing, extrapolates on tokens because context grows, sets attempt timeouts, a sweep-level ceiling and checkpointing, and audits the depth histogram for a spike at the cap.
Decides what the suite is worth per cycle, at what cadence it runs, and what evidence is required before the number is allowed to be quoted outside the team.
**Build the estimate in layers, then throw it away.** The paper estimate is a product, and every factor after the first is one people forget: 1. *Required path* — `tasks x mean required steps`. This is the floor; it is never what you pay. 2. *Turns per step* — agents emit planning turns, malformed calls that get retried, and re-reads of tool output. A multiplier of two to four over the required path is unremarkable. 3. *Repeats* — sampling is stochastic, so one attempt per task yields a number you cannot put an interval on. Harnesses expose this directly (Inspect AI calls a full repeated pass an *epoch*), and it multiplies the whole estimate linearly. 4. *Failure work* — attempts that error out consumed calls before they failed, and retried transport errors consumed them twice. 5. *Tool side* — free when every tool is a local stub; a real cost line and a rate-limit surface the moment one stub proxies a paid API. A worked shape: 400 tasks x 3 repeats x ~10 turns is roughly 12,000 completions. If the mean input across an attempt is ~6k tokens (it starts small and grows) and output ~300, that is ~72M input and ~3.6M output tokens; at illustrative prices of $3 and $15 per million, ~$270 for the sweep. Serialised at 4s a turn that is over 13 hours; at concurrency 8, closer to 100 minutes — assuming the endpoint tolerates the concurrency. **Then pilot, because the arithmetic is always low.** Take ten tasks spanning the shortest and longest chains, run them at the real settings, and read the harness log for turns, input tokens and output tokens per attempt. Extrapolate on **measured tokens, not turn counts**: each turn resends the accumulated transcript plus every tool result so far, so cost per turn climbs through an attempt and a flat per-turn average taken from short pilots badly under-predicts long ones. Watch for the opposite error too — if the provider applies prompt caching to the repeated prefix, a cached pilot extrapolated to a sweep whose cache keeps expiring will under-predict, and an uncached pilot extrapolated to a well-cached sweep will over-predict by a large factor. Check whether cached-input tokens appear in the usage numbers before trusting either. **What the caps do.** A per-attempt limit on messages, tokens or wall-clock is the only hard bound on an agent that loops, and you need one. It is also a measurement instrument, and a badly chosen one manufactures your result: every attempt that would have finished at limit+1 is recorded as unfinished. The diagnostic is the depth histogram — a mass sitting exactly at the configured boundary means the suite is reporting your budget rather than the agent's behaviour. Record the limit value with the run, record limit exhaustion as its own stop reason, and never let it merge into "did not complete". **Where the cost number and the result number both mislead.** Concurrency is the usual culprit. Push it up to fit the sweep into the evening and the endpoint starts returning rate-limit errors; a harness that retries them inflates the bill invisibly, and a harness that records them as sample failures silently converts throttling into "the agent did not complete the task" — a safety result manufactured by your own request rate. The same applies to attempt timeouts: a slow stub or a long generation trips the clock and the attempt lands in the same bucket as a refusal. Any spend forecast built only on the required path will also miss the fact that failed and retried attempts are paid for in full. **Other guards worth setting.** A wall-clock timeout per attempt so a hung stub cannot stall the sweep; a sweep-level spend ceiling that halts and checkpoints rather than killing a single attempt; and per-task checkpointing so an interrupted run resumes instead of being re-paid from zero. **What to check.** Measured tokens per attempt against the pilot's, and the harness's token accounting against the provider's own usage figures — they diverge when caching or retries are in play. Then the share of attempts ending at the limit, the share ending on a harness or transport error broken out from refusals, whether any resumed portion of a sweep ran with the same limits and the same environment build as the first portion, and whether the concurrency setting changed between the pilot and the sweep.
- Why extrapolate from measured tokens rather than from turn counts?Each turn carries the accumulated transcript plus tool output, so the last turns of a long chain cost several times the first. A per-turn average taken from short attempts under-predicts the long ones badly.
- Where do you put a global spend ceiling, and why not just rely on the per-task cap?The per-task cap bounds one attempt; it does nothing about a suite that is simply larger or more expensive than you modelled. A sweep-level ceiling that halts and checkpoints protects the budget as a whole and lets you resume deliberately.
A step cap is a stopwatch stopped at two hours on a marathon: everyone who would have crossed at 2:05 is written down as a did-not-finish, and the results table describes your clock rather than the runners.
saying these in an interview costs you the question
- Estimates only the required-path calls and is surprised by the bill.
- Runs one attempt per task and reports the result as if it had an interval.
- Sets a step cap without recording cap exhaustion as a distinct stop reason.
- Never inspects the depth distribution for a boundary spike at the cap.
- No checkpointing, so an interrupted sweep is re-paid from the start.