skip to content

How do you split an agent eval suite across CI tiers under a fixed cost budget?

level: principalimportance: should knowfreq 38%

answer

  1. tasks times rollouts times tokens
  2. breadth cheap, depth expensive
  3. thirty tasks cannot resolve five points
  4. one looping run can eat the night
  5. enforce the ceiling, do not hope

basics

~20 s

Spend the cheap tier on breadth and the expensive tier on depth. A small stubbed subset at one rollout per task runs on every change to catch breakage; the full suite at several rollouts, with judges and live-ish tools, runs nightly inside a hard spend and wall-clock cap.

solid answer

~50 s

Total cost scales with tasks times rollouts times tokens, so the budget question is how to allocate those three. A per-change tier of roughly 30 representative tasks at k=1, stubbed tools and programmatic verifiers only, finishes in minutes for a few dollars — enough to catch a prompt or tool-schema change that breaks everything, and deliberately not enough to resolve a small quality difference, since 30 tasks carry sampling error of several points. The nightly tier runs the full suite at k=3 to 5 with judged tasks included, parallelized to fit a wall-clock cap and bounded by provider rate limits, under a hard spend ceiling that aborts rather than overruns. A pre-release tier runs high k on the reliability-critical subset and exercises live integrations. Cap per-rollout spend too, so one looping agent cannot consume the whole night's budget.

go deeper

for a junior

Know that eval cost scales with the number of tasks times the number of repeated runs, and that teams therefore run a small fast subset frequently and the full suite less often.

for a middle

Explain the tiering concretely: a stubbed, programmatically-verified subset at one rollout per change; the full suite with repeats and judges on a nightly schedule. Say what each tier can and cannot detect.

for a senior

Demonstrate enforcement and selection: per-rollout token and wall-clock caps, a suite-level spend abort, concurrency bounded by sandbox slots and rate limits, prompt caching on the stable prefix, and a fast subset chosen for coverage rather than cheapness.

for a principal

Own the allocation argument — breadth buys resolution, depth buys reliability, and neither substitutes for the other — and set the budget against product risk. Be candid that the constants are workload-specific and come from measuring your own rollout cost, not from a rule of thumb.

## The cost model Eval spend is roughly *tasks x rollouts x tokens-per-rollout*, plus fixed per-rollout overhead (sandbox startup, fixture restore) and any judge calls, which are themselves *tasks x rollouts x judge tokens*. Agentic tasks are token-hungry because every tool result re-enters the context on the next turn, so a single long-horizon rollout can cost far more than a chat-style eval. Wall clock is a separate, partly independent constraint: it is bounded by rollout latency divided by achievable concurrency, and concurrency is bounded by sandbox isolation slots and provider rate limits, not by ambition. With a fixed budget, the design question is where each of those three factors buys the most information. ## What each factor buys **More tasks buys resolution.** The sampling error on a suite-level pass rate falls roughly as one over the square root of the number of tasks. At 30 tasks, the standard error is around 8 points: that suite can tell you the agent broke, but it cannot tell you 80% from 85%. Distinguishing small differences is a task-count problem, and no amount of repetition fixes it. **More rollouts buys reliability information.** k>1 is how flaky tasks become visible and how pass^k becomes computable. It says nothing about tasks you did not include. **More tokens per rollout — bigger contexts, more turns, higher effort settings — buys capability headroom** but makes every other axis more expensive. ## A tiered allocation **Tier 1, per change.** A curated subset — say 30 tasks chosen to span the distinct capabilities and tool paths, not the 30 easiest — at k=1, stubbed or replayed tools, programmatic verifiers only, no judge. Target: single-digit minutes, a few dollars, and no flakiness. Its job is breakage detection: a malformed tool schema, a prompt edit that destroys instruction following, a broken fixture. Be explicit that this tier is not a quality measurement, or people will over-read its number. **Tier 2, nightly.** The full suite — for a support agent, perhaps 120 golden tasks over a seeded database — at k=3 to 5, judged tasks included, run under an explicit ceiling such as a fixed spend cap and a wall-clock cap. Parallelize to fit, allocating one isolated sandbox per worker, and throttle to stay under provider limits. This is the tier that produces the numbers you actually reason about. **Tier 3, pre-release or weekly.** High k (8 or more) on the reliability-critical subset, plus live contract checks against real integrations. Expensive, infrequent, and the only tier that reports pass^k you would quote externally. ## Enforcing the ceiling A budget that is not enforced is a wish. Three controls matter. *Per-rollout caps.* An agent that loops can consume a disproportionate share of the night. Cap tokens, turns and wall clock per rollout, and record the exhausted runs as a distinct outcome rather than silently as failures — the exhaustion rate is itself a signal. *Suite-level abort.* Track cumulative spend during the run and stop cleanly when the ceiling is hit, reporting partial results with the number of tasks completed. A truncated report you can trust beats an overrun. *Cheapening the tokens.* Stable prefixes (system prompt, tool definitions, fixture description) are identical across every rollout of every task, which makes prompt caching unusually effective for evals. Stubbed tools also cut tokens by returning tight payloads instead of sprawling real ones. ## Choosing the subset The cheap tier's task selection is the highest-leverage decision in the whole scheme, and it is not random sampling. Cover each distinct tool path, each major capability, and the failure modes that have actually bitten you in production — a regression that has happened once will happen again. Keep a couple of tasks that are known-hard and expected to fail; if they suddenly pass, either the agent improved or a verifier broke, and both are worth knowing. Rotate the subset periodically so the fast tier does not become the only thing anyone optimizes for. ## Honest limits None of this is settled practice, and the right numbers are workload-specific: a suite gating an agent that moves money deserves a budget an internal drafting assistant does not. State the principle — the cheap tier detects breakage, the expensive tier measures reliability, and each tier's budget is enforced rather than hoped for — and then say the constants come from measuring your own rollout cost and latency, not from a rule of thumb. ## How to answer Start from the cost identity, explain what each factor buys, propose the tiering, and be specific about enforcement and subset selection. The weak answer runs the whole suite on every commit and is surprised by the bill, or shrinks k and task count together until the suite is cheap and meaningless.

  • Why does adding rollouts not compensate for a small task count?
    They measure different things. Repeats estimate how reliably the agent solves the tasks you chose; they carry no information about tasks you left out. Suite-level sampling error falls with the number of tasks, roughly as one over the square root of that count, so a 30-task suite has several points of noise no matter how many times each task is repeated. Breadth buys resolution; depth buys reliability.
  • What makes prompt caching unusually effective for eval harnesses specifically?
    Evals replay a near-identical prefix thousands of times: the same system prompt, the same tool definitions, often the same fixture description, across every rollout of every task and across every repeat. That is the ideal shape for a cache — a long, stable prefix with variation only at the tail. Ordering the prompt so everything invariant comes first, and keeping tool definitions byte-stable between runs, can cut a substantial fraction of the suite's token bill.
  • How should the harness report rollouts that hit their token or wall-clock cap?
    As a distinct outcome, not as a plain task failure. Exhaustion means the agent did not finish rather than that it answered wrongly, and the rate of exhaustion is a first-class signal about efficiency and looping. Folding it into the failure count hides a trend and makes budget tuning invisible. Report completed, failed and exhausted separately, and track exhaustion over time alongside cost per rollout.

saying these in an interview costs you the question

  • Running the full repeated suite on every commit until the bill lands
  • Reading a 30-task smoke number as a quality measurement
  • Raising repeats to fix a suite that is too small
  • Leaving per-rollout token and time caps off, so one loop dominates spend
  • Choosing the fast subset by picking the cheapest or fastest tasks

context