skip to content

In agent evals, what does pass^k measure that a single run per task hides?

level: middleimportance: must knowfreq 68%

answer

  1. one run flatters an agent
  2. reliability, not best-case ability
  3. every repeat must succeed
  4. 0.9 to the eighth power
  5. the pessimistic sibling of pass@k

basics

~20 s

pass^k is the share of tasks an agent solves on all k repeated attempts, so it measures reliability rather than best-case ability. A single run hides variance: roughly 90% per-attempt success clears all 8 attempts on only about 43% of tasks.

solid answer

~50 s

Agent runs are nondeterministic — sampling, tool timing and ordering, simulated users, and even hosted inference at temperature 0 all vary — so one rollout per task tells you what the agent *can* do, not what it *will* do. pass^k runs each task k times and counts it solved only if every rollout succeeded, which is the metric that matches a user's experience of a product that must work every time. Treat it as the empirical fraction of all-k-success tasks, not as a formula; but the independence approximation is useful for intuition: 0.9^8 is about 0.43, so a suite reporting 90% on a single run can be a coin flip at k=8. The design consequence is that a harness must be built to re-run a task cheaply and identically, because k multiplies every cost you have.

go deeper

for a junior

Know that agents give different answers on repeated runs, so a single pass or fail per task is not evidence. Be able to say that pass^k means the agent succeeded on every one of k attempts.

for a middle

Explain where the nondeterminism comes from and do the arithmetic out loud: about 0.9 to the eighth is 0.43, so a 90% headline can be a coin flip at k=8. Distinguish pass^k from pass@k and say which fits an unattended agent.

for a senior

Show how k reshapes the rig: perfect environment reset between rollouts, isolated sandboxes for concurrency, and cost that scales linearly with k. Explain how you pick k per suite tier and why per-task success counts beat the aggregate for diagnosis.

for a principal

Own the framing that reliability, not peak capability, is the number the business ships on, and defend a budget split between more tasks (resolution) and more repeats (reliability). Be ready to argue which product surfaces genuinely need pass^k at high k and which can tolerate retry-and-filter.

## The problem pass^k exists to solve An agent eval harness runs an agent over a suite of tasks and records, per task, whether it succeeded. The naive report is a single number: how many of the N tasks passed. That number is produced by exactly one rollout per task, and it is close to meaningless for an agentic system, because the same agent given the same task twice can take different paths and reach different outcomes. **Nondeterminism has several independent sources.** Token sampling is the obvious one, but setting temperature to zero does not remove it: hosted inference batches requests, and floating-point reduction order changes with batch composition, so identical prompts can produce different tokens. Beyond the model, tool latency and completion order vary, retries fire or don't, timestamps and generated IDs differ, and — in suites that grade conversational tasks — the simulated user is itself a model with its own variance. Any of these can flip a run. ## What pass^k is **pass^k is the fraction of tasks on which *all* k independent rollouts succeed.** It is the reliability metric, and it is deliberately the pessimistic sibling of pass@k, which asks whether *at least one* of k attempts succeeded. pass@k is the right question when a human filters candidates (you generate five patches and a reviewer picks one). pass^k is the right question for an autonomous agent that will run unattended against a real customer, because there the bad rollout is not discarded — it ships. The arithmetic is what makes the metric bite. If a task's per-attempt success probability is p and rollouts are roughly independent, the chance of surviving k attempts is p^k. At p = 0.9 and k = 8, that is 0.43. A headline of "90% task completion" and a reality of "fails at least once in eight tries on more than half the tasks" are the same agent. Agent benchmarks that grade multi-turn, stateful tasks — tau-bench's retail and airline domains, for instance — report pass^k across repeated rollouts precisely to expose this gap. Treat p^k as intuition only. Real rollouts are not independent: some tasks are deterministic for a given agent (it always passes, or always fails), while others are genuinely coin-flippy. So the empirical pass^k over a suite is usually higher than plugging the mean pass rate into p^k would suggest, and the *distribution* matters more than the average. A useful companion view is the per-task success count out of k: it separates "20 tasks are hard" from "60 tasks are flaky". ## What this forces on the harness Once you commit to k > 1, every design decision in the rig changes. **Repeatability becomes mandatory.** Rollout 2 must start from exactly the state rollout 1 started from — a restored database snapshot, a fresh container, a reset filesystem. If reset leaks, rollout 2 sees rollout 1's side effects and the k runs are not repeats of the same task at all. **Cost multiplies by k.** Tokens, wall-clock, sandbox startup, and judge calls all scale linearly with k. This is why k is a budget decision, and why suites usually run k = 1 on the fast pre-merge tier and reserve high k for a nightly or pre-release run on the subset where reliability is the point. **Parallelism becomes necessary and dangerous.** You will want to run rollouts concurrently to fit the wall-clock budget, which means each concurrent rollout needs its own isolated sandbox and its own fixture instance, and you will hit provider rate limits. **Choosing k is a resolution question.** k = 1 measures capability and detects gross breakage. k = 3 to 5 is the usual working compromise: enough to surface obvious flakiness without a five-fold bill. k = 8 or more is for the small set of tasks whose reliability you actually have to defend. Note that k does not fix the other axis: with only 30 tasks, the sampling error on the pass rate is around 7–9 percentage points, so more repeats will not let you resolve a small quality difference that more tasks would. ## How to talk about it The strong answer names the metric, gives the arithmetic that makes it alarming, identifies where the nondeterminism comes from (including that temperature 0 is not a guarantee), and then connects it to harness design: k repeats are only meaningful if the environment resets perfectly between them, and k is bounded by the eval budget. The weak answer treats pass^k as a synonym for pass@k, or claims that pinning temperature and a seed makes the whole problem go away.

  • How is pass^k different from pass@k, and when is pass@k the metric you actually want?
    pass@k counts a task solved if at least one of k attempts succeeded; pass^k requires all k. pass@k fits workflows where a human or a verifier filters candidates — generate several patches, keep the one whose tests go green. pass^k fits autonomous operation, where the bad rollout is not discarded but delivered to a user. Reporting pass@k for an unattended agent systematically overstates what customers will experience.
  • You set temperature to 0 and pin the model version, but repeated runs still diverge. What explains that?
    Hosted inference is not bit-reproducible: requests are batched, and the reduction order in floating-point kernels depends on batch composition, so the same prompt can yield different tokens. On top of that, the harness itself introduces variance — tool latency and completion ordering, retries, timestamps and generated IDs, network errors, and any model-driven user simulator. Pinning sampling narrows the distribution; it does not collapse it.
  • How would you use per-task rollout counts rather than the aggregate pass^k number?
    Record successes out of k per task and look at the histogram. Tasks at k/k and 0/k are deterministic-pass and deterministic-fail; anything in between is flaky and is where reliability work pays. That split tells you whether a low pass^k means the agent is incapable on a subset or unreliable across the board, which are different fixes — capability work versus determinism, retries, or guardrails.

saying these in an interview costs you the question

  • Treating pass^k and pass@k as the same metric
  • Claiming temperature 0 makes runs fully deterministic
  • Reporting one rollout per task as the agent's reliability
  • Raising k to fix a suite that is simply too small
  • Averaging per-run scores instead of requiring all k to pass

context