skip to content

How do you split an LLM eval suite between a per-PR gate and a nightly run on a fixed budget?

level: principalimportance: should knowfreq 38%

answer

  1. the bill and the wait are design inputs
  2. not every case has to run every time
  3. stratify, do not randomly sample
  4. the slow tier blocks the release, not the merge
  5. score the subset against what the full run finds

basics

~20 s

Treat cost and wall-clock as gate design constraints. Put a small stratified subset covering every failure mode on pull requests so it finishes in minutes, run the full suite nightly and before release, and measure how many nightly regressions the subset would have caught.

solid answer

~50 s

Cost and feedback latency are first-class design inputs, not afterthoughts. If the full suite is 900 cases at roughly fourteen dollars and twenty minutes a run, it cannot sit on every pull request — the spend scales with merge volume and the wait pushes engineers to merge around it. So stratify: pick a subset of around 120 cases that covers every failure mode and every critical slice, weighted toward the cases that historically move when something breaks, and size it so a run finishes inside about four minutes with parallel execution. That is the blocking gate. The full suite runs nightly and on release candidates, blocking the release rather than the merge. Then measure the subset itself: of the regressions the nightly catches, what share would the PR subset have caught? That recall number is the real quality of your gate, and it tells you whether to grow the subset or re-weight it.

go deeper

for a junior

Know that eval runs cost money and time, so teams usually run a small fast set on each pull request and the full suite on a schedule rather than everything everywhere.

for a middle

Explain the tiering — fast blocking subset, nightly full suite, release-candidate run — and the concrete levers that shrink a run: bounded concurrency, a cached shared prefix, and deterministic graders instead of extra model calls.

for a senior

Show how you would build the subset by stratifying over failure modes, critical slices, and historical discrimination, and how you would keep both cost and duration reported on every build so the budget stays visible.

for a principal

Own the trade explicitly: earlier detection versus a merge loop engineers will actually keep. Measure the fast tier's recall against the nightly, set the spend policy, and re-derive the subset as failure modes shift.

## The constraint is real and it is two-dimensional An LLM eval suite costs money per run and takes wall-clock time. Both scale with case count, both are paid on every execution, and they pull in opposite directions from coverage. A 900-case suite at around fourteen dollars per run is a rounding error once a night and a serious line item at forty merges a day; twenty minutes of wall clock is fine overnight and is a productivity tax on every pull request. Engineers respond to slow gates predictably — they batch changes, they context-switch away, and eventually they look for the flag that skips it. A gate that people route around protects nothing, so latency is a correctness property of the gate design, not a nicety. ## Tiering The standard shape is three tiers with different jobs: **The pull-request tier** is small, fast, and blocking. Its job is to catch the obvious, high-frequency regressions in the time a reviewer spends reading the diff. Target a few minutes end to end. **The nightly tier** is the full suite, run on the main branch. Its job is coverage. It does not block anyone's merge; it opens a ticket, pages the owning team, and — crucially — blocks the release candidate. **The release tier** is the full suite plus anything too slow or expensive for either of the above: long multi-turn scenarios, large-context cases, expensive judge rubrics. It runs on a candidate build before promotion. ## Choosing the pull-request subset The subset is not a random sample. Stratify it: - **Cover every failure mode.** One or more cases per known failure category, so no category is invisible at merge time. - **Cover every critical slice.** Anything with a hard guard — safety, legal, schema compliance — belongs in the fast tier by definition, because those must never merge broken. - **Weight by historical discrimination.** Some cases move whenever anything breaks; others have passed unchanged for a year. Score cases by how often they have flipped alongside a genuine regression and keep the discriminating ones. - **Prefer deterministic graders.** Cases graded by code are faster, cheaper, and quieter than judged ones, so they earn their place in a latency-constrained tier at a much better rate. - **Exclude the quarantined.** Advisory cases do not belong in a blocking tier. ## Engineering the wall clock down Most of the four minutes is waiting on the network, so concurrency is the main lever: run cases in parallel up to the provider's rate limit, with bounded workers so a burst does not trigger throttling that costs more time than it saves. Beyond that, share a cached prompt prefix across cases where the system prompt is common, keep reasoning effort at the level the case actually needs rather than the maximum, and run deterministic graders locally rather than as extra model calls. Report the run's cost and duration on every build so the budget stays visible instead of surfacing as a surprise invoice. ## Measure the gate, not just the product The question that makes this a design problem rather than an arbitrary trim is: **what does the fast tier miss?** You can answer it empirically. Every regression the nightly catches is a labelled example; replay it against the pull-request subset and ask whether the subset would have caught it. The resulting recall — say, the subset catches 80% of what the full suite catches — is the honest description of the gate. If it is low, the subset is mis-weighted, not merely too small, and the fix is usually to swap in cases from the categories that keep slipping through rather than to double the case count. Re-derive it periodically, because as failure modes shift, yesterday's discriminating cases become today's permanent passes. ## The trade to state out loud A fast blocking tier plus a comprehensive non-blocking tier means some regressions land on the main branch and are caught hours later. That is a deliberate choice: the cost is a revert or a follow-up fix; the benefit is a merge loop people actually use and a bill that scales sub-linearly with team size. The alternative — full coverage on every pull request — buys strictly earlier detection at a price that rises with headcount and a wait that erodes the gate's legitimacy. State the trade explicitly, and pick the split from your actual regression frequency and revert cost rather than from a preference for thoroughness. The teams that get this wrong usually do so by being too thorough at the merge point and then quietly disabling the whole thing three months later.

  • How do you decide which cases make it into the fast pull-request tier?
    Stratify rather than sample. Take at least one case per known failure mode, every critical-guard case, and then weight by historical discrimination — cases that have flipped alongside real regressions earn a slot, cases that have passed unchanged for a year do not. Prefer deterministic graders because they are cheaper and quieter per unit of coverage, and exclude anything currently quarantined.
  • How do you know the fast subset is good enough rather than just small?
    Measure its recall against the nightly. Every regression the full suite catches is a labelled example: replay it through the subset and record whether the subset would have caught it. A subset catching 80% of full-suite regressions is a defensible gate; one catching 40% is mis-weighted, and the fix is usually swapping in cases from the categories that keep slipping rather than simply adding more cases.
  • What levers reduce the wall clock of an eval run without cutting cases?
    Concurrency is the main one, since most of the time is network wait — run cases in parallel with bounded workers sized under the provider's rate limit so throttling does not cost more than it saves. Then share a cached common prompt prefix across cases, set reasoning effort per case rather than globally at maximum, and run deterministic graders locally instead of as additional model calls.
  • Is it acceptable that some regressions reach the main branch and are only caught overnight?
    Yes, if it is a stated trade rather than an accident. The cost is a revert or follow-up fix; the benefit is a merge loop engineers actually use and a bill that does not scale with headcount. Choose the split from measured regression frequency and revert cost. Full coverage on every pull request buys earlier detection at a price that grows with team size and a wait that eventually gets the gate disabled.

saying these in an interview costs you the question

  • Running the full suite on every pull request regardless of cost
  • Choosing the fast subset by random sampling
  • Never measuring what the fast tier misses
  • Treating eval spend as invisible infrastructure cost
  • Blocking merges on a twenty-minute gate people will skip

context