How do you set an eval gate threshold that catches regressions without failing on noise?
answer
- a number pulled from thin air fails twice
- measure before you decide
- compare against the last accepted run
- one sample is not a verdict
- averages hide the cases that matter
basics
~20 sMeasure the suite's run-to-run spread on unchanged code first, then set the fail margin above that band — for example baseline minus two points — and decide on repeated runs with a majority rule so a single unlucky sample cannot block a release.
solid answer
~50 sStart empirically: run the unchanged suite several times and record the spread of the aggregate. That band is your noise floor, and any threshold tighter than it guarantees false reds. Then set the rule as a **delta against the baseline**, not an absolute score — say, fail if the aggregate falls more than two points below the last accepted run — and evaluate it over three runs with a majority rule, so one unlucky sample cannot stop a merge. Aggregate alone is not enough, though: a suite can hold its average while a small critical slice collapses. So layer on per-slice guards — any case tagged safety-critical or contractual flips from pass to fail and the build goes red regardless of the total. Publish both numbers in the run report. And re-derive the noise band whenever the model, judge, or case set changes, because the band moves with them.
code
python · 7 linesdef gate(run_scores, baseline, margin=2.0):
"""Fail only if a majority of runs drop more than `margin` below baseline."""
breaches = sum(1 for s in run_scores if baseline - s > margin)
return "FAIL" if breaches * 2 > len(run_scores) else "PASS"
print(gate([86.0, 88.5, 87.0], 89.0)) # PASS - one unlucky sample
print(gate([84.0, 85.5, 86.0], 89.0)) # FAIL - real regressiongo deeper
Know that the gate compares the run against a stored baseline and fails on a drop bigger than an agreed margin, and that a single noisy run is not enough to declare a regression.
Explain how the margin is derived from a measured noise band, why the rule is a delta rather than an absolute score, and how a majority-of-three decision rule cuts false reds without hiding real drops.
Show the layered rule in practice: aggregate delta plus critical-case guards plus per-slice reporting, with a fixed repeat count and no unbounded retries. Be ready to diagnose whether a red build is noise, a bad margin, or a real regression.
Own the false-red versus missed-regression trade for the organisation: validate thresholds by replaying history, set who may widen a margin, and treat repeated widening as evidence the suite needs repair rather than tuning.
## Why a threshold is a design decision, not a constant An eval gate answers one question: is this change bad enough to stop? Set the bar too tight and every build is red for reasons nobody can act on, the team starts merging past it, and the gate is dead. Set it too loose and quality erodes a fraction of a point per pull request until, forty merges later, the feature is visibly worse and no single commit is to blame. Threshold design is the discipline that keeps a gate on the useful side of that line. ## Step one: measure the noise, do not guess it Before any threshold exists, run the suite several times against **unchanged** code and pins. The spread of the aggregate across those runs is your noise band — the score movement that means nothing. It comes from the residual serving nondeterminism that survives greedy decoding, plus any judge-model variance. A threshold set inside that band is a coin flip dressed up as a quality signal. Re-measure the band whenever anything structural changes: a new model snapshot, a new judge, a materially different case set. A band measured six months ago against a retired model is not evidence about today. ## Step two: delta, not absolute Gate on the *change* relative to the last accepted baseline rather than a fixed target score. Absolute targets rot: an 85-point bar is meaningless once the suite is at 92, and it silently permits a seven-point collapse. A delta rule — "fail if the aggregate is more than two points below the baseline" — asks the only question a reviewer cares about, which is whether this diff made things worse. Pair it with baseline ratcheting: when a change improves the score, the pull request updates the baseline, so improvements are locked in and cannot be quietly given back. ## Step three: a decision rule over repeated runs A single run is one sample from a noisy distribution. Repeating the suite and taking a majority verdict — fail only if most of, say, three runs breach the margin — collapses the false-red rate dramatically while still catching any regression large enough to matter, because a real regression breaches the margin in every run, not one. Two things make this honest rather than a fudge: the repeat count is fixed in advance, and re-running on failure is not permitted outside it. An unbounded retry loop will eventually make any regression green, which is how a gate becomes theatre. ## Step four: guard the slices the aggregate hides Aggregates average away the failures you least want to average away. A 300-case suite can lose every one of its eight refund-policy cases and move the total by under three points. So the rule is layered: - **aggregate delta** against the baseline, with the noise-derived margin — catches broad quality drift; - **critical-case guards** — a named subset (safety, legal, contractual, schema compliance) where any pass-to-fail flip fails the build on its own, no margin, no averaging; - **slice deltas** — per-segment scores reported for the largest slices, so a reviewer can see a single segment collapsing even when the total looks fine. The critical set should stay small enough that every member is genuinely non-negotiable. If everything is critical, nothing is. ## Step five: report enough to act on A red build must say which rule fired, by how much, with the per-case diff of newly failing cases attached. "Aggregate 84.9 vs baseline 89.0, margin 2.0, failing in 3 of 3 runs; new failures concentrated in the refund-policy slice" is a build a reviewer can act on in a minute. "Eval gate failed" is a build they will re-run. ## Tuning by looking backwards Once there is history, tune the threshold against it. Replay past pull requests through the current rule: how many known-good changes would it have blocked (false reds), and how many changes later found to be regressions would it have let through (false greens)? That converts an argument about a number into an argument about consequences. Most teams discover their instinct was too tight, and that widening the margin while adding critical-case guards catches strictly more of what matters with far fewer wasted builds. ## What the threshold cannot do A threshold governs a fixed case set. It cannot see regressions on inputs nobody wrote a case for, and a suite whose margin has been widened repeatedly to keep builds green is telling you the cases are unstable or the quality is genuinely sliding. Treat repeated margin-widening as a signal to fix the suite, not as tuning.
- Why gate on a delta from the baseline instead of a fixed minimum score?A fixed minimum rots in both directions. Once the suite scores well above the bar, the bar permits a large collapse before it fires; and when a legitimately harder case set lowers the ceiling, the bar blocks everything. A delta asks the question that matters for a pull request — did this change make things worse than what we already accepted — and pairs naturally with ratcheting the baseline upward on improvements.
- Your aggregate is within margin but three safety-tagged cases went from pass to fail. Should the build go red?Yes. That is exactly what critical-case guards exist for: a small named subset where any pass-to-fail flip fails the build on its own, regardless of the aggregate. Averages are designed to absorb small movements, and a handful of safety or contractual failures is small in aggregate terms and unacceptable in product terms. Keep the critical set small enough that every member truly is non-negotiable.
- How do you tell a badly chosen threshold from genuinely unstable cases?Look at where the variance lives. If the run-to-run spread is spread thinly across many cases, the margin is simply too tight for the suite's noise floor and should be widened. If it concentrates in a handful of cases that flip while everything else is rock-steady, the threshold is fine and those specific cases need triage — tightening, replacing with a deterministic check, or quarantining out of the blocking tier.
- How would you validate the threshold you picked?Replay history. Run past pull requests through the proposed rule and count how many known-good changes it would have blocked and how many later-confirmed regressions it would have let through. That turns the choice from an argument about a number into an argument about false reds versus missed regressions, and it usually shows the first instinct was too tight.
saying these in an interview costs you the question
- Picking a round threshold with no measurement behind it
- Gating on a fixed absolute score instead of a delta
- Deciding on a single run of the suite
- Allowing unlimited re-runs until the build goes green
- Trusting the aggregate while a critical slice collapses
- Widening the margin every time the build goes red