skip to content

How do you triage eval cases that keep flapping in CI at temperature 0?

level: seniorimportance: should knowfreq 45%

answer

  1. red must stay informative
  2. name the cause before you exempt the case
  3. most flaps live in the grader, not the model
  4. still runs, still reports, cannot block
  5. the exempt list is a metric with a cap

basics

~20 s

Diagnose the flap first — an ambiguous expected answer, a judge scoring near a rubric boundary, or an unpinned fixture. Fix what is fixable, move the irreducible remainder into an advisory tier that reports but cannot block, and delete cases that stabilise nothing.

solid answer

~50 s

Treat a flapping case as a bug in the case until proven otherwise. Most flaps have a nameable cause: the expected answer is genuinely ambiguous, so two correct outputs land on opposite sides of the grader; the judge is scoring right at a rubric boundary; the check is a fuzzy text match where a deterministic assertion would do; or a fixture — retrieval index, tool response, current date — was never pinned. Each of those has a fix: tighten the case, decompose the rubric into binary sub-checks, replace the grader with code, pin the fixture. What is left after that is irreducible serving nondeterminism, and the right move is **quarantine**: the case moves to an advisory tier that still runs and still reports, but cannot block a merge. Quarantine needs an owner and an expiry, and the quarantine count must be visible — a growing advisory tier is quality leaving the gate. Cases that cannot be stabilised and do not encode a real user-visible property get deleted.

go deeper

for a junior

Know that a case failing at random undermines trust in the whole gate, and that the first step is finding out why it flaps rather than re-running the build until it passes.

for a middle

Name the common causes — ambiguous expectations, a judge near a rubric boundary, a fuzzy grader where code would do, an unpinned fixture — and the matching repair for each before any exemption is considered.

for a senior

Show the full policy: diagnose, repair what is repairable, quarantine the irreducible remainder into an advisory tier with an owner and expiry, bound the re-runs, and track flap rate as a suite-level health signal.

for a principal

Own the coverage budget. Set the cap on how much of the suite may sit advisory, decide who may delete a case and on what recorded reasoning, and treat a growing advisory tier as a quality decision the team is making by default.

## Why a flapping eval case is different from a flaky test A flapping eval case is not just an annoyance; it directly attacks the gate's meaning. A gate is only useful if red means "a human should look at this". Every case that goes red at random trains the team to hit re-run, and once re-running is reflexive, a genuine regression gets re-run away too. So the goal of triage is not to make the build green — it is to restore the property that red is informative. ## Diagnose before you quarantine Quarantining first and asking questions later is how a suite hollows out. Run the flapping case many times in isolation, with the pins fixed, and capture the outputs and the grader's reasoning. Almost always the cause is one of five: **1. The case is ambiguous.** Two outputs are both defensible and the grader has to pick. "Summarise this incident report" with a reference summary is the archetype: the model's summary is fine, but overlap with the reference hovers around the pass line. Fix by making the expectation checkable — assert the required facts appear rather than scoring similarity to one blessed wording. **2. The judge sits on a rubric boundary.** A holistic 1-to-5 rubric with a pass line at 4 will flip whenever the true quality is around 3.5. Fix by decomposing into binary sub-checks that a judge answers reliably ("does it name the affected service?", "does it avoid promising a refund?") and scoring the conjunction, rather than asking for a global number near a cliff. **3. The grader is fuzzier than the property.** Embedding or overlap similarity applied to something a regex or schema check would settle exactly. Fix by moving the case down to a deterministic grader — this is the highest-value repair available, because it removes both noise and cost. **4. A fixture is not pinned.** A retrieval index refreshing on a schedule, a live tool call, a case whose answer depends on today's date. Fix by snapshotting the fixture or freezing the clock. These masquerade as model nondeterminism and are entirely under your control. **5. Irreducible serving variance.** After all of the above, some cases still flip because a near-tied token goes the other way in someone else's inference batch. Nothing in your repository fixes this. ## Quarantine as a policy, not a shrug Only category five earns a quarantine. The mechanics matter: - The case keeps running and keeps reporting. A quarantined case that stops executing is a deleted case with extra bookkeeping. - It moves to an **advisory tier** whose result is visible on the build but never blocks a merge. - It gets an **owner and an expiry date**. Unowned quarantine is permanent by default. - The **quarantine count is a tracked metric** with a cap. When the advisory tier grows, the blocking tier is quietly shrinking and the gate is covering less than the team believes. A cap forces the conversation. ## Bounded re-runs, never unbounded The tempting shortcut is "retry the case until it passes". A case that genuinely passes 60% of the time will go green within a few attempts, and so will a case that a real regression just pushed from 95% to 60%. Unbounded retry is indistinguishable from deleting the case, except that it costs money and creates the illusion of coverage. A fixed repeat count with a majority rule, decided in advance, is the honest version of the same instinct: it tolerates one unlucky sample and still fails on a case that has genuinely got worse. ## Deleting cases is allowed Suites accumulate cases nobody can stabilise and nobody can justify — written to chase a one-off complaint, encoding a preference the product no longer holds, or simply not measuring anything a user would notice. Deleting them is a legitimate, even healthy, outcome of triage. The test to apply: if this case failed permanently, would we ship anyway? If yes, it is not a gate case. Delete it in a reviewed pull request with the reasoning recorded, so it is a decision rather than an erosion. ## The signal in the aggregate A rising flap rate across many cases is not a case-level problem at all — it usually means a pin came loose, a judge model moved, or the provider changed something underneath you. Track flap rate as a suite-level health metric alongside the score. When it jumps, look at the pins before you open a triage ticket on individual cases.

  • Why not simply retry a flapping case until it passes?
    Because retry-until-green cannot distinguish a case that always passed 60% of the time from one a real regression just pushed down to 60%. Both eventually go green, so the case stops detecting anything while still costing money and implying coverage. A fixed repeat count with a majority rule, agreed in advance, gives the same tolerance for one unlucky sample without erasing the signal.
  • What stops a quarantine tier from becoming a graveyard?
    An owner, an expiry, and a cap on the count, all visible on the build report. Quarantine without those defaults to permanent, and every case that drifts into it silently shrinks the blocking suite while the team still believes the gate covers everything. A cap forces an explicit decision — fix it, delete it, or accept a smaller gate — instead of letting coverage erode by accident.
  • Several cases start flapping at once. Where do you look?
    At the suite level, not the case level. A simultaneous jump in flap rate usually means a shared input moved: a model or judge snapshot changed, a retrieval index was rebuilt, a fixture stopped being pinned, or the provider altered something upstream. Check the recorded pins against the last stable run first; opening per-case triage tickets for a shared root cause wastes days.
  • When is deleting an eval case the right answer?
    When it cannot be stabilised and it does not encode a property a user would notice. The test is simple: if this case failed permanently, would we ship anyway? If yes, it was never a gate case. Delete it in a reviewed pull request with the reasoning recorded, so the coverage decision is deliberate rather than a slow erosion nobody signed off on.

saying these in an interview costs you the question

  • Quarantining a case before diagnosing why it flaps
  • Retrying until green as the standard policy
  • Quarantine with no owner, expiry, or count cap
  • Assuming every flap is model nondeterminism
  • Blaming the model when the fixture was never pinned

context