Your DeepEval suite in CI goes red at random. How do you stabilise it?
answer
- separate errors from low scores
- measure the spread before moving the line
- repeats and caching do not compose
- a per-case flag that warns instead of raising
- gate on a rate, not on every case
basics
~20 sSeparate metric errors from genuine low scores with the runner's ignore-errors flag, measure each case's score variance with repeated runs, move thresholds off the noise floor, mark genuinely unstable cases flaky so they warn instead of raising, and gate on an aggregate rather than every case.
solid answer
~50 sStart by classifying the redness. A metric that *errored* — a judge timeout, a rate limit, an unparsable response — is an infrastructure failure and should not read as a quality regression; `deepeval test run -i` stops those from raising, and `-s` skips cases missing a field a metric needs. A metric that scores 0.68 one run and 0.72 the next is judge variance, and the fix is to measure it: run `-r 3` (which reruns each test; note caching is disabled whenever repeats are on) and look at the spread before choosing a threshold that sits clearly below it rather than on top of it. For cases that stay unstable, `LLMTestCase(..., flaky=True)` makes a failing metric emit a warning instead of an `AssertionError`, so it stays visible without blocking. Beyond that: pin the judge model version, keep the judge's own settings fixed, and consider gating on a pass rate computed from `evaluate()` instead of demanding every case pass.
code
python · 13 linesfrom deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
def test_known_unstable_case():
test_case = LLMTestCase(
input="Summarise the refund policy.",
actual_output="You can request a refund within 30 days.",
flaky=True,
)
# A failing metric warns instead of raising, so CI stays green
assert_test(test_case=test_case, metrics=[AnswerRelevancyMetric(threshold=0.8)])go deeper
Understand that an eval score comes from a model and will move between runs, so an eval test is not deterministic the way a unit test is. Do not expect an exact number twice.
Distinguish a metric error from a low score, and name the runner flags for each: ignore-errors for the former, threshold and repeat work for the latter. Explain why caching is off during repeats.
Show the full diagnosis-first routine: classify the redness, measure the spread, place thresholds off the noise floor, mark truly unstable cases flaky, and pin the judge so score movement means product movement.
Own the policy: which gate blocks a merge, what pass rate is acceptable, who re-baselines after a judge upgrade, and when a suite has become expensive noise that should be moved off the merge path entirely.
## Diagnose before you tune A red eval build has three quite different causes, and treating them alike is what leaves teams disabling the suite. 1. **Infrastructure noise.** The judge model call timed out, hit a rate limit, or returned something the metric could not parse. DeepEval records this as a metric *error*, and an errored metric fails the assertion by default. 2. **Judge variance.** The output did not change materially but the score moved across the threshold. This is inherent: the metric is itself a model call. 3. **A real regression.** The application genuinely got worse. Only the third should stop a merge. Everything below is about routing the first two elsewhere. ## Handling infrastructure noise `deepeval test run -i` (ignore errors) makes an errored metric non-fatal — it does *not* make a low-scoring metric pass, which is exactly the distinction you want. `-s` (skip on missing params) handles the adjacent case where a test case lacks a field a metric requires, skipping it rather than erroring. Neither is a licence to ignore the errors: if a third of your judge calls error out, you have a concurrency or quota problem, and lowering the runner's process count (`-n`) or the batch concurrency is the real fix. Because `deepeval test run` forwards unrecognised arguments to pytest, and `pytest-rerunfailures` is one of deepeval's dependencies, a transient-failure rerun can also be pushed down to pytest for the genuinely flaky-transport case. ## Handling judge variance The honest move is to measure it before you tune anything. `-r N` repeats each test N times, which turns a single score into a small distribution. Note that DeepEval deliberately disables the result cache whenever repeats are requested — a repeated test that returns a cached score measures nothing, so the two flags do not compose. With a distribution in hand, three levers apply. - **Threshold placement.** A threshold sitting inside the observed spread is a coin flip. Put it below the bottom of the range you are willing to accept, and accept that this makes the gate less sensitive. - **Judge stability.** Pin the judge to a specific model version rather than a moving alias, and keep its configuration constant, so a score change means an application change. A judge upgrade should be a deliberate, reviewed event with a re-baselining pass, not something that lands mid-sprint. - **Per-case tolerance.** `LLMTestCase(..., flaky=True)` is DeepEval's own escape hatch: a failing metric on such a case emits a Python warning (visible in pytest's warnings summary) rather than raising. It keeps a known-unstable case in the report without letting it block the pipeline, which is far better than deleting the case or quietly widening its threshold to zero. ## Choosing what the gate actually asserts Demanding that all 300 cases clear their threshold on every run is a strict gate, and strictness multiplies variance: if each case independently flips 1% of the time, a 300-case suite is red most runs. Two common alternatives: - **Aggregate gating.** Run `evaluate()` over the suite and assert on the share of successful results, so a single wobbling case cannot block a merge while a broad regression still does. - **Tiered suites.** A small, stable, high-signal set gates pull requests; the full sweep runs on a schedule and produces a report a person reads. Marks plus the runner's `-m` flag express the split. ## Making a red build diagnosable Whatever you gate on, the log has to say why. Run with a run identifier (`-id`) so a build maps to a specific reviewable run, keep `include_reason` behaviour in mind so the judge's justification reaches the failure message, and use `-d failing` to keep the end-of-run report focused. When the reason field says the metric was confused by a missing field rather than by a bad answer, you have a test bug, not a model regression. ## What not to do Do not lower thresholds until the suite goes green — that is deleting the signal and keeping the cost. Do not silently retry until a pass; a retry-until-green loop turns a quality gate into a random number generator with extra billing. Do not let a judge model float on a moving alias while treating score movement as product signal. And do not respond to flakiness by removing the suite from CI entirely and running it manually, because a suite nobody runs is worth less than a noisy one.
- Why does DeepEval disable the result cache when you pass the repeat flag?Because the point of repeating is to observe variance, and a cached result is the same number every time. The runner computes cache usage as "use cache AND no repeat requested", so `-c -r 3` silently behaves as `-r 3`. If you want both cheap reruns and variance data, do them in separate invocations.
- What exactly changes for a test case marked flaky=True?When a metric fails on that case, DeepEval emits a Python warning naming the failing metrics, their scores and thresholds, instead of raising AssertionError. The case still appears in the run report with its real score, so the signal survives while the pipeline stays green. It is a documented tolerance, not a deletion.
- A colleague proposes retrying failed eval tests until they pass. What is wrong with that?It converts a quality gate into a sampling exercise: with enough retries any case eventually clears any threshold, so the suite stops distinguishing a good build from a bad one while still paying for every judge call. Retries are legitimate for transport errors — a timeout, a 429 — but never for a score below threshold.
- How do you keep a judge-model upgrade from looking like a product regression?Pin the judge to an explicit model version so it never moves on its own, and treat changing it as its own change: re-run the suite on the old and new judge over the same cases, compare the score distributions, re-baseline thresholds if they shifted, and land that as a separate reviewed commit rather than mixed into an application change.
saying these in an interview costs you the question
- Lowers thresholds until the suite turns green
- Retries failing evals until one passes
- Treats a judge timeout as a quality regression
- Assumes -c and -r compose
- Lets the judge model float on a moving alias