skip to content

How do you turn a Ragas evaluation into a CI gate on a pull request?

level: principalimportance: should knowfreq 45%

answer

  1. the assert is the easy part
  2. one sample's judge verdict is noise
  3. compare against main, not a magic number
  4. pin every part of the instrument
  5. provider errors are not quality regressions

basics

~20 s

Run evaluate() over a small pinned dataset, read aggregate scores, and fail the job when a metric falls below its floor. The hard part is not the assert — it is choosing thresholds that catch regressions without going red at random, and keeping the run cheap enough to sit on every pull request.

solid answer

~50 s

Mechanically it is short: run `evaluate()` in the job, aggregate the per-sample table into means, and exit non-zero when a metric drops below its floor. Everything interesting is the surrounding judgement. Gate on aggregates over a curated set of a few dozen cases, never on individual samples, because a single sample's judge score is noisy. Prefer a delta against the current main baseline with a tolerance band over absolute magic numbers, so the gate tracks the suite instead of the suite being tuned to the gate. Pin the judge model snapshot and set temperature to zero, or the gate moves under you. Keep the pull-request suite small — judge calls per metric times metrics times samples is real money and real wall clock on every push — and run the full dataset nightly. Finally, separate infrastructure failure from quality failure: a provider 429 or timeout should retry or mark the job unstable, not report a quality regression that never happened.

code

python · 27 lines
python
import sys

from langchain_openai import ChatOpenAI
from ragas import EvaluationDataset, SingleTurnSample, evaluate
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import Faithfulness, LLMContextRecall

judge = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini", temperature=0))
dataset = EvaluationDataset(samples=[
    SingleTurnSample(
        user_input="Who wrote Hamlet?",
        retrieved_contexts=["Hamlet was written by William Shakespeare."],
        response="William Shakespeare.",
        reference="William Shakespeare.",
    )
])

result = evaluate(
    dataset=dataset,
    metrics=[Faithfulness(), LLMContextRecall()],
    llm=judge,
)
scores = result.to_pandas().mean(numeric_only=True)
gates = {"faithfulness": 0.85, "context_recall": 0.75}
failed = [name for name, floor in gates.items() if scores[name] < floor]
if failed:
    sys.exit(f"ragas gate failed: {failed}")

go deeper

for a junior

Know the shape: run evaluate() in the job, aggregate the scores, and exit non-zero below a threshold. Be able to say that a single sample's score is too noisy to gate on.

for a middle

Explain the mechanics end to end — dataset in the repo, aggregate the per-sample table, compare against floors — and identify the obvious cost driver: metrics times samples times judge calls per metric.

for a senior

Show you have operated one. Pin judge, version and dataset; express thresholds as a delta from main's baseline; retry provider failures and report them as unstable rather than as regressions; keep the fast suite small and the full suite nightly.

for a principal

Own the policy: which failures must never reach main, which metrics earn blocking status and on what evidence, how baselines are refreshed when the judge changes, and how the evaluation key's spend is capped and kept away from untrusted workflows.

## The mechanical part is five lines Run `evaluate()` over a dataset checked into the repo, take the per-sample table, average it, compare each metric against a floor, and exit non-zero if any fails. Nothing about that is hard, and it is not what an interviewer is probing. The question is whether you have run one of these gates long enough to know how they go wrong. ## Aggregate, never per-sample An LLM judge's verdict on any one sample carries noise. Gate on a single case and you will block a pull request because a judge changed its mind about one borderline claim. Gate on the mean over a few dozen cases and the noise mostly cancels. Per-sample results are still worth printing in the job output — they are how a developer debugs a genuine failure — but they must not be the pass/fail signal. ## Relative thresholds beat absolute ones An absolute floor like "faithfulness must exceed 0.85" is a number someone invented on a Tuesday. It will be either so loose it never fires or so tight it fires constantly, and it does not survive a dataset change. The more durable design is a baseline recorded from main plus a tolerance band: fail when the pull request is more than a few points below the recorded baseline, and refresh the baseline when main moves deliberately. That gate detects *regressions*, which is what you actually care about, and it degrades gracefully as the suite grows. ## Pin the instrument Judge model snapshot pinned, temperature zero, ragas version pinned, dataset version pinned. Any of these drifting turns the gate into a random number generator whose failures nobody can reproduce locally. When you deliberately change one — a judge upgrade especially — re-run main's baseline under the new instrument before you let the gate compare across the change, or every pull request that week will look like a regression. ## Budget the run The cost model is metrics × samples × judge-calls-per-metric, and metrics issue more than one call each. On every push, across every contributor, that adds up in both dollars and minutes. A defensible split is a small fast suite on pull requests — a few dozen carefully chosen cases, the two or three metrics you would actually block a merge on — and the full dataset with the full metric list on a nightly or pre-release schedule. Deciding what belongs in each is a product decision about which failures must never reach main, not a technical one. ## Separate infrastructure red from quality red This is the distinction that decides whether people keep the gate. Three different things can turn the job red: quality genuinely regressed, the judge provider rate-limited or timed out, or the run itself crashed. Only the first is a signal about the pull request. Provider failures should be retried and, if they persist, reported as an unstable job rather than a quality verdict; a gate that blames the developer for someone else's 429 gets disabled within a month. Ragas's run configuration gives you the concurrency and retry knobs to absorb the common transient case; the policy decision — retry, skip, or fail — is yours. ## Blocking or advisory Not every eval gate should block. A useful progression: start advisory, posting scores as a pull-request comment for a few weeks while you learn the suite's natural variance; measure how often it would have blocked and whether those blocks were right; then promote the one or two metrics with a demonstrated signal to blocking and leave the rest advisory. Promoting everything on day one produces a gate the team routes around, and a routed-around gate is worse than none because it also produces false confidence. ## The secrets and access question CI now needs a judge API key. That key can spend money proportional to how often the job runs, which means a fork-based pull request or a runaway loop is a financial event, not just a security one. Scope the key to the evaluation project, put a spend cap on it, and think carefully before exposing it to workflows triggered by untrusted contributors. ## What to say The strong answer sounds like: "A small pinned dataset, aggregate thresholds expressed as a delta from main's baseline, a pinned judge at temperature zero, advisory first and blocking only for the metrics that have proven signal, with infrastructure failures explicitly distinguished from quality failures — and the full suite nightly rather than on every push." That is a set of decisions, which is what the level is asking for.

  • Your ragas gate goes red on a pull request that only changed a README. What do you conclude?
    That the gate is measuring the instrument, not the change. Check first whether the judge model, ragas version or dataset moved, then whether the failure was a provider timeout rather than a score drop. If the scores genuinely wobbled on unchanged inputs, the thresholds are inside the suite's natural variance and need widening or the gating set needs to be larger.
  • How do you keep a per-push ragas gate affordable?
    Shrink the surface rather than the rigour: a few dozen high-signal cases and only the two or three metrics you would actually block a merge on, with a cheap pinned judge. Run the full dataset and full metric list nightly. Cost is metrics times samples times judge calls per metric, so cutting any factor cuts the bill proportionally.
  • Would you make the gate blocking from day one?
    No. Run it advisory for a few weeks, posting scores on the pull request, and measure how often it would have blocked and whether those blocks were correct. Promote to blocking only the metrics with demonstrated signal. A gate that blocks on noise gets disabled or routed around, which is worse than no gate because it also creates false confidence.

saying these in an interview costs you the question

  • Gates on individual sample scores instead of aggregates
  • Invents absolute thresholds with no baseline from main
  • Lets a floating judge alias drift under the gate
  • Reports provider timeouts as quality regressions
  • Runs the full dataset and metric list on every push

context