How do you gate an LLM prompt change in CI when outputs are not exact-match?
answer
- no single right string to assert
- score, then compare to something
- the last accepted run is the reference
- fail on delta, not on wording
- cheap deterministic graders first
basics
~20 sRun a fixed set of scored eval cases as a build step and compare the aggregate score against the last accepted baseline. The build fails on a score drop beyond an agreed margin, not on any difference in wording.
solid answer
~50 sAn LLM feature has no single correct string, so the CI check is a **scored suite**, not an assertion. You keep a frozen set of inputs with graded expectations, run every case on the candidate prompt or model, and score each one — exact match or a regex where the output is structured, a programmatic check where the property is checkable (valid JSON, cites a real ticket id, under N words), and a judge model only where the property is genuinely subjective. The run produces an aggregate score plus per-case results. CI compares that aggregate to the **baseline** recorded from the last accepted change: if it drops more than the agreed margin, or if a case marked critical flips from pass to fail, the build goes red. Storing the baseline in the repo alongside the prompt is what makes 'did this change help or hurt?' answerable at review time.
go deeper
Know that LLM outputs vary, so you score a fixed set of cases and compare the total to a saved baseline instead of asserting exact text. Be able to name one deterministic check, such as schema validity.
Explain the grader tiers and why deterministic checks come first, how the baseline is stored and ratcheted in the repository, and why the per-case pass/fail diff is more useful to a reviewer than the aggregate number.
Show how you would scope the trigger, keep the run fast enough for a pull request, and read a failing diff to decide whether a drop is a real regression or an acceptable trade. Be ready to say what the suite does not prove.
Own the argument that a scored gate is the cheapest quality control an LLM project can buy, and the counter-argument that a suite nobody trusts is worse than none. Set the policy for who may move a baseline and on what evidence.
## The problem with a normal assertion Ordinary test code asserts equality: given this input, the function returns exactly this value. An LLM feature breaks that contract. Two runs of the same prompt can differ in word choice, ordering, or punctuation while both being perfectly good answers, and a prompt edit that genuinely improves quality will change almost every output string. An exact-match assertion over free text therefore fails constantly for reasons nobody cares about, and the team learns to ignore it — which is worse than having no check at all. The replacement is a **regression suite**: a fixed collection of cases, each scored on a scale, whose *aggregate* is compared to a stored baseline. ## Anatomy of a case A case is an input (plus any fixtures it needs — retrieved documents, tool responses, conversation history) and a grader. Graders come in three tiers, and you should reach for them in this order: 1. **Programmatic / deterministic.** Anything checkable by code: the output parses as JSON against the schema, it contains the order id that was in the input, it never mentions a competitor, it is under 200 tokens, the extracted date equals the expected date. These are free, instant, and perfectly stable — the backbone of any gate that has to be fast. 2. **Reference-based similarity.** A reference answer plus a text-overlap or embedding-similarity score. Cheap, but noisy and only weakly correlated with human judgement on open-ended text. 3. **Model-graded.** A judge model applies a rubric where the property is genuinely subjective ("is the tone appropriate for a customer apology?"). Powerful, but it adds cost and a second source of variance to your gate, so use it for the cases that need it rather than by default. A well-built suite is mostly tier 1 with a minority of tier 3, because every tier-3 case makes the gate slower, pricier, and noisier. ## The baseline The suite alone tells you a number; a gate needs a comparison. The baseline is the score the suite produced on the last change that was accepted into the main branch — stored in the repository next to the prompt, so it is versioned, reviewable, and moves only through a pull request. A CI run then reports three things: - the **aggregate** score on the candidate; - the **delta** against the baseline; - the **per-case diff** — which cases newly fail and which newly pass. The per-case diff is what reviewers actually read. "Overall 87.4, down 0.6" is a shrug; "three refund-policy cases that passed yesterday now hallucinate a 60-day window" is a decision. ## Where it sits in the pipeline Treat the eval run as a build step triggered by changes to anything that can move quality: the prompt templates, the model configuration, retrieval settings, tool descriptions, post-processing code. It runs like any other job, publishes its report as a build artifact, and either passes or fails the pull request. The essential discipline is that a prompt edit cannot reach the main branch without producing a number — the failure mode this whole practice exists to prevent is someone "tidying up the system prompt" on a Friday and nobody noticing that structured-output compliance fell from 99% to 91% until a customer reports it. ## Ratcheting the baseline When a change improves the score, you update the baseline in the same pull request. That turns the gate into a ratchet: quality can only be given up deliberately, in a diff someone approved, with the new number written down. If a change legitimately trades one dimension for another — better tone, slightly worse conciseness — the reviewer accepts the new baseline explicitly rather than the gate quietly absorbing the loss. ## What it does not prove A green suite says the candidate did not regress *on the cases you wrote*. It says nothing about inputs your set does not represent, and offline agreement is not a guarantee of production behaviour — real traffic drifts and users do things your fixtures never anticipated. So the suite is a floor, not a proof, and it needs to be refreshed as production reveals new failure shapes. It is still the cheapest quality signal you will ever add to an LLM project, and the one that makes every later prompt change safe to attempt.
- What do you do when a change improves the score — leave the baseline alone?No. Update the baseline in the same pull request, so the gate ratchets. Once the higher number is committed, any later change that gives that quality back has to fail the build or be accepted explicitly by a reviewer editing the baseline. Leaving the old, lower baseline in place means the improvement can be silently undone later without anyone seeing a red build.
- Which parts of the repository should trigger the eval job?Anything that can move output quality: prompt templates, model and decoding configuration, retrieval or index settings, tool and function descriptions, output post-processing, and the graders themselves. A change to unrelated application code does not need the suite. Scoping the trigger keeps spend and wall-clock down while ensuring no quality-bearing edit reaches the main branch without a number attached.
- Why prefer a programmatic check over a judge model when both could grade a case?A programmatic check is free, instantaneous, and perfectly stable, so it adds no noise and no cost to the gate. A judge model adds latency, spend, and a second source of run-to-run variance that you then have to control. Reserve model grading for properties that genuinely cannot be expressed as code, such as tone or helpfulness, and express everything else as an assertion.
saying these in an interview costs you the question
- Asserting exact output strings and calling it an LLM test
- Treating any wording change as a regression
- Running the suite without comparing to a stored baseline
- Judging every case with a model when code would do
- Never updating the baseline after a genuine improvement