skip to content

What must a prompt regression suite pin so a score change is attributable?

level: seniorimportance: must knowfreq 55%

answer

  1. one variable, or no conclusion
  2. the ruler must not move either
  3. aliases rotate under your suite
  4. store which items flipped
  5. threshold outside measured noise

basics

~20 s

Pin everything except the prompt: the eval dataset version, the model version, decoding parameters, and — if scoring is automated by a model — the judge and its rubric. Anything unpinned becomes an alternative explanation for every score move.

solid answer

~50 s

A regression suite is a controlled experiment, and it only works if the prompt is the sole variable. That means pinning the **dataset version**, the **model version** (an alias that a provider rotates underneath you is not a pin), the **decoding parameters** including temperature and any reasoning-effort setting, the **scoring code**, and, when a model grades the outputs, the **judge model and its rubric text**. It also means the suite must store per-item outputs and verdicts, not just the aggregate, so a drop can be traced to specific items that flipped rather than argued about. Gate the build on a tolerance band derived from measured run-to-run variance rather than on any decrease, or sampling noise will red the pipeline on changes that regressed nothing. As of mid-2026 the most common untracked variable is a silently upgraded model behind a stable alias.

code

python · 18 lines
python
GOLD = [("a", "refund"), ("b", "payoff"), ("c", "escrow"), ("d", "escrow")]
HARD_CASES = {"c"}

def score(predict, dataset):
    verdicts = {item: predict(item) == gold for item, gold in dataset}
    accuracy = sum(verdicts.values()) / len(dataset)
    return accuracy, verdicts

def gate(baseline, candidate, tolerance=0.03):
    acc_b, verdicts_b = baseline
    acc_c, verdicts_c = candidate
    flipped = [k for k in verdicts_b if verdicts_b[k] != verdicts_c[k]]
    hard_broken = [k for k in flipped if k in HARD_CASES and not verdicts_c[k]]
    return acc_c >= acc_b - tolerance and not hard_broken, flipped

old = score(lambda x: {"a": "refund", "b": "payoff", "c": "escrow", "d": "escrow"}[x], GOLD)
new = score(lambda x: {"a": "refund", "b": "payoff", "c": "taxes", "d": "escrow"}[x], GOLD)
print(gate(old, new))

go deeper

for a junior

Know that a prompt change should be checked against a fixed set of test cases, and that the model, the data and the scoring must stay the same between runs for the comparison to mean anything.

for a middle

Be able to list the pinned components — dataset version, model version, decoding parameters, scorer, judge and rubric — and explain why an unpinned one becomes a competing explanation for every score movement.

for a senior

Demonstrate operating one: setting a tolerance from measured variance rather than gating on any drop, keeping incident-derived cases at zero tolerance, storing per-item verdicts so a red build is triaged in minutes, and re-baselining deliberately when a pinned component moves.

for a principal

Own the economics and the credibility: what the suite costs per run, which tier runs per commit versus nightly, who may re-baseline and with what justification, and how you keep a blocking gate trusted so teams do not route around it.

## What the suite is for A prompt regression suite runs a fixed set of scored cases against the current prompt on every change, and fails the build when quality drops beyond a threshold. Its value is not the headline number — it is **attribution**. When the score moves, you want exactly one candidate explanation: the diff someone just made. Every unpinned element of the rig is a second explanation, and a suite with three plausible explanations for a drop produces arguments rather than decisions. ## The pin list - **Dataset version.** Items added, removed or re-labelled change the score independently of the prompt. Give the set a version and record it with every result; a comparison across dataset versions is not a regression signal. - **Model version.** Provider aliases move. If your suite calls a stable-sounding alias, the model underneath can change without any commit in your repo, and the suite will attribute the resulting shift to whatever was merged that day. Pin an explicit dated or numbered version, and treat an intentional model change as its own re-baselining event with its own recorded score. - **Decoding parameters.** Temperature, top-p, max output length and any reasoning-effort or thinking-budget setting all move outputs. Fix them, and record them alongside the score. - **Scoring implementation.** Exact-match, normalisation rules, tokenisation of the comparison, how ties and refusals are counted — all of it. A change to the normalizer is a scoring change masquerading as a quality change. - **Judge model and rubric.** If outputs are graded by a model — common for tone, helpfulness or summary quality — that judge is part of the instrument. Pin its version and freeze the rubric text; a rubric edit re-scales every historical number. When you must change either, re-run the previous baseline under the new judge so you have both prompts measured on both instruments. - **Prompt identity itself.** The suite must record which prompt version produced the run, so results can be joined back to a diff. ## Nondeterminism and the gate threshold Even with temperature at zero, hosted models are not perfectly reproducible run to run, and many tasks genuinely require sampling. A gate that fails on *any* decrease will red the build constantly and be disabled within a month. The disciplined version: 1. Measure the suite's own variance first — run the unchanged prompt several times and record the spread. 2. Set the gate tolerance outside that spread, so it fires on regressions rather than on noise. 3. For a noisy suite, run each case k times and compare means, or score with a criterion that is stable across samples. 4. Complement the aggregate threshold with **hard cases that must never fail** — a handful of items encoding past incidents, gated at zero tolerance. Those catch specific regressions that a 3-point band would absorb. ## Per-item results are the deliverable Storing only the aggregate makes the suite nearly useless during an incident. A score that holds steady at 84% while eleven items broke and eleven others started passing is a real change hiding behind a stable mean. Store each item's input id, output, verdict, and ideally the judge's reasoning, and diff runs item by item. The first question after a red build is "which items flipped?", and the suite should answer it without a rerun. ## Cost, latency and where it runs A full suite on every commit can be slow and expensive, especially with a judge in the loop. The usual shape is a fast subset — a few dozen cases, cheap scoring — on every push, and the full set plus judge-scored dimensions nightly or on release candidates. Budget the suite explicitly: if it costs more per run than the team is willing to pay, it will be sampled down informally and silently, which is worse than a smaller suite run honestly. ## The failure that motivates all of this The canonical incident: a tone rubric scored by a judge model reports a four-point drop overnight. The prompt was untouched. Hours go into reviewing the previous day's merges before someone notices that the judge alias now resolves to a newer model that reads the same outputs slightly more harshly. Nothing regressed; the ruler changed. Pinning the judge, versioning the rubric, and recording both alongside every score turns that day into a one-line explanation. ## Making it a real gate The suite only changes behaviour if it blocks something. Wire it into CI so a prompt change cannot merge without a run attached, report the delta and the flipped items in the change itself, and require a re-baseline commit — with a stated reason — whenever a pinned component moves. That combination is what turns "we evaluated it" into a claim someone can check six months later.

  • The suite drops four points and nobody changed the prompt. How do you triage?
    Check the pinned components before the code. Confirm the model and judge versions resolved to the same thing as the last green run, that the dataset version is unchanged, and that decoding parameters and the scorer are identical. Then diff per-item results against the last green run: a broad, shallow shift across many items points at a model or judge change, while a cluster of related items points at data or a real behaviour change.
  • How do you handle a task where sampling is required, so runs are never identical?
    Measure the variance instead of pretending it away. Run each case several times, compare distributions or means rather than single scores, and set the gate tolerance from the observed spread. Keep a small set of deterministic hard cases at zero tolerance for the regressions that matter most, and increase repetitions only for the metrics whose noise actually threatens the decision.
  • The pinned model is being deprecated. What does migrating the suite involve?
    Treat it as re-baselining, not a routine bump. Run the current prompt on both the old and new model on the same dataset version, record both numbers, and check whether any per-item flips indicate behaviour you depend on. If a judge is involved, re-validate it against the new outputs before trusting the delta. Then commit the new baseline explicitly with the reason recorded, so future comparisons start from a documented point.
  • Should the gate block a merge or only warn?
    Block on the primary metric with a noise-aware tolerance, and block at zero tolerance on the handful of incident-derived cases. Warn on secondary slices and on cost or latency drift, where a small move rarely justifies stopping a change. A gate that fires on noise gets bypassed, so the credibility of the blocking tier matters more than its coverage.

saying these in an interview costs you the question

  • Calls a floating model alias and calls that a pinned version
  • Fails the build on any decrease, then disables the gate
  • Stores only the aggregate score, not per-item results
  • Edits the judge rubric and compares against old numbers
  • Assumes temperature 0 makes hosted runs perfectly reproducible

context