skip to content

How do you keep LLM-judge scores comparable across judge upgrades and months?

level: principalimportance: should knowfreq 38%

answer

  1. the judge is a versioned instrument
  2. pin the model, hash the rubric
  3. temperature 0 is not determinism
  4. dual-run old and new over an overlap
  5. re-validate agreement after any change

basics

~20 s

Treat the judge as a versioned instrument: pin the judge model version and rubric text, record both with every score, and never compare numbers across a change without a bridge. When either changes, re-run a frozen human-labelled set and dual-run old and new judges over an overlap to measure the offset.

solid answer

~60 s

A judge score is only comparable to another score produced by the same instrument, and the instrument is the judge model version plus the rubric text plus the prompt scaffolding. So the operating discipline is version control for all three: pin the judge model to an explicit version rather than a floating alias, hash the rubric, and stamp every stored score with both plus the temperature. Temperature 0 reduces variance but is not determinism — serving-side batching and routing mean identical prompts can still yield different verdicts, so quantify residual variance by re-running a sample rather than assuming it away. When you must change the judge, run old and new over an overlap window, measure the offset and the rank correlation, and re-check agreement against a frozen human-labelled set before switching; a silent judge upgrade will move every dashboard and look like a product regression. Cost is the other binding constraint: judging long outputs with a reasoning-capable model can rival or exceed generation cost, so sample rather than judge everything, use cheap deterministic checks for what does not need semantics, and reserve the expensive judge for the criteria that do.

go deeper

for a junior

Know that judge scores are only comparable if the judge model and the rubric stayed the same, and that both should be recorded alongside every score.

for a middle

Explain what makes up the instrument — judge model version, rubric text, prompt scaffolding — and why temperature 0 reduces variance without guaranteeing identical verdicts across runs.

for a senior

Describe the migration mechanics: re-validate agreement on a frozen labelled set, dual-run old and new judges over an overlap to get an offset and rank correlation, annotate the seam, re-measure bias diagnostics. Put real numbers on the judge cost budget.

for a principal

Own the instrument across the organisation — one rubric registry, one pinned judge configuration, one calibration set, and a rule that no quality claim is admissible without its provenance. Be willing to say a migration broke the historical series rather than papering over it with an offset.

## The judge is infrastructure, not a script Once judge numbers are used to make decisions, they become a time series that people reason about across months. That immediately raises the question every metrology discipline has faced: is today's reading comparable to last quarter's? For a judge, the answer is yes only if nothing about the instrument changed — and there are four things that silently can. ## The four things that move a judge score without quality moving **The judge model version.** Providers ship updates behind stable-sounding aliases. If your evaluation points at a floating alias, your baseline moves underneath you and you cannot tell a product regression from a judge change. Pin an explicit version, and treat a version change as a scheduled migration, not a background event. **The rubric text.** Any edit — even clarifying a sentence — is a new instrument. Version the rubric alongside code and store its hash with each score. **The prompt scaffolding around the rubric.** Output format instructions, few-shot examples, whether the judge is asked to reason before verdicting: all of it shifts behaviour. It is part of the instrument. **The distribution of what is being judged.** If your generator starts producing longer or more structured outputs, judge behaviour shifts even with the instrument frozen, because of verbosity and formatting sensitivity. This one is not a versioning problem, but it is the reason a moving score must be diagnosed rather than believed. ## Determinism, honestly Temperature 0 is the standard setting and it does substantially reduce run-to-run variance, so use it. But it is not a determinism guarantee. Serving-side batching, floating-point non-associativity across differing batch compositions and expert-routing effects mean the same prompt can produce different verdicts on different runs. The correct posture is empirical: re-run a fixed sample several times, measure how many verdicts move, and report that residual as the noise floor of your instrument. A reported difference smaller than the noise floor is not a difference. Self-consistency — sampling the judge several times and taking a majority — trades cost for stability and is worth it on high-stakes verdicts. It does not remove bias; it removes some variance. ## Migrating the judge without breaking the series When the judge model must change — deprecation, cost, a genuinely better model — do it as an instrument migration: 1. **Re-validate against the frozen human-labelled set.** Agreement is a property of the pairing, so the new judge's agreement must be measured, not inherited. A better model with worse agreement on your rubric is a real and common outcome. 2. **Dual-run over an overlap window.** Score the same items with both judges. Compute the mean offset and, more importantly, the rank correlation — if the two judges order candidates the same way, historical comparisons survive with an offset annotation; if they reorder, your history is not translatable and you should say so. 3. **Annotate the seam.** Every dashboard and report that spans the change carries a marker. Nothing is worse than a step change in a quality metric that no one can attribute. 4. **Re-measure the bias diagnostics** — swap-flip rate, score-versus-length relationship — because a new judge has a new bias profile. ## Cost, which is the constraint people underestimate Judging is not free and is often not cheap. A judge that reads a long output, is given the source material for grounding, and reasons before verdicting can consume more tokens per item than the generation being graded — and if you judge on several criteria separately, multiply by that. Levers, in rough order of leverage: judge a stratified sample rather than the whole population; push everything expressible as an assertion into deterministic checks; use a small model for the easy binary criteria and a strong one only for the criteria that need judgement; cache aggressively where the rubric prefix is shared across items; and drop judge reasoning depth to the minimum that preserves agreement — measure this, do not guess it. The honest framing for an interview: a judge that costs more than the human review it replaced is not automation, it is a more expensive process with a worse standard. ## Governance in the wider sense At organisational scale the failure is not technical but social — several teams each running their own rubric, judge and baseline, then arguing about whose number is right. The fixes are boring and effective: one owned rubric registry with versions, one pinned judge configuration per criterion family, one frozen human-labelled calibration set that everyone re-validates against, and a rule that no quality claim is admissible without naming the judge version, rubric version and agreement figure it rests on. ## What a strong answer sounds like Name the instrument explicitly, describe pinning and stamping, be honest that temperature 0 is variance reduction rather than determinism, describe the dual-run bridge and re-validation on judge change, and put a number-shaped constraint on cost. Admit that some judge migrations simply break the historical series, and that saying so is better than pretending an offset fixes it.

  • Your judge model is deprecated and you must migrate. What is the migration plan?
    Re-validate the new judge against the frozen human-labelled set first — agreement belongs to the pairing, not the model. Then dual-run both judges over an overlap window and compute the mean offset and the rank correlation: matching order means history survives with an annotation, reordering means the series breaks and you say so. Re-measure the bias diagnostics for the new judge, then mark the seam on every dashboard spanning it.
  • If temperature 0 is not deterministic, how do you know whether a score change is real?
    Establish the noise floor empirically. Re-run a fixed sample through the unchanged instrument several times and measure how many verdicts move; that spread is the smallest difference the instrument can resolve. Report changes against it, and for high-stakes verdicts use self-consistency — sample the judge several times and take the majority — to buy stability with tokens. Anything inside the noise floor is not a finding.
  • How would you cut judge cost by an order of magnitude without losing the signal?
    Stratified sampling instead of judging the whole population, which usually recovers most of the statistical power at a fraction of the volume. Push every criterion expressible as an assertion — schema validity, length, forbidden strings, tool-call correctness — out of the judge entirely. Use a small model for the easy binary criteria and reserve a strong judge for the ones that genuinely need reading. Then measure agreement again, because each of these changes the instrument.
  • Several teams each run their own judge and rubric and disagree about quality. What do you fix?
    Not the models — the governance. One versioned rubric registry, one pinned judge configuration per criterion family, one frozen human-labelled calibration set everyone validates against, and a standing rule that no quality claim is admissible without naming the judge version, rubric version and agreement figure behind it. The disagreement is usually three different instruments reporting three different quantities, all called quality.

saying these in an interview costs you the question

  • Pointing evaluation at a floating model alias instead of a pinned version
  • Assuming temperature 0 makes judge verdicts fully deterministic
  • Swapping in a newer judge and comparing the scores to historical numbers
  • Editing rubric wording without treating it as a new instrument version
  • Judging every production item when a stratified sample carries the same signal

context