skip to content

How do you validate an LLM judge against human labels before trusting its scores?

level: seniorimportance: must knowfreq 55%

answer

  1. compare the instrument to the standard
  2. raw percent agreement misleads on skewed sets
  3. chance-corrected coefficient
  4. Cohen's kappa on a labelled sample
  5. disagreements usually expose rubric ambiguity

basics

~20 s

Have humans label a sample against the same rubric, run the judge on that sample, and measure chance-corrected agreement such as Cohen's kappa rather than raw percent agreement. Below roughly moderate agreement the judge is not usable, and the usual cause is an ambiguous rubric, not a weak model.

solid answer

~50 s

Validation means treating the judge as an instrument and measuring it against the standard it is meant to replace. Humans label a sample using the exact rubric text the judge will see; the judge then scores the same items blind to those labels; you compare. Use a **chance-corrected** statistic — Cohen's kappa for two raters, Krippendorff's alpha for more or for ordinal scales — because raw percent agreement is inflated whenever one label dominates: on a set that is 90% pass, a judge that always says pass scores 90% agreement and zero kappa. Interpretation is blunt: around 0.4 is moderate and not shippable, above roughly 0.75-0.8 is usable, and human-to-human agreement is your realistic ceiling — a judge cannot beat the consistency of the raters defining truth. When agreement is low, read the disagreements. Nearly always they cluster on cases the rubric never resolved, and rewriting the rubric into decomposed binary criteria moves kappa far more than swapping in a bigger judge model.

code

python · 14 lines
python
def cohens_kappa(judge_labels, human_labels):
    n = len(judge_labels)
    observed = sum(j == h for j, h in zip(judge_labels, human_labels)) / n
    categories = set(judge_labels) | set(human_labels)
    expected = sum(
        (judge_labels.count(c) / n) * (human_labels.count(c) / n)
        for c in categories
    )
    return (observed - expected) / (1 - expected)

# 90% of items are "pass"; a judge that always says pass:
human = ["pass"] * 90 + ["fail"] * 10
judge = ["pass"] * 100
print(cohens_kappa(judge, human))  # 0.0 despite 90% raw agreement

go deeper

for a junior

Know that a judge must be checked against human labels on a sample before its numbers mean anything, and that the comparison uses an agreement statistic rather than accuracy against a truth.

for a middle

Explain why chance-corrected agreement is needed: on a 90/10 set a judge that always says pass hits 90% raw agreement with zero information. Be able to state roughly what kappa values are usable.

for a senior

Run the whole loop and show judgement in it: freeze the rubric for both sides, measure human-human agreement as your ceiling, read disagreements individually, and fix the rubric rather than upgrading the model. Say when agreement expires and must be re-measured.

for a principal

Own the standard itself. If your raters disagree with each other, the construct is underspecified and no judge will rescue it — the work is defining what quality means before automating its measurement, and deciding which decisions may rest on judge evidence at what agreement level.

## Validation is a measurement problem, not a modelling problem The question a judge must answer is not "is this output good in some absolute sense" but "would our raters have called this output good". That reframing makes validation tractable: you have a standard (your human raters), you have an instrument (judge model plus rubric plus prompt), and you want to know how closely the instrument reproduces the standard. ## The procedure 1. **Freeze the rubric.** Humans and judge must grade against identical criteria text. If your raters are working from tribal knowledge and the judge from a paragraph, you are measuring the gap between two rubrics, not the judge. 2. **Label a sample by hand.** Use more than one rater on at least an overlapping subset, so you can also compute human-to-human agreement. 3. **Run the judge blind** on the same items — no access to human labels, no few-shot examples drawn from this sample, or you have contaminated the estimate. 4. **Compute chance-corrected agreement.** 5. **Read the disagreements individually.** This is the step teams skip and the step that produces the fix. ## Why chance-corrected, and what kappa is Cohen's kappa compares observed agreement against the agreement you would expect if both raters assigned labels independently at the same base rates: kappa = (p_observed - p_expected) / (1 - p_expected). A kappa of 0 means the judge is doing no better than base-rate guessing; 1 means perfect agreement. The reason this matters is class imbalance, which is the normal state of an eval set. If 90% of outputs pass, a judge that emits pass unconditionally hits 90% raw agreement — a number that looks like success on a dashboard and carries no information. Its kappa is 0. Any time someone quotes raw agreement on an imbalanced set, ask for kappa. Rough reading of the scale: below 0.2 is negligible, 0.2-0.4 fair, 0.4-0.6 moderate, 0.6-0.8 substantial, above 0.8 near-ceiling. For ordinal scales use weighted kappa or Krippendorff's alpha so that a judge saying 4 where a human said 5 is penalised less than one saying 1. Mid-2026 practice on decomposed binary criteria targets agreement above roughly 0.85 before a judge is used unsupervised; softer holistic rubrics rarely reach it. ## The human ceiling Measure agreement between your own raters first. If two experienced teachers agree only 70% of the time on whether essay feedback is "specific and encouraging", no judge will exceed that, and a judge reporting 0.9 agreement with one of them should be regarded with suspicion rather than delight. Low human-human agreement is itself the finding: it means the construct is underspecified and the rubric needs work before any automation is worth building. ## Diagnosing a low score A worked shape of this: a judge grading middle-school essay feedback against a four-level rubric scores kappa 0.41 against two teacher raters. Moderate — not usable. Reading the 40-odd disagreements shows they concentrate in one place: the rubric says feedback should be "specific", and the judge counts a reference to "your second paragraph" as specific while the teachers require a quoted phrase and a concrete revision. The rubric never said which. Rewriting the criterion into three binary checks — quotes at least one phrase from the submission; names one concrete revision; contains at least one affirming statement about the work — moves kappa to 0.72 with the same judge model. Nothing about the model changed; the ambiguity that both sides were resolving differently was removed. That is the general pattern. Ranked by how often they are the true cause: rubric ambiguity, scale granularity (five-point scales invite disagreement that three-point ones do not), missing context in the judge prompt, then finally judge capability. ## Keeping validation alive Agreement is a property of a specific judge-model version, rubric version and item distribution, so it expires. Re-measure when you change the judge model or its version, when you edit the rubric at all, and when the population of outputs shifts — a prompt change that alters output style can move the judge's behaviour without touching the rubric. Keep a small frozen human-labelled set specifically for this re-check so the comparison across time is apples to apples. ## What to report A credible validation report states: number of items, number of human raters, human-human agreement, judge-human agreement with the statistic named, the judge model and version, the rubric version, and a couple of representative disagreements. That is what makes a judge number something a reviewer can weigh rather than take on faith.

  • Your judge reaches kappa 0.41. What do you try before reaching for a stronger judge model?
    Read every disagreement first — they usually cluster on one criterion the rubric left ambiguous. Then, in order: rewrite that criterion into concrete binary checks, coarsen an over-granular scale, and add whatever context the human raters had that the judge prompt lacks (the source document, the user's original request). Model capability is the last hypothesis because it is the one that costs money and usually is not the cause.
  • Your judge agrees with human raters more often than they agree with each other. How do you read that?
    As a warning, not a win. Human-human agreement is the ceiling for a well-defined construct, so exceeding it usually means the judge has latched onto a surface regularity that one rater happens to follow — often length or formatting — or that the labelled sample leaked into the judge prompt as examples. Check for contamination, then check whether the judge's agreement holds against each rater separately.
  • When does judge-human agreement need re-measuring?
    Whenever any part of the instrument or its input distribution changes: a new judge model or model version, any edit to the rubric text, a change to the judge prompt scaffolding, or a shift in the outputs being graded — a prompt change that makes answers longer or more structured can move judge behaviour without the rubric changing at all. Keep a frozen human-labelled set for exactly this re-check.
  • Would you use weighted kappa instead of plain Cohen's kappa? When?
    Yes, whenever the scale is ordinal. On a four-level rubric, a judge saying 3 where a human said 4 is a much smaller error than 1 versus 4, and plain kappa treats both as simple disagreement. Weighted kappa (or Krippendorff's alpha, which also handles more than two raters and missing labels) credits near-misses proportionally and gives a fairer picture of a graded scale.

saying these in an interview costs you the question

  • Reporting raw percent agreement on a heavily imbalanced eval set
  • Assuming a low kappa means the judge model is too weak
  • Validating once and never re-checking after a judge upgrade
  • Ignoring human-to-human agreement as the achievable ceiling
  • Drawing the judge's few-shot examples from the same labelled sample

context