skip to content

What is LLM-as-judge evaluation, and why is a judge score not ground truth?

level: juniorimportance: must knowfreq 70%

answer

  1. a model grading another model
  2. instrument, not ground truth
  3. errors are correlated, not random
  4. judge shares the generator's blind spots
  5. validate against human labels first

basics

~20 s

LLM-as-judge means prompting a model with a rubric to score another model's output. The score is a noisy estimate, not truth: the judge has its own biases and blind spots, so it must be validated against human labels before anyone trusts the number.

solid answer

~50 s

An LLM-as-judge is a model you prompt with the task, the candidate output and a rubric, and ask for a verdict — a score, a label, or a preference between two candidates. Teams reach for it because most generative tasks have no single correct string, so exact-match and n-gram overlap metrics score good answers as failures. The catch is that the judge is a **measuring instrument built from the same technology being measured**. Its errors are systematic rather than random: it prefers longer answers, it is swayed by the order candidates appear in, and it can favour text that looks like its own. That means a judge score is only meaningful once you have shown it agrees with human raters on a labelled sample, and only for the rubric it was validated on. Treat a judge number the way you would treat a sensor reading from an uncalibrated sensor.

go deeper

for a junior

Be able to say plainly what an LLM judge is — a model prompted with a rubric to grade another model's output — and why teams use it where exact-match metrics fail. Add that the score needs checking against human labels before it is trusted.

for a middle

Explain why judge errors are systematic rather than random, so averaging does not wash them out, and why the rubric carries most of the quality. Interviewers expect you to describe the concrete shape of a judge prompt.

for a senior

Show that you treat the judge as an instrument with a validity claim attached to a specific rubric and task distribution. Talk about the hybrid pipeline — deterministic checks, then judge, then human review of what it flags — and about reporting scores with their provenance.

for a principal

Own the framing that a judge metric becomes the team's optimisation target, so a biased judge quietly steers the product. Be ready to argue about which decisions may rest on judge numbers at all and which still require human adjudication.

## The problem judges exist to solve Most things an LLM produces have no single correct answer. A summary of an incident report, feedback on a student essay, a rewritten support reply — all have many acceptable forms. Metrics that compare a candidate string against a reference string reward surface overlap, so they punish a good answer that used different words and reward a bad answer that echoed the reference's vocabulary. Human review does not have that problem, but it costs money and time and cannot run on every commit. LLM-as-judge sits between the two. You give a model the input, the candidate output, and a written rubric, and ask it to return a verdict: a numeric score, a set of pass/fail criteria, or a preference between two candidates. It runs in seconds, costs cents, and scales to thousands of items. ## The shape of a judge call A judge prompt has four parts: the original task or question, the candidate output being graded, the rubric that defines what good means, and a required output format. The rubric is the part that carries almost all of the quality. A rubric that says "rate helpfulness 1-5" leaves the definition of helpfulness entirely to the judge model's priors, and the score will move when you swap judge models. A rubric that decomposes into concrete binary criteria — "does the feedback name at least one specific sentence from the essay?", "does it avoid a bare negative judgement with no suggestion?" — pins down what is being measured and produces far more stable verdicts. ## Why the output is not ground truth Three properties separate a judge score from a true label. **The errors are correlated, not random.** If a judge is noisy in an unbiased way, averaging over many items recovers the true mean. Judge errors are not like that: they lean consistently in one direction. A judge that mildly prefers longer answers will report that the verbose variant of your prompt is better on every item in the set, and the average will confidently report a nonexistent improvement. **The judge shares failure modes with the generator.** A judge from the same model family often shares the same factual gaps, so a plausible-sounding wrong claim can pass unremarked. This is why a judge is a weak detector of subtle factual error unless it is given the reference material to check against. **A verdict is only valid for the rubric it was validated on.** Agreement measured on a rubric for essay feedback says nothing about the same judge scoring code review comments. Validity is a property of the judge-plus-rubric-plus-task-distribution, not of the model. ## What makes a judge usable anyway A judge does not need to be right in an absolute sense; it needs to agree with the humans whose judgement you would otherwise be buying. So the standard workflow is: write the rubric, have humans label a modest sample, run the judge over that same sample, and measure agreement. If agreement is high enough, you can run the judge at scale and treat its aggregate as a proxy for what your raters would have said. If it is not, the fix is almost always a sharper rubric rather than a bigger judge model. As of mid-2026 the mature production shape is a hybrid pipeline: cheap deterministic checks first (does it parse, is it within length, does it contain a forbidden phrase), the judge for the parts that need semantic understanding, and human review of the small fraction the judge flags as borderline or fails confidently. The judge is used to *route human attention*, not to replace it. ## How to talk about judge numbers honestly Report judge scores with their provenance: which judge model and version, which rubric version, and what agreement with humans that pairing achieved. A statement like "quality went from 3.9 to 4.2" is close to meaningless on its own — the scale is arbitrary, the judge might have been upgraded underneath, and the difference may be inside the noise. A statement like "on our 200-item labelled set, the new prompt passes the specificity criterion on 74% of items versus 61%, with the judge agreeing with our raters on 88% of those labels" is a claim someone can act on. ## The common mistake The failure that actually happens in teams is not using a judge — it is forgetting the calibration step, shipping the judge as if it were an oracle, and then discovering months later that the metric everyone optimised was measuring answer length.

  • When would you not reach for a judge at all?
    When a cheaper deterministic check answers the same question. If the criterion is "valid JSON matching this schema", "under 200 words", "contains no customer email address", or "the generated SQL returns the expected rows", write the assertion. A judge costs tokens, adds variance and needs validating; spend it only on criteria that genuinely require reading for meaning.
  • Why does an LLM judge struggle to catch subtle factual errors?
    Because it is checking a claim against its own parametric knowledge, which shares gaps with the generator's. A confident, well-formed false statement looks like a good answer on every stylistic dimension the judge can see. The mitigation is to give the judge the source material and ask a narrower question — does every claim in this output appear in the provided text — rather than asking it whether the output is true.
  • Does using a stronger judge model remove the need for validation?
    No. A stronger judge usually agrees with humans more often, but the biases are structural rather than capability-limited: presentation order, verbosity and formatting still move its verdicts. And validation measures the judge against *your* raters' standard, which no model knows a priori. A stronger model raises the ceiling; it does not tell you where you are under it.

A judge model is an uncalibrated sensor. It produces a confident reading every time, but until you have compared it against a known reference you have no idea what the reading means.

saying these in an interview costs you the question

  • Treating the judge score as an objective measure of quality
  • Believing averaging over many items cancels judge bias
  • Assuming a judge reliably detects factual errors without source material
  • Reusing a validated rubric on a different task and assuming it transfers
  • Reporting a score without naming the judge model or rubric version

context