skip to content

Which biases distort LLM-as-judge scores, and how do you control each one?

level: middleimportance: must knowfreq 68%

answer

  1. the judge has preferences of its own
  2. order of presentation moves verdicts
  3. longer is not better
  4. same family, kinder scores
  5. swap, decompose, cross-family

basics

~20 s

The recurring four are position bias (order of presentation sways pairwise verdicts), verbosity bias (longer wins), self-preference (a judge rates its own family higher), and formatting halo (bullets and confident tone read as quality). Controls: swap orders, normalise or cap length, judge across families, and strip presentation from the rubric.

solid answer

~50 s

Four biases show up in essentially every judge deployment. **Position bias**: in pairwise judging, the candidate shown first is favoured — run each pair in both orders and count only verdicts that survive the swap; the flip rate is itself a bias metric. **Verbosity bias**: a 300-word critique beats a sharper 80-word one on the same rubric — control it by scoring decomposed binary criteria instead of an overall impression, stating explicitly that length is not a criterion, and comparing scores against a length-matched baseline. **Self-preference**: a judge from the same model family as the generator tends to rate its own outputs higher, so use a judge from a different family, or an ensemble, when comparing two systems. **Formatting halo**: headings, bullets and confident register inflate scores independently of content — a rubric that names content criteria explicitly, and spot checks on plain-text-stripped outputs, keep it honest. None of these are removed by a stronger judge; they are structural, so you measure them rather than assume them away.

code

python · 20 lines
python
RUBRIC = (
    "Which response better satisfies the rubric? "
    "Reply with exactly one word: first, second, or tie."
)

def build_prompt(question, first, second):
    return (
        f"{RUBRIC}\n\nTask: {question}\n\n"
        f"[first]\n{first}\n\n[second]\n{second}"
    )

def pairwise_verdict(judge, question, a, b):
    """Only a verdict that survives the swap counts as signal."""
    v1 = judge(build_prompt(question, a, b))
    v2 = judge(build_prompt(question, b, a))
    if v1 == "first" and v2 == "second":
        return "A"
    if v1 == "second" and v2 == "first":
        return "B"
    return "tie"

go deeper

for a junior

Know the names and the direction of the four biases: first position wins, longer wins, same family wins, prettier formatting wins. Being able to list them and say why they matter is already a good junior answer.

for a middle

Pair each bias with a concrete control — swap the order and count only consistent verdicts, decompose the rubric into binary criteria, judge across model families, name content criteria explicitly. Explain why these are structural rather than capability limits.

for a senior

Show that you measure the biases rather than assume the controls worked: flip rate under swap, self-versus-self ties, score-versus-length relationship. Be able to say what a 22% flip rate means for the effect size you were about to report.

for a principal

Own the second-order risk: a biased judge becomes the team's optimisation target and quietly reshapes the product toward long, formatted, confident output. Decide which comparisons are admissible on judge evidence at all, especially cross-provider ones.

## Why bias is the central topic here A judge is only useful if its verdicts track the thing you care about. Every bias in this list is a case where the verdict tracks something else — presentation order, output length, family resemblance, typography — while looking exactly like a quality signal. Because the errors are consistent rather than random, they do not average out, and because the judge writes a plausible rationale for every verdict, they are invisible in the output. You find them by designing experiments, not by reading judge explanations. ## Position bias In pairwise judging, where the judge sees two candidates and picks one, the position a candidate occupies changes its odds. A judge with strong position bias will pick the first-presented candidate substantially more often than chance even when the two candidates are equivalent. The diagnostic is the **swap test**: run every pair twice, once as (A, B) and once as (B, A), and count how often the verdict flips. A run where swapping the order flips 22% of verdicts is telling you that roughly a fifth of your comparisons carry no signal at all. The control follows directly: score both orders and accept a winner only when both runs agree; everything else is recorded as a tie. That halves your judge calls and shrinks your apparent effect size, which is the honest outcome — the effect was never there. A useful secondary check: run a candidate against a copy of itself. Any systematic winner is pure position bias, with the true answer known to be a tie. ## Verbosity bias Judges reward length. Given the same rubric, a 300-word critique that restates the question, lists three considerations and adds a caveat tends to beat an 80-word critique that says the sharp thing once. Length correlates with thoroughness often enough in training data that the judge has learned it as a proxy, and it applies that proxy even when your rubric explicitly values concision. Controls, roughly in order of effectiveness: - **Decompose the rubric into binary criteria.** "Does it cite a specific line from the submission?" is hard to satisfy by padding; "rate the feedback 1-5 for quality" is easy to satisfy by padding. - **Say length is not a criterion**, and where concision is genuinely valued, make it its own scored criterion so it is measured rather than assumed. - **Length-normalise or length-match.** Compare candidate scores against a baseline of similar length, or check whether the score difference between two variants survives when you regress out length. If your "improved" prompt only wins because it produces 2.4x more text, you want to know. - **Cap or report length alongside score** so a reviewer can see the confound. ## Self-preference A judge tends to rate outputs from its own model family more favourably than outputs from a rival family, on the order of a fraction of a point on a five-point scale. The mechanism is plausibly familiarity with its own stylistic distribution rather than any intent, but the effect is enough to reverse a close comparison. This matters most in exactly the situation teams use judges for: deciding whether to switch model providers. If you judge candidate model X against incumbent model Y using a judge from X's family, the result is not admissible. Controls: pick a judge from a third family, run a panel of judges from different families and require agreement, or validate the judge on a human-labelled subset that includes outputs from both systems and check the agreement is comparable for both. ## Formatting halo Headings, bullet lists, bold text and a confident, hedge-free register raise scores independently of whether the content improved. This is closely related to verbosity bias but distinct: a well-formatted short answer also benefits. It is the most dangerous bias for prompt iteration, because "add markdown structure to the output" is a change anyone can make in thirty seconds and it will show up as a quality win on almost any impression-based rubric. Controls: name content criteria explicitly in the rubric, and spot check by running the judge on plain-text-stripped versions of both candidates to see how much of the delta survives. ## What does not fix these A stronger judge model reduces some of these effects but removes none. Chain-of-thought in the judge improves criterion-following but can also produce a more persuasive rationalisation of a biased verdict. Averaging more items does not help, for the reason given above. The only durable posture is to treat each bias as a measurable quantity: report your swap-flip rate, report the score-versus-length relationship, name the judge family relative to the systems being compared, and re-measure them whenever the judge or rubric changes. ## Interview framing The strongest answer names the bias, names its diagnostic, and names its control, in that order — and adds that the diagnostics are cheap enough to run continuously, so there is no excuse for not knowing your numbers.

  • How would you measure your judge's position bias rather than just mitigating it?
    Two cheap experiments. Run each pair in both orders and record the flip rate — the share of pairs whose winner changes when order changes; that is a direct read of how much of your signal is positional. Second, judge a candidate against an exact copy of itself, where the true answer is known to be a tie; any systematic winner is pure position preference. Track both numbers whenever the judge model or rubric changes.
  • Your rubric already says "do not reward length". Why does verbosity bias persist?
    Because the preference is baked into the judge's learned notion of a good answer, not into the instruction it is following. The instruction competes with a prior rather than overriding it. What actually works is removing the room for length to help: decompose into binary criteria that padding cannot satisfy, score concision as its own criterion, and check whether your score delta survives length-matching between the candidates.
  • Is a panel of judges from different families always better than one?
    It reduces self-preference and averages out family-specific quirks, and requiring agreement gives you a natural abstain signal. But it multiplies cost and latency, and shared biases — verbosity, formatting — survive the panel because every family has them. Use a panel when the decision is a cross-provider comparison or otherwise high stakes; a single cross-family judge with the swap control is usually enough for routine iteration.
  • Which of these biases most threatens a prompt-iteration loop, and why?
    Formatting halo, closely followed by verbosity. Both are trivially easy to induce with a one-line prompt edit, both raise impression-based judge scores immediately, and neither necessarily improves the output for a reader. A loop optimising against such a judge will converge on long, bulleted, confidently worded answers and report steady progress the whole way.

saying these in an interview costs you the question

  • Claiming a stronger judge model eliminates position and verbosity bias
  • Running pairwise comparisons in one fixed order only
  • Judging a model against a rival using a judge from its own family
  • Reading the judge's written rationale as evidence the verdict was unbiased
  • Treating an added markdown structure that raises scores as a real quality gain

context