Two reviewers disagree on 31 of 400 eval labels — what is your adjudication protocol?
answer
- disagreement points at the rubric
- blind, independent, then third-reviewer tiebreak
- cluster the splits, generalize the rule
- some items are genuinely unlabelable
- labels are the ceiling on measurable difference
basics
~20 sRoute the disputed items to a third, independent adjudicator, then read the resolved cases together. Most disagreement is a symptom of an underspecified rubric, so the real output is a sharper rubric plus a re-label of the affected category — not just 31 settled labels.
solid answer
~50 sFirst, the disagreement is data, not an inconvenience. On a prior-authorization set, two pharmacists labeled 400 cases independently and blind to each other and to the model's output; 31 splits went to a third reviewer whose decision is recorded with a one-line reason. Then I read those 31 as a group. If they cluster — say, most involve whether a missing lab value should count as an incomplete submission or a denial — that is a rubric gap, and the fix is to write the rule explicitly and re-label every item in that category, not only the disputed ones. A residue of genuinely ambiguous items always survives; I mark those as unlabelable and keep them in their own slice rather than forcing a coin-flip label into the pass rate. And I treat the human agreement level as a ceiling: a metric cannot be more reliable than the labels underneath it.
go deeper
Know that eval labels come from people, that two people can label the same case differently, and that disagreements get resolved by a third reviewer rather than by guessing.
Be ready to describe the loop end to end: blind independent labeling, a pilot batch, third-reviewer adjudication with recorded reasons, then a rubric fix and a re-label of the affected category.
Demonstrate the judgment: an unlabelable bucket instead of forced labels, blinding reviewers to the system version and its rationale to avoid anchoring, and treating the disagreement rate as the ceiling on what any metric can detect.
Own the economics and the governance — how much of the set is double-labeled, who owns the rubric and its versioning, and the fact that a persistently ambiguous slice is a signal the product's definition of correct needs a decision, not more annotators.
## Why disagreement is the useful signal When two qualified reviewers label the same eval item differently, one of three things is true: one of them made a mistake, the rubric does not decide the case, or the case is genuinely ambiguous and no rubric could decide it without an arbitrary convention. Only the first is noise. The second is the most common and the most valuable, because it points at a definition your whole evaluation depends on and has never actually written down. Teams that treat disagreement purely as a quality problem in their annotators — retraining people, tightening SLAs — usually leave the underlying ambiguity in place, where it keeps producing inconsistent labels forever. ## The protocol **1. Label blind and independently.** Both reviewers see the same input and the same candidate output, and neither sees the other's label. Critically, reviewers should not know which system version produced the output, and ideally should not see the model's own justification before deciding — reading a fluent rationale anchors a reviewer toward accepting it. Where you can, show outputs in randomized order and strip version identifiers. **2. Pilot on a small batch first.** Run twenty or thirty items through both reviewers before committing to four hundred. The disagreements from that pilot are the cheapest rubric feedback you will ever get, and fixing the rubric before the bulk pass avoids re-labeling the whole set later. **3. Adjudicate the splits.** A third reviewer, again blind to who labeled what, resolves each disputed item and records a short reason. The reason is the artifact that matters; a bare tiebreak label teaches you nothing. **4. Cluster and generalize.** Read the adjudicated items as a set. Do they concentrate in one category, one failure mode, one boundary condition? A cluster means a rule is missing. Write it into the rubric with a worked example on each side of the line, then re-label everything in the affected category — because the agreed labels in that category were agreed for the wrong reason as often as the disputed ones were disputed. **5. Version the rubric with the set.** A label is only meaningful under the rubric that produced it. When the rubric changes, the set version changes, and scores from before and after are not directly comparable until the previous system version is re-scored. ## What to do with irreducible ambiguity Some items resist every rubric: two defensible answers, a request whose intent is genuinely unclear, a case where the correct behaviour depends on information the system was never given. Forcing a label on those items injects noise into every future comparison and, worse, injects it in a way that looks like signal. The cleaner move is a third category — unlabelable, or ambiguous — kept in the set but excluded from the headline pass rate and reported as its own slice. Its size is informative on its own: a growing ambiguous slice means the product's own definition of correct is drifting, which is a design conversation, not a labeling one. Some teams instead label such items with an accepted-set of valid answers, which works when the ambiguity is over form rather than over what the right decision is. ## The label-noise ceiling Human labels are the measuring instrument, and no instrument reports differences smaller than its own error. If independent reviewers agree on roughly 92 percent of items in a category, then two system versions whose true quality differs by two points in that category cannot be distinguished on that set no matter how many items you add, because the disagreement is not random noise that averages out — it is systematic, tied to the same hard cases every time. This is the most common reason a carefully built eval suite fails to reproduce a difference the team is sure exists. The practical consequences: measure the disagreement rate and publish it next to the scores; spend labeling budget on rubric clarity before spending it on volume; and be explicit that any automated grader you later validate against these labels inherits the same ceiling — it can at best reproduce the humans, including their inconsistency. ## Cost discipline Double-labeling everything is expensive. A common compromise is to double-label a random sample — perhaps 15 to 20 percent — plus every item in high-stakes categories, and single-label the rest. That gives an ongoing measurement of reviewer consistency without paying twice for the whole set. Re-measure it after every rubric change, because the number you are quoting was earned under the old definition.
- What do you do with items that stay ambiguous after adjudication?Mark them unlabelable and hold them in a separate slice, excluded from the headline pass rate. Forcing an arbitrary label injects noise that looks like signal in every future comparison. The size of that slice is itself a metric: if it grows, the product's own definition of a correct answer is drifting and that is a design problem, not a labeling one.
- How does the human agreement level limit what your evals can detect?It is a ceiling. If independent reviewers disagree on eight percent of a category, differences smaller than that in the same category cannot be resolved on that set, because the disagreement is systematic — the same hard cases every time — and does not average out with more items. Publish the disagreement rate beside the scores so nobody over-reads a small win.
- Double-labeling 400 items twice is expensive. Where would you cut?Double-label a random 15 to 20 percent plus everything in high-stakes categories, and single-label the rest. That keeps a live measurement of reviewer consistency at a fraction of the cost. Re-measure after any rubric change, since the old number was earned under the old definition.
saying these in an interview costs you the question
- Disagreement just means one reviewer was careless
- Take the senior reviewer's label and move on
- Drop every disputed item and keep the clean ones
- Force a label on ambiguous cases so counts stay round
- Show reviewers the model's rationale before they label