skip to content

Before a suite trusts a model-judged check, how do you calibrate it against human-labelled examples?

level: seniorimportance: nice to knowfreq 24%

answer

  1. Show it outputs whose verdict you know
  2. Read the disagreements, not the score
  3. Obvious examples teach you nothing
  4. Near-misses are where defects live
  5. Re-run it when anything changes

basics

~20 s

Run the judged check over outputs people have already labelled against the written standard, then read every disagreement. Trust it only once it catches known past defects and the subtly-wrong near-misses, not just obviously good and obviously broken examples.

solid answer

~50 s

Collect real outputs of the kind the check will meet, have people label each against the written standard, run the check blind over the same set, and go through the disagreements one at a time. Most disagreements turn out to be an unclear standard rather than a wrong verdict, which is itself the finding. The set has to contain real past defects, legitimate variants on the good side (other phrasings, other locales, empty states), and above all the **near-misses** — the right amount with the wrong currency word, the accurate summary missing a required disclosure — because that is where real defects live and where a judged check is weakest. Hold some examples back from tuning, label subjective cases with more than one person, and re-run the whole set whenever the standard, the judge or the output's shape changes.

code

yaml · 19 lines
yaml
# calibration-set/refund-message.yaml  (labelled by hand, held in the repo)
standard_version: refund-message-v4
examples:
  - id: known-defect-2024-11
    label: unacceptable
    why: names a second customer in the closing line
  - id: near-miss-currency
    label: unacceptable
    why: amount correct, currency word wrong
  - id: near-miss-missing-disclosure
    label: unacceptable
    why: accurate but omits the arrival-time statement
  - id: variant-locale-short-form
    label: acceptable
    why: different phrasing, all required parts present
  - id: variant-empty-state
    label: acceptable
    why: nothing to refund, message correctly says so
held_back: [near-miss-rounding, variant-long-name]

go deeper

for a junior

Recall that a judged check has to be tried against outputs whose correct verdict a person already decided, before anyone relies on it. Knowing that trust is earned on labelled evidence is the point at this level.

for a middle

Explain how the run works: label real outputs against the written standard, run the check blind over the same set, and read every disagreement as either a wrong verdict or an unclear standard.

for a senior

Show judgement about the set's composition: known past defects, deliberate near-misses, legitimate variants on the good side, enough of the failing class, held-back examples, and re-calibration when the standard, the judge or the output shape changes.

for a principal

Own the economics: which judged checks are worth a maintained labelled set at all, who keeps it current, and what the team does when a judge update moves verdicts across several suites at once.

## Why calibration comes before trust A judged check is a decision procedure you did not write and cannot read. The only way to learn what it actually does is to run it over outputs whose correct verdict you already know and compare its answers with yours. That is calibration: not a statistical exercise, but the plain act of showing the check a body of labelled evidence before letting it report on anything real. The sequence is simple and the discipline is in doing it at all. 1. Collect real outputs of the kind the check will see in the suite. 2. Have people label each one against the written standard: acceptable, or not, and why. 3. Run the judged check over the same set without the labels. 4. Read the disagreements one by one. Each is either the check being wrong or the standard being unclear, and the second is far more common than teams expect. 5. Rewrite the standard, and re-run on examples you held back from the tuning. ## What the set has to contain A calibration set assembled from convenient examples teaches you nothing, because convenient examples are the ones every check gets right. - **Real past defects.** The failures this check exists to catch, reproduced as outputs. If it cannot recognise the problems you already know about, it will not recognise the next one. - **The near-misses.** This is the part teams skip and the part that decides everything. An amount that is right with the wrong currency word; a summary that is accurate but omits the one required disclosure; a message that is polite and actionable but names a different customer. Subtly wrong outputs are where a judged check is weakest and where real defects actually live. - **Legitimate variety on the good side.** Different phrasings, different locales, empty states, long values, the seasonal template. Without these you learn nothing about how often the check raises alarms on output that was fine, and a check that cries wolf gets ignored within a sprint. - **Enough of the failing class.** Sixty examples containing two bad ones tells you essentially nothing about missed defects: the check can agree with you on fifty-eight and still be blind to the very thing it was built for. Weight the set deliberately toward the failure modes you care about rather than mirroring their natural rarity. - **Genuine outputs, not idealised ones.** Examples hand-written to demonstrate the standard are cleaner than production output and hide exactly the messiness that causes trouble. - **Held-back examples.** Keep part of the set out of the loop while you tune the standard. Otherwise you are only measuring how well you fitted the wording to those specific cases. Where the standard is subjective, label with more than one person. Human disagreement is not noise to be averaged away; it is a direct measurement that the standard is under-specified, and it is cheaper to discover it here than in a disputed build failure. | The set contains | What you learn | If you omit it | | --- | --- | --- | | Known past defects | Whether it catches what it was built for | You trust it on faith | | Near-misses | Its behaviour where real defects live | It looks excellent and is blind | | Legitimate variants | How often it flags acceptable output | It cries wolf and gets ignored | | Held-back examples | Whether the standard generalises | You have tuned to the examples | ## Reading the result, and re-doing it Two directions of error, and they do not cost the same. A judged check that passes a genuinely broken output has silently removed a check somebody believed in; a judged check that flags good output costs attention and, if it happens often, credibility. For a test suite the first is usually the more expensive, which is why the near-misses and a decent number of known-bad examples matter more than a large agreeable set. Calibration is not a one-off. Re-run the set whenever the judging standard is edited, whenever the model doing the judging is updated, and whenever the product's output changes shape — a new field, a restructured message, a new language. Each of these can move verdicts on outputs you have already labelled, and the labelled set is the only cheap way to find out before the suite tells you in a confusing way. Keeping the set in the repository next to the standard, and re-running it in the pipeline as its own case, turns all of that from a good intention into something that actually happens.

  • Two people label the same output differently. What do you do with that disagreement?
    Treat it as a measurement, not noise. Human disagreement means the written standard does not decide the case, so no judged check built on it can be stable either. Resolve the case explicitly, write the resolution into the standard as a clarifying clause or example, then re-label. Averaging the two labels away hides the only cheap warning you get that the criterion is under-specified.
  • When does a calibrated judged check need re-calibrating?
    Whenever any of its three inputs move: the judging standard is edited, the model doing the judging is updated, or the product's output changes shape with a new field, a restructured message or a new language. Each can shift verdicts on outputs you already labelled. Keeping the labelled set in the repository and re-running it as its own pipeline case makes that automatic rather than aspirational.
  • Why not build the calibration set from hand-written ideal examples?
    Because they are cleaner than anything the suite will actually see. Real output carries odd lengths, unusual names, empty states, truncated fields and locale differences, and those are what push a judged verdict over the line. A set of tidy specimens produces a check that looks excellent in calibration and behaves unpredictably on the first real run.

Hiring a proofreader by handing them pages you have already marked up: the point is not the pages they get right, it is the errors you planted that they walked past.

saying these in an interview costs you the question

  • Calibrating only on obviously good and obviously broken outputs
  • Labelling a set that contains almost no bad examples
  • Tuning the standard on the same examples used to judge it
  • Averaging away disagreement between two human labellers
  • Calibrating once and never repeating it