How do you validate an LLM judge used to filter synthetic fine-tuning data?
answer
- the filter is a classifier
- measure it against expert labels
- false accepts teach the model errors
- false rejects skew the surviving distribution
- re-measure when generator or rubric changes
basics
~20 sTreat the filter as a classifier and measure it. Hand-label a few hundred generated rows with a domain expert, then compute the judge's precision and recall against those labels at the exact threshold you plan to ship, and re-measure whenever the generator or the rubric changes.
solid answer
~50 sOn the veterinary triage set we scored every generated pair 1-5 on correctness and clinical specificity and dropped everything below 4. That decision rule is a binary classifier, and shipping an unmeasured classifier over your training data is the whole risk. Have a vet label a stratified sample of a few hundred rows as keep or drop, then compute precision and recall of the judge against those labels — and sweep the threshold rather than assuming 4 was right. The two errors cost different things: a false accept injects a wrong clinical instruction that the model will learn as truth, while a false reject mostly costs you a row, but systematically, because the judge rejects whatever it dislikes and quietly skews the surviving distribution toward its own preferences. Use programmatic checks wherever the property is actually checkable, keep a continuous spot-check of a few percent, and re-run the validation when anything upstream changes.
go deeper
Know that scoring generated training rows with a model and dropping the low scorers is a filter, and that somebody has to check whether the filter's verdicts match a human expert's.
Explain the measurement: label a sample by hand, compute precision and recall at the shipping threshold, and describe why false accepts and false rejects cost different things.
Show the second-order reasoning — that a filter reshapes the surviving distribution toward whatever the judge prefers — and that you would sweep the threshold, prefer programmatic checks, and keep a continuous expert spot-check.
Own the filter as governed infrastructure: a versioned rubric, a standing expert-labelled gold sample, revalidation triggered by any upstream change, and an explicit position on which errors the organisation is willing to train on.
## The filter is a classifier, so measure it like one Running a judge model over generated training rows and keeping the ones it scores highly is a filtering decision, not a scoring exercise. Every row is being classified keep or drop, and the resulting dataset is whatever the classifier let through. If you would not deploy an unmeasured classifier into a production path, you should not deploy one over your training data either — it decides what the model learns. The measurement is the ordinary one. Draw a sample of generated rows — stratified across the score bands, not just the borderline ones — and have a domain expert label each keep or drop under the same written rubric the judge was given. Then compute, at the threshold you intend to use: - **Precision** — of the rows the judge kept, what fraction the expert also would have kept. Low precision means bad rows are entering training. - **Recall** — of the rows the expert would keep, what fraction the judge kept. Low recall means you are discarding good data, and paying for it twice: once in generation cost, once in coverage. A few hundred labelled rows is usually enough to see whether the filter is trustworthy at all; sweep the threshold across that same labelled sample rather than accepting the first cut-off someone picked. ## The two errors are not symmetric **False accepts** are the dangerous ones. A wrong row that survives the filter becomes a training target, and the model learns it as the correct behaviour. In a clinical-adjacent domain, one confidently wrong triage instruction repeated across a few dozen surviving rows is a real harm, not a statistical nuisance. **False rejects** look benign — you lose a row you paid to generate — but they are not random. The judge rejects according to its own preferences, so the rows it drops share properties: they are terser, less conventionally formatted, more hedged or less hedged, further from the register the judge finds "good". The survivors are therefore a biased sample of what you generated, and the fine-tuned model inherits that bias. This is the second-order effect people miss: a filter does not merely subtract quality problems, it reshapes the distribution. Where the filter also removes the harder or more unusual cases, it can undo the coverage work that went into generating them. ## Correlated blind spots When the generator and the judge are the same model, or close relatives, their failure modes overlap: the judge is least likely to notice exactly the mistakes the generator is most likely to make, and it will favour text in its own style. Using a different model family for judging, or a rubric written from the domain expert's failure list rather than from general notions of quality, reduces the overlap. It does not eliminate it, which is why the human-labelled sample remains the ground truth for the filter. ## Prefer checks that do not need a judge Anything mechanically verifiable should be verified mechanically, before the judge sees the row: schema validity, required fields present, the urgency band drawn from the permitted set, no contradiction between the stated band and the recommended action, length within bounds, no leakage of the generation prompt. Programmatic checks are deterministic, free, and never drift. Reserve the judge for the genuinely subjective residue — is the clinical detail specific enough, does the reasoning actually support the conclusion — and write its rubric as concrete criteria rather than "rate the quality". ## Ongoing spot-checks and drift One validation is a snapshot. Keep a continuous sample — a few percent of accepted rows reviewed by the expert — because the thing being filtered changes: new generation prompts, a new seed batch, a new judge model version, a rubric edit. Any of those invalidates the earlier precision and recall figures. Cheap monitoring signals help between full validations: the accept rate per generation batch, the score distribution shape, and the share of accepted rows falling in each coverage cell. ## Cost, and when to skip the filter Judging every generated row costs a model call per row, which for a large synthetic set is a real line item. Cheaper arrangements: run programmatic checks first so the judge only sees rows that could plausibly pass; judge a sample to estimate batch quality and regenerate whole batches that fail rather than filtering row by row; use a smaller judge for an initial pass and a stronger one on the borderline band. Whichever you choose, the validation requirement does not change — you still need to know the precision and recall of the rule you are applying.
- How would you choose the score threshold rather than defaulting to 4 out of 5?Sweep it on the human-labelled sample. Compute precision and recall at each cut-off and pick by which error you can afford: for clinical content, favour precision even at the cost of discarding good rows, then regenerate to replace the volume. Also check what each threshold does to coverage — if raising it empties particular case types, the quality gain is being paid for with a distribution hole.
- The judge and the generator are the same model. What is the concrete risk?Their blind spots are correlated, so the judge is least likely to catch precisely the errors the generator makes most, and it will systematically prefer text in its own style. That inflates the accept rate without inflating quality. Using a different model family for judging, and grounding the rubric in the domain expert's real failure list, reduces the overlap — but the human-labelled sample remains the only trustworthy check.
- Which quality checks should never reach the judge at all?Anything mechanically decidable: schema validity, required fields, enum values such as the urgency band, contradictions between a stated band and the recommended action, length bounds, prompt leakage. Deterministic checks are free, never drift, and cannot be talked into a wrong verdict. Running them first also cuts judging cost, since the judge only sees rows that could plausibly pass.
saying these in an interview costs you the question
- A judge filter needs no validation of its own
- A high average accepted score proves the filter works
- Filtering only removes bad rows, nothing else changes
- Judge and generator from the same model is fine
- One validation holds after the rubric is rewritten