skip to content

Eval Dataset Curation

You learn to build the golden set your evals run against: where the examples come from, how they are labeled and sliced, and how many you need before a score difference means anything. A weak dataset makes every downstream metric decorative.

on this pageshow

questions

5

How do you sample production traces into a golden eval set without losing rare cases?

level: middleimportance: must knowfreq 60%

answer

  1. uniform draw buries the tail
  2. quota per cell, not per corpus
  3. strata: segment crossed with failure mode
  4. keep production frequencies as weights
  5. provenance and de-duplication per item

basics

~20 s

Sample by strata, not uniformly. Bucket traces by the dimensions you care about — request type, customer segment, known failure mode — then take a quota from each bucket, so rare but costly cases appear in numbers large enough to score.

solid answer

~50 s

A uniform draw from traffic reproduces the traffic's own imbalance: on a pharmacy prior-authorization assistant, routine refills would swamp the specialty-drug denials that actually cause harm. So I stratify. I take 90 days of traces, bucket them on the axes that matter — six drug classes crossed with three payer plan types, plus a bucket per known failure mode — and draw a fixed quota per cell, capping how many items any single high-volume account contributes and de-duplicating near-identical requests. That gives each cell enough items to carry its own score. I still record the real production frequency of each cell, so a headline number can be computed as a frequency-weighted average while per-cell scores stay unweighted. Every item keeps provenance: source trace id, date, who labeled it, and whether it came from traffic, an escalation, or was written by hand.

code

python · 25 lines
python
import random
from collections import defaultdict


def stratified_sample(traces, key, per_stratum, seed=7):
    buckets = defaultdict(list)
    for t in traces:
        buckets[key(t)].append(t)
    rng = random.Random(seed)
    out = []
    for name in sorted(buckets):
        items = buckets[name]
        rng.shuffle(items)
        out.extend(items[:per_stratum])
    return out


traces = [
    {"id": i, "drug_class": c, "plan": p}
    for i, (c, p) in enumerate(
        [(c, p) for c in "ABCDEF" for p in "XYZ"] * 40
    )
]
sample = stratified_sample(traces, lambda t: (t["drug_class"], t["plan"]), 8)
print(len(sample))

go deeper

for a junior

Know that an eval set is a fixed, labeled collection of inputs you re-run after every change, and that where the items come from decides what the score can tell you.

for a middle

Be ready to explain stratified sampling concretely: choose the axes, set a per-cell quota, de-duplicate, and keep production frequencies so a weighted headline is still available. Name the failure mode of a plain random draw.

for a senior

Show the operational side: provenance metadata, per-account caps, a held-out portion so prompt tuning cannot leak into the measurement, and a coverage check against a real incident taxonomy.

for a principal

Own the tradeoff between representativeness and sensitivity — an over-sampled set detects cohort regressions but needs weights to speak about user impact. Decide who owns the taxonomy, what labeling budget it justifies, and which cohorts are worth their own quota at all.

## What a golden eval set is A golden eval set (also called a golden set or eval suite) is a frozen, versioned collection of inputs paired with an expected outcome or a grading rubric. Every downstream number — a prompt-version comparison, a judge-model score, a regression check — is computed over it. That makes its composition the highest-leverage decision in the whole evaluation stack: if the set does not contain the situations you care about, no metric computed on it can tell you about them. A weak set makes every score decorative. ## Why a uniform sample of traffic fails Production traffic is heavy-tailed. On a pharmacy prior-authorization assistant, the overwhelming majority of requests are routine refills of common generics, which the system handles almost perfectly. The requests that cause real damage — specialty-drug denials, plan-specific step-therapy rules, edge-case dosing — might be one percent of volume. Draw 300 traces uniformly and you get roughly three of them: too few to score, and a change that breaks that whole cohort moves the aggregate by a fraction of a point. The set is technically representative and practically blind. The opposite mistake is sampling only from complaints and thumbs-down feedback. That set is dense in failures, but it is a biased sample of the world: it cannot tell you whether the ordinary path still works, and it drifts toward whatever users bother to report. ## Stratify on the axes that matter Stratified sampling means partitioning the population into strata and drawing a quota from each, rather than drawing from the pool as a whole. Two families of axis are usually worth stratifying on: - **Who or what the request is** — segment, plan type, product line, language, document length, channel. For the prior-auth assistant: six drug classes crossed with three payer plan types gives eighteen cells. - **What can go wrong** — a failure-mode taxonomy built from real incidents: missing clinical criterion, wrong plan rule applied, hallucinated formulary entry, refusal on a valid request. Draw a fixed quota per cell — say eight to thirty items depending on how precisely you need to score that cell. The quota, not the population share, controls how much signal each cohort carries. ## Keep the production mix as metadata, not as the sample Over-sampling rare cells means the unweighted mean of the set no longer estimates production quality. That is fine, provided you keep each cell's true production frequency alongside it. Then you can report two numbers: a frequency-weighted average that approximates what users experience, and the raw per-cell scores that tell you which cohort is broken. Losing the weights is what makes teams either over-sample and then misread the headline, or refuse to over-sample at all. ## Hygiene that decides whether the set is usable - **De-duplicate.** Traces from a single integration often repeat almost verbatim; a hundred copies of one request is one item of information with a hundred votes. Near-duplicate detection plus a per-account cap fixes it. - **Strip and pseudonymize identifiers** before the set leaves the production boundary, since eval sets get copied into notebooks, CI logs and vendor dashboards. - **Record provenance** per item: source (trace, escalation, hand-written adversarial, synthetic perturbation), date captured, labeler, rubric version. You will need to report trace-derived and synthetic slices separately. - **Hold out.** Keep a portion of the set untouched during prompt iteration. A set you tune against stops being a measurement and becomes a training signal — the same overfitting problem as a leaked test split. ## Filling gaps traffic cannot Some situations have never occurred in production but will: a new payer plan, a drug class you are about to launch, an adversarial input someone will eventually send. Hand-written and synthetically perturbed items cover those, and they are legitimate — as long as they are marked as such and reported in their own slice, because they carry no evidence about frequency. ## Judging the set itself A good set discriminates. If the current system passes every item, the set has no headroom and cannot detect improvement; if it fails every item, it cannot detect regression either. Look at the spread of scores, check coverage against the failure taxonomy, and prune items that neither version ever gets right into a backlog slice rather than leaving them to drag the headline. Sizing — how many items each cell actually needs before a difference means anything — is a separate calculation, and it usually pushes the quotas higher than instinct suggests.

  • If you deliberately over-sample rare denials, how do you still report a number that reflects what users experience?
    Keep each stratum's true production frequency alongside the set and report a frequency-weighted average as the headline, while per-stratum scores stay unweighted. Two numbers, clearly labeled: one estimates the user-visible rate, the other tells you which cohort is broken. Losing the weights is what makes an over-sampled set unreadable.
  • One enterprise account generates 40% of your traces. How does that change the sample?
    Cap what any single account contributes, and de-duplicate near-identical requests before sampling. Otherwise the set measures one customer's phrasing and integration quirks, and a prompt change that helps them but hurts everyone else scores as a win. I would also give that account its own slice so its behaviour is visible rather than dominant.
  • How do you cover situations that have never appeared in production traffic?
    Write them by hand or generate perturbations of real items — new plan types, adversarial inputs, formats you are about to support. Mark them as synthetic and report them as a separate slice, because they carry no information about how often the situation occurs. They are a coverage instrument, not a frequency estimate.

saying these in an interview costs you the question

  • Just take the last 500 requests, that is representative
  • Bigger is always better, size beats composition
  • Sample only from thumbs-down feedback and call it golden
  • Keep every near-duplicate trace, more data is more signal
  • Tune prompts against the same items you report scores on

context

open as a page

Why report sliced eval scores by segment and failure mode instead of one aggregate?

level: seniorimportance: must knowfreq 50%

basics

~20 s

An aggregate averages cohorts together, so a large healthy segment hides a small broken one. Slicing by who the request came from and by what went wrong exposes concentrated regressions, and lets you gate on the worst slice rather than the mean.

open as a page

Two reviewers disagree on 31 of 400 eval labels — what is your adjudication protocol?

level: middleimportance: should knowfreq 42%

basics

~20 s

Route the disputed items to a third, independent adjudicator, then read the resolved cases together. Most disagreement is a symptom of an underspecified rubric, so the real output is a sharper rubric plus a re-label of the affected category — not just 31 settled labels.

open as a page

Your 60-item eval suite shows a 4-point win — is that difference real?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Almost certainly not. On 60 pass/fail items scoring around 80 percent, the 95 percent interval is roughly plus or minus 10 points, so a 4-point gap is well inside noise. Detecting it needs pairing, graded scores, or several hundred items.

open as a page

How often should a golden eval set be refreshed, and what should trigger a refresh?

level: principalimportance: should knowfreq 28%

basics

~20 s

On a scheduled cadence plus event triggers. Schedule a review each quarter; trigger immediately when the rules the labels encode change, when the product gains a capability the set never covers, or when the traffic mix shifts. Version the set, never edit it silently.

open as a page