How do you sample production traces into a golden eval set without losing rare cases?
answer
- uniform draw buries the tail
- quota per cell, not per corpus
- strata: segment crossed with failure mode
- keep production frequencies as weights
- provenance and de-duplication per item
basics
~20 sSample by strata, not uniformly. Bucket traces by the dimensions you care about — request type, customer segment, known failure mode — then take a quota from each bucket, so rare but costly cases appear in numbers large enough to score.
solid answer
~50 sA uniform draw from traffic reproduces the traffic's own imbalance: on a pharmacy prior-authorization assistant, routine refills would swamp the specialty-drug denials that actually cause harm. So I stratify. I take 90 days of traces, bucket them on the axes that matter — six drug classes crossed with three payer plan types, plus a bucket per known failure mode — and draw a fixed quota per cell, capping how many items any single high-volume account contributes and de-duplicating near-identical requests. That gives each cell enough items to carry its own score. I still record the real production frequency of each cell, so a headline number can be computed as a frequency-weighted average while per-cell scores stay unweighted. Every item keeps provenance: source trace id, date, who labeled it, and whether it came from traffic, an escalation, or was written by hand.
code
python · 25 linesimport random
from collections import defaultdict
def stratified_sample(traces, key, per_stratum, seed=7):
buckets = defaultdict(list)
for t in traces:
buckets[key(t)].append(t)
rng = random.Random(seed)
out = []
for name in sorted(buckets):
items = buckets[name]
rng.shuffle(items)
out.extend(items[:per_stratum])
return out
traces = [
{"id": i, "drug_class": c, "plan": p}
for i, (c, p) in enumerate(
[(c, p) for c in "ABCDEF" for p in "XYZ"] * 40
)
]
sample = stratified_sample(traces, lambda t: (t["drug_class"], t["plan"]), 8)
print(len(sample))go deeper
Know that an eval set is a fixed, labeled collection of inputs you re-run after every change, and that where the items come from decides what the score can tell you.
Be ready to explain stratified sampling concretely: choose the axes, set a per-cell quota, de-duplicate, and keep production frequencies so a weighted headline is still available. Name the failure mode of a plain random draw.
Show the operational side: provenance metadata, per-account caps, a held-out portion so prompt tuning cannot leak into the measurement, and a coverage check against a real incident taxonomy.
Own the tradeoff between representativeness and sensitivity — an over-sampled set detects cohort regressions but needs weights to speak about user impact. Decide who owns the taxonomy, what labeling budget it justifies, and which cohorts are worth their own quota at all.
## What a golden eval set is A golden eval set (also called a golden set or eval suite) is a frozen, versioned collection of inputs paired with an expected outcome or a grading rubric. Every downstream number — a prompt-version comparison, a judge-model score, a regression check — is computed over it. That makes its composition the highest-leverage decision in the whole evaluation stack: if the set does not contain the situations you care about, no metric computed on it can tell you about them. A weak set makes every score decorative. ## Why a uniform sample of traffic fails Production traffic is heavy-tailed. On a pharmacy prior-authorization assistant, the overwhelming majority of requests are routine refills of common generics, which the system handles almost perfectly. The requests that cause real damage — specialty-drug denials, plan-specific step-therapy rules, edge-case dosing — might be one percent of volume. Draw 300 traces uniformly and you get roughly three of them: too few to score, and a change that breaks that whole cohort moves the aggregate by a fraction of a point. The set is technically representative and practically blind. The opposite mistake is sampling only from complaints and thumbs-down feedback. That set is dense in failures, but it is a biased sample of the world: it cannot tell you whether the ordinary path still works, and it drifts toward whatever users bother to report. ## Stratify on the axes that matter Stratified sampling means partitioning the population into strata and drawing a quota from each, rather than drawing from the pool as a whole. Two families of axis are usually worth stratifying on: - **Who or what the request is** — segment, plan type, product line, language, document length, channel. For the prior-auth assistant: six drug classes crossed with three payer plan types gives eighteen cells. - **What can go wrong** — a failure-mode taxonomy built from real incidents: missing clinical criterion, wrong plan rule applied, hallucinated formulary entry, refusal on a valid request. Draw a fixed quota per cell — say eight to thirty items depending on how precisely you need to score that cell. The quota, not the population share, controls how much signal each cohort carries. ## Keep the production mix as metadata, not as the sample Over-sampling rare cells means the unweighted mean of the set no longer estimates production quality. That is fine, provided you keep each cell's true production frequency alongside it. Then you can report two numbers: a frequency-weighted average that approximates what users experience, and the raw per-cell scores that tell you which cohort is broken. Losing the weights is what makes teams either over-sample and then misread the headline, or refuse to over-sample at all. ## Hygiene that decides whether the set is usable - **De-duplicate.** Traces from a single integration often repeat almost verbatim; a hundred copies of one request is one item of information with a hundred votes. Near-duplicate detection plus a per-account cap fixes it. - **Strip and pseudonymize identifiers** before the set leaves the production boundary, since eval sets get copied into notebooks, CI logs and vendor dashboards. - **Record provenance** per item: source (trace, escalation, hand-written adversarial, synthetic perturbation), date captured, labeler, rubric version. You will need to report trace-derived and synthetic slices separately. - **Hold out.** Keep a portion of the set untouched during prompt iteration. A set you tune against stops being a measurement and becomes a training signal — the same overfitting problem as a leaked test split. ## Filling gaps traffic cannot Some situations have never occurred in production but will: a new payer plan, a drug class you are about to launch, an adversarial input someone will eventually send. Hand-written and synthetically perturbed items cover those, and they are legitimate — as long as they are marked as such and reported in their own slice, because they carry no evidence about frequency. ## Judging the set itself A good set discriminates. If the current system passes every item, the set has no headroom and cannot detect improvement; if it fails every item, it cannot detect regression either. Look at the spread of scores, check coverage against the failure taxonomy, and prune items that neither version ever gets right into a backlog slice rather than leaving them to drag the headline. Sizing — how many items each cell actually needs before a difference means anything — is a separate calculation, and it usually pushes the quotas higher than instinct suggests.
- If you deliberately over-sample rare denials, how do you still report a number that reflects what users experience?Keep each stratum's true production frequency alongside the set and report a frequency-weighted average as the headline, while per-stratum scores stay unweighted. Two numbers, clearly labeled: one estimates the user-visible rate, the other tells you which cohort is broken. Losing the weights is what makes an over-sampled set unreadable.
- One enterprise account generates 40% of your traces. How does that change the sample?Cap what any single account contributes, and de-duplicate near-identical requests before sampling. Otherwise the set measures one customer's phrasing and integration quirks, and a prompt change that helps them but hurts everyone else scores as a win. I would also give that account its own slice so its behaviour is visible rather than dominant.
- How do you cover situations that have never appeared in production traffic?Write them by hand or generate perturbations of real items — new plan types, adversarial inputs, formats you are about to support. Mark them as synthetic and report them as a separate slice, because they carry no information about how often the situation occurs. They are a coverage instrument, not a frequency estimate.
saying these in an interview costs you the question
- Just take the last 500 requests, that is representative
- Bigger is always better, size beats composition
- Sample only from thumbs-down feedback and call it golden
- Keep every near-duplicate trace, more data is more signal
- Tune prompts against the same items you report scores on