How often should a golden eval set be refreshed, and what should trigger a refresh?
answer
- a frozen set ages against a moving world
- schedule quarterly, trigger on rule changes
- new version, never a silent edit
- re-baseline the old version after refresh
- name an owner and fund the churn
basics
~20 sOn a scheduled cadence plus event triggers. Schedule a review each quarter; trigger immediately when the rules the labels encode change, when the product gains a capability the set never covers, or when the traffic mix shifts. Version the set, never edit it silently.
solid answer
~50 sA frozen set is only as good as the world it froze. When a formulary update made 18 percent of a prior-authorization set's expected answers wrong overnight, the suite did not just lose accuracy — it started penalizing correct behaviour, which is worse than measuring nothing. My policy is a standing quarterly review plus hard triggers: an external rule change, a new product capability the set does not exercise, a traffic-mix shift, or a run of incidents in situations the set never contained. Refresh means cutting a new **version** of the set, not editing items in place, and re-scoring the previous system version on the new set so the baseline is restored — otherwise the next comparison confuses a moved yardstick with a real regression. I also keep a small frozen archival slice for long-term trend, and I budget the labeling cost of refresh up front, because an eval set that nobody is funded to maintain quietly stops being true.
go deeper
Know that an eval set can go out of date when the underlying rules or product change, and that stale expected answers make the score wrong rather than merely old.
Be ready to name the staleness sources — external rule changes, new capabilities, traffic shift, over-iteration — and to explain why a refresh means a new versioned set with a changelog rather than an in-place edit.
Show the operational routine: event triggers alongside a calendar cadence, re-scoring the prior system version to restore the baseline, and a concrete check that separates a moved yardstick from a genuine regression.
Own the governance — a named owner, a funded refresh budget, a churn allowance per cycle, and a frozen archival slice for long-run trend. Argue explicitly for how much continuity of measurement is worth against how fast the domain moves.
## Four ways a golden set goes stale **The world changed.** The expected answers encode facts and rules that live outside your system: a formulary update, a policy revision, a pricing change, a regulation. When those move, the labels are wrong while the items still look perfectly fine. This is the most dangerous form of staleness because it inverts the metric — the system now loses points for being right. **The product changed.** New capabilities, new tools, new supported formats. The set exercises the old surface and says nothing about the new one, so coverage silently drops even though the score holds steady. **The traffic changed.** The set was stratified against a distribution that has since shifted — a new customer segment, a new channel, a seasonal pattern. Slice weights that made the headline representative no longer do. **The set was overused.** Items iterated against for months stop being a measurement. Every prompt tweak that was accepted or rejected on this set has fitted the system to it, and the residual headroom on the set no longer predicts headroom in production. ## Cadence plus triggers A purely scheduled refresh is too slow for rule changes and too aggressive for stable domains. The workable policy is both: - **Scheduled review** — quarterly is a common default. The review is not automatically a rewrite; it is a checkpoint that asks whether the labels are still correct, whether the slice list covers current failure modes, and whether the traffic mix still matches the strata. - **Event triggers** that force a refresh regardless of the calendar: an external rule or policy change that touches the labeled decisions; a shipped capability the set does not exercise; a measurable shift in the request mix; a cluster of production incidents in situations absent from the set; a rubric change from an adjudication round. ## Refresh means a new version, not an edit Silently editing items is the most common process failure. Scores from before and after the edit are then quietly incomparable, and nobody can tell whether a step change came from the system or from the set. The discipline: 1. Cut a new set version with a version identifier, a changelog of items added, removed and re-labeled, and the rubric version it was labeled under. 2. **Re-score the previous system version on the new set** to re-establish a baseline. Without that, the first comparison after a refresh measures the set change, not the system change. 3. Keep the old version retrievable, so historical claims remain auditable. A useful supplement is a small, deliberately frozen archival slice — items chosen because their correct answer is timeless — kept unchanged for years so long-run trend has at least one stable anchor. ## Distinguishing a moved yardstick from a real regression When a score drops, the first question is which of the two moved. Cheap discriminators: did the drop concentrate in exactly the categories the external change touched; do the failing items' *expected* answers still look right to a reviewer; does the previous system version, re-run today, also score lower on the same items. If the old version regresses too, the set changed, not the system. Building that check into the refresh routine prevents the classic incident where a team spends a week debugging a model for a formulary update. ## Churn budget and the maintenance economics Refresh is not free — it is re-labeling, re-adjudicating, and re-running baselines. A useful governance device is a churn budget: an expected fraction of the set that may change per cycle, say 10 to 20 percent, with anything larger requiring an explicit decision. Too little churn and the set ossifies; too much and no metric survives long enough to show a trend. The deeper point is ownership. An eval set with no named owner and no funded maintenance decays exactly like documentation, except its decay is invisible until it produces a wrong decision. Whoever owns the product decision the set informs should own the set, and the refresh cost should be a line item in the same budget as the feature work it gates. ## Interaction with over-fitting Because months of iteration fit the system to the set, refresh also serves a hygiene role: rotating in fresh items restores headroom and re-exposes weaknesses the team has been implicitly tuning around. Keeping a portion of each refresh untouched by any iteration — added, labeled and reserved — gives you a clean surface for the decisions that actually matter, at the cost of measuring less often on it.
- After a refresh, a score drops. How do you tell whether the model regressed or the set moved?Re-run the previous system version on the new set. If the old version also scores lower on the same items, the yardstick moved. Then check whether the drop concentrates in the categories the external change touched, and have a reviewer confirm that the failing items' expected answers are still correct. Building that check into every refresh prevents debugging a model for a policy update.
- How much of a golden set should change per refresh cycle?Treat it as a budget — something like 10 to 20 percent per cycle, with larger churn requiring an explicit decision. Too little and the set ossifies against a world that has moved; too much and no metric survives long enough to show a trend. Pair it with a small frozen archival slice of timeless items so long-run comparisons keep one stable anchor.
- Why does long-term iteration against one eval set degrade it even when nothing external changes?Every prompt change accepted or rejected on that set fits the system to those specific items, so remaining headroom on it stops predicting headroom in production. Rotating in fresh items each refresh restores signal, and reserving part of each intake untouched by iteration gives you a clean surface for decisions that matter.
saying these in an interview costs you the question
- The set is golden, it never changes
- Just fix the wrong labels in place, quietly
- Score dropped after refresh, so the model regressed
- Refresh whenever someone notices a bad label
- No owner needed, the set maintains itself