Your prompt search scores every candidate with a judge model — how do you keep that affordable and trustworthy?
answer
- Cost scales with candidates times examples
- Cheap checks first, judge last
- Freeze the judge for the run
- Anchor set detects scale drift
- Validate agreement against human labels
basics
~20 sTreat the judge as expensive infrastructure: cascade cheap programmatic checks in front of it so it only ranks survivors, pin its model version and rubric for the whole campaign so scores stay comparable across generations, and re-validate it against human labels on a fixed anchor set.
solid answer
~50 sJudge scoring costs one model call per candidate per example, so 40 candidates against 250 examples is 10,000 calls per generation before any repeats — the objective becomes the dominant line item in a long campaign. My first move is a cascade: hard gates and programmatic checks run first for free, and the judge only scores what survives, or only ranks finalists. My second is pinning: the judge model version, its decoding settings and the rubric text are frozen for the run, because a judge upgrade mid-campaign silently rebases the scale and makes generation five incomparable to generation one. Third is validation over time: keep an anchor set of frozen outputs with human labels, rescore it periodically, and watch for drift in agreement. I also avoid using the same model family as generator and judge where I can, since self-preference systematically favours candidates that write in the judge's own idiom.
go deeper
Know that using a model to score outputs costs a model call per scored item, and that those calls multiply quickly when a search evaluates many candidates against many examples.
Be able to do the arithmetic — candidates times examples times repeats — and to explain why the judge's model version, decoding settings and rubric must stay fixed for scores to be comparable across generations.
Show a concrete cost-control design: cascade cheap checks ahead of the judge, subsample per generation, cache repeats, and keep an anchor set for drift detection. Be ready to describe how you validated the judge against human labels before trusting it.
Own the call on whether a judge belongs in the inner loop at all. Frame it as coverage bought with cost, latency, noise and bias, argue for a judge-light architecture where a programmatic proxy is validated to track it, and set the campaign's versioning and re-validation policy.
## The judge is not a metric, it is a service In a hand-run evaluation a judge model is called a few hundred times and its cost is invisible. Inside an optimization loop it is called on every candidate, every example, every generation, and it becomes both the largest cost centre and the single point of failure for the objective. Reasoning about it the way you would reason about a production dependency — capacity, versioning, validation, failure modes — is what separates a campaign that produces a trustworthy prompt from one that produces a number nobody can interpret. This reflects how teams run these campaigns as of mid-2026; the economics shift as inference prices fall, but the structural concerns do not. ## The arithmetic Candidates × examples × repeats = judge calls per generation. Forty candidates over 250 examples is 10,000 calls; three repeats to reduce variance makes it 30,000; twenty generations makes it 600,000. Each call carries the graded output plus a rubric, so input tokens are not small either. Latency compounds the same way and often matters more than money — a generation that takes six hours changes how many experiments a team can run in a quarter. The levers, roughly in order of payoff: - **Cascade.** Run every free check first — schema validity, required fields, length caps, exact-match on the subset of examples that have one right answer — and let those gate or partially score. Reserve the judge for the open-ended residue, or for ranking only the top handful of candidates. This routinely removes an order of magnitude. - **Subsample per generation.** Judge each candidate on a stratified subset and only put finalists through the larger set. - **Pairwise only where it earns its cost.** Pairwise comparison is more reliable than absolute scoring but requires comparisons that grow with the candidate count. A practical split is pointwise scoring for the broad screen and pairwise for the final few. - **Right-size the judge.** A smaller judge validated against human labels on your specific rubric may agree as well as a frontier one at a fraction of the cost. That is an empirical question, answered once per campaign. - **Cache.** Identical candidate-output/example pairs recur across generations when candidates converge; caching by content hash is free savings. ## Comparability: pin everything The subtle failure is not cost, it is drift. A search loop compares scores *across* generations — that is how it knows it is improving. Every input to the judge that changes mid-run breaks that comparison: - **Model version.** Providers update models. If the judge silently changes, the scale rebases and your improvement curve becomes an artifact of the upgrade. Pin an explicit version for the campaign and treat changing it as starting a new campaign. - **Decoding settings.** Judge at the lowest available randomness so the same output gets the same score twice; otherwise part of every measured gap is sampling noise. - **The rubric.** Editing the rubric mid-run invalidates earlier scores just as thoroughly as swapping the model. If a rubric fix is genuinely needed, rescore the retained history under the new rubric or restart. When you must change the judge, an **anchor set** makes the change auditable: keep a fixed collection of frozen candidate outputs spanning the quality range, score them under both the old and new judge, and publish the mapping between scales. That anchor set doubles as your drift detector — rescoring it monthly shows whether the judge's behaviour has moved even with a pinned version. ## Trustworthiness A judge is a proxy for human preference, and inside a loop it is a proxy under sustained optimization pressure, which is the worst case for any proxy. - **Validate against humans.** Label a slice by hand and measure agreement between the judge and the human labels at the start of the campaign, and again periodically. If agreement is poor, no amount of search sophistication saves the result. - **Self-preference.** Judges favour text produced by their own model family and their own stylistic conventions. When the generator and the judge are the same model, the loop can select for house style rather than quality. Using a different family for judging, or validating that the ranking survives a second judge, mitigates it. - **Verbosity and position effects.** Judges tend to reward longer, more structured answers and, in pairwise mode, are sensitive to ordering. Length penalties and order randomization are the standard counters. - **Injection.** The graded text is candidate-controlled, so the rubric must instruct the judge to score against the criteria and ignore claims made inside the text about its own correctness. ## The strategic call The honest framing for a lead is that a judge in the inner loop buys coverage of qualities nothing else can measure, at the cost of a slower, pricier, noisier and biased objective. Many mature setups end up judge-light: programmatic scoring drives the search, and the judge is used at checkpoints to confirm that the programmatic proxy has not drifted from what humans want. Deciding where on that spectrum a given campaign should sit — and being willing to say the judge is not worth its cost for this task — is the judgment being tested.
- When would you drop the judge from the inner loop entirely?When a programmatic scorer correlates well enough with the judge on your task to drive the search. Measure that correlation once, then let the cheap scorer rank candidates every generation and bring the judge in only at checkpoints and for finalists. It cuts cost by orders of magnitude and removes a biased, noisy signal from the tightest feedback path.
- How would you notice that the judge itself has drifted mid-campaign?Keep an anchor set of frozen candidate outputs spanning the quality range and rescore it on a schedule. Scores on frozen text should not move; if they do, the judge or its serving stack changed, and every cross-generation comparison since then is suspect. The same anchor set lets you map old scores onto a new judge's scale if you deliberately upgrade.
- Why is using the same model family as both generator and judge risky here?Self-preference. Judges systematically rate text from their own family and stylistic conventions higher, so the search selects for house style rather than task quality — and because the loop optimizes hard, that bias compounds over generations. Use a different family for judging, or at minimum confirm that the final ranking holds under a second, independent judge.
- How do you keep the judge from being manipulated by the text it grades?Score against explicit criteria and a reference rather than open-ended quality, and instruct the judge to disregard any assertions inside the graded text about its own correctness. Keep programmatic gates that no framing can pass, and review the top candidates' outputs — a candidate addressing the grader is obvious to a human and invisible in the aggregate score.
saying these in an interview costs you the question
- Ignoring that judge cost scales with candidates times examples
- Letting the provider's default model version float during a run
- Editing the rubric mid-campaign and comparing scores across it
- Judging with the same model that generated the candidates
- Never checking judge agreement against human labels