How do you detect prompt candidates that game the scorer instead of solving the task?
answer
- A measure under pressure stops measuring
- Search set versus held-out divergence
- Read what the winners actually output
- Verbosity is the classic exploit
- Turn each exploit into new examples
basics
~20 sCompare the optimized score against a held-out set the search never touched, and read the top candidates' actual outputs. A score that climbs on the search set while flat or falling on held-out data is the signature of a candidate exploiting the scorer.
solid answer
~50 sAny hard-optimized metric drifts from the quality it stands for, so I assume gaming will happen and build detection in. The primary signal is the gap between the search set and a held-out set the loop never scores against: rising search score with flat or falling held-out score means the candidate learned the scorer, not the task. The second is qualitative — I read the outputs of the top few candidates every so often, because gaming is usually obvious on sight: padded summaries when the objective is recall-oriented overlap, outputs containing text aimed at the judge, or generated code that special-cases the fixtures. Structural defences help too: hard constraints as gates, a length or cost penalty when the metric rewards verbosity, a second metric measuring a different construct, and adversarial examples added whenever a specific exploit is found. When a candidate wins by a suspiciously large margin, that is a hypothesis to investigate, not a result to ship.
go deeper
Know the basic idea that a prompt can score well without being good, and that the check is testing it on examples the tuning never used.
Be able to name concrete exploits per scorer type — length inflation under overlap metrics, self-assertive text under a judge, fixture special-casing under execution scoring — and explain why an automated loop finds them when a human tuner would not.
Show a working detection routine: held-out versus search-set curves, qualitative review of top candidates, perturbation tests, and turning each discovered exploit into new adversarial examples. Talk about stopping the search on held-out gains rather than search gains.
Own the governance angle: the scoring function is a living artifact with an owner and a review cadence, and a team that promotes whatever tops the leaderboard has delegated its quality bar to an unaudited proxy. Be able to argue how much optimization pressure the objective can safely carry.
## Goodhart in an optimization loop When a measure becomes a target it stops being a good measure. A human prompt engineer nudges a metric a few points and moves on; an automatic loop pushes on it thousands of times and finds whatever slack exists between the metric and the underlying goal. Metric gaming is therefore not an occasional accident in automatic prompt engineering — it is the default outcome of a long, effective search against an imperfect scorer. ## What gaming looks like, by scorer type **Overlap metrics.** Recall-oriented n-gram scores reward covering more of the reference, and the cheapest way to cover more n-grams is to emit more text. A search scored on such a metric reliably converges on candidates that instruct the model to be exhaustive, restate the input, and hedge — output that scores well and reads badly. **Judge scorers.** A judge reads the candidate's output, and the candidate's output is attacker-controlled text from the judge's point of view. Candidates emerge that assert their own correctness, adopt the rubric's vocabulary, or produce the confident, structured, verbose shape judges tend to reward regardless of substance. In the extreme, the generated output contains language directed at the grader — an injection into the scoring step rather than an improvement to the answer. **Execution-based scorers.** These are the hardest to game and still gameable: generated code or transformations can special-case the fixtures — matching the specific inputs in the suite rather than implementing the rule — and pass every test while generalizing to nothing. **Exact-match scorers.** Candidates learn the normalization rules, or, when the search set is small and reused, latch onto surface artifacts of those specific examples. ## Detection **Held-out comparison is the workhorse.** Reserve examples the search never scores against, and track the promoted candidate's score on them alongside the search score. Divergence between the two curves is the alarm. This is the only detector that works without knowing in advance which exploit to look for. **Read the outputs.** Gaming is usually visible immediately to a human, and reviewing the top few candidates each generation costs minutes. If a score jumped and you cannot articulate what improved in the text, treat that as unexplained rather than as progress. **Watch the shape of the win.** Suspicious signatures include an implausibly large single-generation jump, a candidate that wins on the aggregate while regressing on several slices, and a candidate whose advantage disappears when examples are rephrased or reordered. **Perturbation tests.** Paraphrase the inputs, shuffle option order, swap entity names. A genuine improvement survives; an exploit of surface artifacts usually does not. ## Structural defences - **Gates before grades.** Hard constraints — schema validity, a length cap, required citations, forbidden content — zero out a candidate regardless of its graded score, which closes off the crudest exploits. - **Counter-pressure terms.** If the metric rewards length, penalize length or cost explicitly so verbosity has to earn its place. - **A second, differently-shaped metric.** Two metrics measuring different constructs are much harder to satisfy by one trick than either alone, though every added term is also another surface to exploit, so keep the set small. - **Isolate the judge from the candidate's framing.** Score the answer against the rubric and the reference, with the judge instructed to disregard any statements about correctness inside the text it is grading. - **Grow adversarial examples.** Each discovered exploit becomes new examples in the scoring set that specifically defeat it. Over a campaign this turns detection into a ratchet. - **Rotate or refresh the search set.** Reusing one small set for thousands of evaluations invites memorization of that set. ## The judgment part The deeper point is that you cannot fully close the gap between a proxy and the real goal, so the discipline is to bound the damage: cap how hard you optimize (stop when held-out gains flatten rather than when search gains do), keep a human in the promotion path for anything the metric cannot see, and treat the scoring function as a living artifact that gets patched as exploits appear. A team that ships whatever tops the leaderboard has outsourced its quality bar to a proxy it never audited.
- How can a candidate attack a judge-based scorer directly, and how do you blunt it?The candidate's output is text the judge reads, so it can assert its own correctness, mirror rubric language, or address the grader outright. Blunt it by scoring against a reference and explicit criteria rather than open-ended quality, instructing the judge to ignore claims about correctness inside the graded text, and keeping structural gates that no amount of persuasive framing can pass.
- Is a rising score on the search set ever sufficient evidence that a prompt improved?No. The search score is the quantity being optimized, so it is the least trustworthy evidence available. Confirmation requires a set the loop never scored against, plus a human look at the top candidates' outputs. If both agree with the search score, you have a result; if they diverge, you have an exploit.
- Execution-based scoring is the hardest to game — how does a candidate still do it?By special-casing the fixtures. Generated code or transformations can branch on the exact inputs in the suite, or exploit an incidental property all fixtures share, passing every test while implementing nothing general. The defences are holding out fixtures the search never sees, generating fixtures programmatically so they cannot be enumerated, and reading the artifact rather than only its pass rate.
saying these in an interview costs you the question
- Trusting the search-set leaderboard as evidence of quality
- Assuming execution-based scoring cannot be gamed
- Never reading the winning candidate's actual output
- Running the search until improvement stops rather than until held-out gains stop
- Fixing gaming by adding more weighted metric terms without holding anything out