skip to content

In automatic prompt optimization, how do you pick the scoring metric for a task?

level: middleimportance: must knowfreq 70%

answer

  1. The metric is the objective
  2. Pick by output shape
  3. If it runs, execute it
  4. Overlap and judges are proxies
  5. Gates for constraints, grade on top

basics

~20 s

Match the scorer to the output shape: exact match or F1 when one label is right, execution-based checks when the output can be run and verified, overlap scores only as a rough proxy, and a judge model only for genuinely open-ended text.

solid answer

~50 s

The scoring function is the objective, so the search will only ever optimize what it measures. I pick it from the shape of the output. If there is a single correct label or span, exact match or accuracy works; if classes are skewed or the answer is a set, precision/recall/F1 on the class I care about. If the output is executable or checkable — SQL, a regex, JSON against a schema, code — I run it against a fixture suite and score pass rate, which is the most objective and cheapest signal available. Overlap scores like ROUGE are usable for summaries but only weakly track human preference, so I treat them as a coarse screen. A judge model is the last resort, for open-ended text where nothing programmatic exists. In practice I often use a composite: hard constraints as gates (parses, obeys format, no forbidden content) and a graded metric on top.

code

python · 20 lines
python
import re

FIXTURES = [
    ("  Acme,  Inc. ", "Acme, Inc."),
    ("Globex   Ltd", "Globex Ltd"),
]

def score_candidate(pattern: str, replacement: str) -> float:
    try:
        compiled = re.compile(pattern)
    except re.error:
        return 0.0
    passed = sum(
        1
        for raw, expected in FIXTURES
        if compiled.sub(replacement, raw).strip() == expected
    )
    return passed / len(FIXTURES)

print(score_candidate(r"\s+", " "))

go deeper

for a junior

Know the basic metric families by name — accuracy, precision/recall/F1, overlap scores like ROUGE, and running the output against tests — and be able to say which one fits a classification task versus a summarization task.

for a middle

Be ready to justify a choice from the output shape and to explain the mechanics: what exact match normalizes away, why F1 exists, and why executing generated artifacts against fixtures is a stronger signal than comparing their text.

for a senior

Show that you treat the metric as the objective the search will exploit. Talk about hard-constraint gates, cost per evaluation, metric variance versus the effect size you need to detect, and how you would notice the metric drifting away from real quality.

for a principal

Own the tradeoff between measurability and validity: the most automatable metric is rarely the one the business cares about, and committing an optimization campaign to a cheap proxy sets the direction of the product for months. Be able to argue when to invest in a better construct instead of more search.

## The scorer is the objective An automatic prompt-engineering loop generates candidate prompts, scores each one on a dataset, keeps the winners and repeats. Whatever function produces that score *is* the definition of "better" for the entire run — the search has no other access to quality. Picking it is therefore the most consequential decision in the setup, and it is driven by the shape of the task's output rather than by preference. ## Families of scorer **Exact match and accuracy.** When each example has one correct answer — a label, a numeric result, a normalized span — you compare the candidate prompt's output to the reference after normalization (case, whitespace, punctuation, formatting). This is cheap, deterministic and un-gameable at the string level. Its weakness is brittleness: a right answer phrased differently scores zero, so normalization rules quietly become part of your metric. **Precision, recall and F1.** When the label distribution is skewed, or when the output is a set of extracted items rather than one value, per-class precision and recall are the honest measurement and F1 combines them into one number the search can rank on. Which of precision or recall matters more is a business decision, not a metric decision, and it is legitimate to optimize a weighted F-beta instead. **Overlap metrics.** ROUGE, BLEU and chrF compare n-grams between the candidate output and a reference text. They are cheap and fully automatic, which makes them tempting as a search objective for summarization or translation, but they measure surface overlap, not usefulness. On meeting-minutes summaries it is common for ROUGE rankings and human preference rankings over the *same* candidate prompts to disagree substantially, because a summary can share many n-grams with the reference while omitting the one decision that mattered. Use them as a coarse screen and know their bias: recall-oriented variants reward length. **Execution-based scoring.** If the output can be run or mechanically verified, run it. Generated SQL is executed against a fixture database and the result set compared; generated code is run against unit tests; generated JSON is validated against a schema; generated data-cleaning regexes are applied to a table of raw-to-expected string pairs and scored by how many pairs they normalize correctly. This is the strongest option when it exists: it is objective, costs no model calls, has no rater bias, and it scores the *behaviour* rather than the wording. It still has a hole — a candidate can special-case the fixtures instead of generalizing — which is why the fixture suite must contain examples the search never sees. **Judge models.** For open-ended text where nothing programmatic applies, an LLM scores the output against a rubric. It is the most flexible and the most expensive scorer, and it introduces the scorer's own errors into the objective. ## Composite objectives Real setups rarely use one number. A common structure is a gate plus a grade: hard constraints (output parses, matches the required schema, stays under a length cap, contains no disallowed content) are pass/fail and zero out a candidate; a graded metric then ranks the survivors. Weighted sums of several metrics are also used, but every weight you add is another dial the search will exploit, so keep the combination small and justify each term. ## Cost and noise are part of the choice The scorer runs once per candidate per example, so it is invoked thousands of times in a campaign. A programmatic scorer costs microseconds and lets you evaluate widely; a judge costs a model call and constrains how many candidates you can afford. Noise matters equally: if the metric is high-variance — a judge at nonzero temperature, or a sampled generation on a small set — the search will happily promote candidates whose apparent lead is sampling noise. Deterministic scoring (fixed decoding where the provider allows it, fixed fixtures) and adequate example counts are how you keep the signal above the noise floor. ## Failure modes to name The metric can be *valid but noisy* (right construct, too few examples), *invalid but stable* (a clean number that measures the wrong thing, like ROUGE for decision-capture), or *gameable* (a candidate raises the score without improving the task). These are three different diseases with three different treatments: more data, a better construct, and adversarial held-out checks respectively. A strong answer distinguishes them instead of saying "use a better metric". ## What interviewers listen for They want you to reach for the most objective scorer the task admits, to be explicit that overlap and judge scores are proxies, to mention hard constraints as gates, and to acknowledge that the metric you optimize hard against will eventually be the metric you get — including its defects.

  • What do you do when a task has one mechanically checkable part and one open-ended part?
    Score them separately and combine deliberately. The checkable part becomes a hard gate or a programmatic sub-score (schema valid, figures match the source), and the open-ended part gets a graded score on the survivors. Keeping them separate lets you see which half a candidate improved, and it stops a strong score on the cheap half from masking regression on the expensive half.
  • Why is an execution-based scorer usually preferred when it is available?
    Because it measures behaviour rather than wording: it is deterministic, costs no model calls, has no rater bias, and it cannot be won by rephrasing. That lets you evaluate far more candidates per unit budget. The caveat is that a candidate can special-case the fixtures, so the suite needs held-out cases and enough variety that memorizing it is harder than solving the task.
  • How does the scorer choice constrain how many candidates you can search over?
    Directly. Programmatic scorers are effectively free, so you can score hundreds of candidates against thousands of examples and let a broad search run. A judge costs one model call per candidate per example, which usually means a narrower candidate pool, fewer examples, or a cascade where a cheap scorer screens and the judge only ranks finalists.

saying these in an interview costs you the question

  • Reaching for a judge model when a deterministic check exists
  • Treating ROUGE as a measure of summary usefulness
  • Optimizing overall accuracy without checking per-class behaviour
  • Assuming a metric that is easy to compute is the right construct
  • Combining many weighted metrics without asking what the search will exploit

context