skip to content

When does verifier-based best-of-n reranking beat majority voting over samples?

level: seniorimportance: should knowfreq 42%

answer

  1. voting needs no judge
  2. one good candidate can beat the crowd
  3. checking is cheaper than solving
  4. the mode is wrong on hard problems
  5. scoring harder eventually finds the scorer's flaws

basics

~20 s

Reranking wins when a scorer can recognise a correct answer that only one sample produced — hard problems where the modal answer is wrong, or answer spaces too open for counting. Voting wins when no trustworthy verifier exists, because it needs none.

solid answer

~60 s

Majority voting is verifier-free: it assumes the correct answer is the most frequent one, which is a good assumption only while the model is right more often than it is wrong in any single consistent way. Best-of-n reranking scores every candidate independently — with a trained reward model, or better, a deterministic check — and returns the top scorer even if it stood alone. On a probate inheritance split you can *check* candidates programmatically: do the shares sum to the estate, are they non-negative, do the fractions match the will's stated proportions. That kind of verifier dominates voting, because it can rescue the single correct answer from nineteen plausible wrong ones. Learned reward models are weaker: they are noisy and hackable, and pushing n high enough turns the search into optimisation against the scorer's flaws, so measured score rises while true quality falls. The robust middle ground is verifier-weighted voting — score each candidate but sum scores per answer bucket — which keeps the ensemble's variance reduction while using the verifier's signal.

code

python · 5 lines
python
def verifier(shares, estate):
    return abs(sum(shares) - estate) < 0.01 and all(s >= 0 for s in shares)

candidates = [[50000.0, 30000.0, 20000.0], [60000.0, 30000.0, 20000.0]]
print([verifier(c, 100000.0) for c in candidates])

go deeper

for a junior

Know the distinction: voting picks the most common answer and needs no judge, while reranking scores each candidate and picks the best, so it needs something that can tell good from bad.

for a middle

Explain when each bet pays — reranking when a lone correct candidate exists or answers are unbounded, voting when no trustworthy verifier is available — and note that verification is often much cheaper than generation.

for a senior

Demonstrate the design judgement: hunt for a deterministic check before reaching for a reward model, name overoptimisation as the failure mode of large n against a learned scorer, and propose verifier-weighted voting or verify-then-vote as the robust default.

for a principal

Own the accountability angle — a learned scorer becomes an optimisation target and a thing you must monitor, retrain and defend, so weigh the accuracy gain against owning a second model in the decision path of a consequential outcome.

## Two different bets Both methods start from the same place: several sampled answers, one to return. - **Voting bets on frequency.** No external judgement is needed. The premise is that correct reasoning converges and incorrect reasoning scatters, so the mode is the answer. - **Reranking bets on a scorer.** Each candidate is evaluated on its own merits and the best-scoring one wins, regardless of how many others agreed. The entire choice comes down to one question: do you have something that recognises a correct answer better than frequency does? ## The verifier spectrum **Deterministic checks** are the strongest and most underused. Many domains admit a cheap correctness or plausibility test that is not a model at all: unit tests for generated code, a constraint check on an allocation, dimensional analysis, re-substitution of a solution into the original equation, a schema or type check, a database lookup confirming a cited entity exists. On a probate inheritance split, verifying that the shares sum to the estate and respect the stated fractions eliminates most wrong candidates outright. These checks are sound where they apply — they never reward a wrong answer for looking good — though usually incomplete, since passing the check does not prove the interpretation of the will was right. **Learned scorers** — a reward model trained to rate final answers — apply where no deterministic test exists. They are useful and noisy in equal measure, and their errors are correlated with the generator's, because both learned from similar data. **The generator judging itself** is the weakest option. A model asked to score its own candidates tends to favour what it would have produced anyway, which recovers little beyond the vote. ## Where reranking clearly wins 1. **The mode is wrong.** On problems near the edge of the model's competence, the most frequent answer is often a shared, plausible mistake. Voting can never escape that; a verifier can, by recognising the lone correct candidate. 2. **The answer space is unbounded.** Where no two candidates are ever equal — generated code, a query plan, a document — buckets have size one and counting is meaningless, but a scorer still ranks. 3. **Checking is cheaper than solving.** The classic asymmetry. Verifying a proposed split, running a test suite, or validating a constraint costs a fraction of producing the answer, so the verifier is nearly free. 4. **Recall matters more than precision.** If you would rather find the right answer among many candidates than be safe, ranking beats counting. ## Where voting wins 1. **No verifier, or a bad one.** A scorer worse than frequency actively harms you. 2. **Small closed answer spaces.** With four options, counting is a strong, transparent estimator and reranking adds little. 3. **Explainability.** "Fourteen of twenty agreed" defends itself; "the reward model scored this 0.83" invites the question of what the reward model is. 4. **Adversarial or high-stakes settings.** Any learned scorer is a target, and optimising hard against it is exactly how you find its blind spots. ## Overoptimisation — the failure to name With a learned scorer, increasing n does not improve results indefinitely. Past some point the search stops finding better answers and starts finding candidates that exploit the scorer's imperfections: the measured score keeps rising while true quality, measured by humans or by ground truth, declines. This is reward hacking in miniature, and the practical consequence is that n for best-of-n has an empirical optimum you must find on held-out data — it is not "more is better". Deterministic verifiers are far more resistant, which is another reason to prefer them where the domain allows. ## The hybrid worth proposing Verifier-weighted voting keeps both signals: canonicalise candidates into buckets as usual, but sum verifier scores within each bucket instead of counting members. An answer that many samples reached *and* the verifier likes wins comfortably; a single candidate with a suspiciously high score cannot carry the decision alone. In practice this is more robust than pure argmax over scores, because it degrades gracefully when the scorer is noisy, and it still returns a meaningful agreement number. A further refinement is to use the verifier as a filter rather than a ranker: discard candidates that fail a hard constraint, then vote among the survivors. That combines a sound check with a transparent decision rule and is often the easiest version to defend to a reviewer. ## What to say in an interview Lead with the asymmetry — verification is often cheaper than generation, and where a real check exists it beats counting. Then be honest that most interesting tasks have no such check, that learned verifiers bring their own failure mode, and that the safe default in that case is voting, optionally weighted by whatever partial signal you trust.

  • Why can raising n in best-of-n with a learned reward model make results worse?
    More candidates means a harder search against the same imperfect scorer. Past an empirical optimum, the extra candidates that win are the ones exploiting the scorer's blind spots rather than genuinely better ones, so measured score climbs while true quality, judged by ground truth or humans, falls. Find the optimum on held-out data; deterministic verifiers are far more resistant to this.
  • How would you combine the two methods rather than choosing one?
    Two good hybrids. Use the verifier as a filter — discard candidates that fail a hard constraint, then vote among the survivors, which keeps a transparent decision rule. Or use verifier-weighted voting: canonicalise into buckets as usual but sum verifier scores within each bucket, so an answer needs both support and score to win and a single high-scoring outlier cannot carry the decision.
  • What makes a deterministic check a stronger verifier than a trained reward model?
    It is sound where it applies: an allocation that does not sum correctly, code that fails its tests, or a solution that does not satisfy the original equation cannot be scored highly for looking plausible. Its errors are also uncorrelated with the generator's, unlike a learned scorer trained on similar data. It is usually incomplete — passing does not prove correctness — so use it to eliminate, then vote.

saying these in an interview costs you the question

  • Assumes the most frequent answer is correct on hard problems
  • Treats a learned reward model as ground truth
  • Believes larger n always improves best-of-n results
  • Overlooks cheap deterministic checks the domain already offers
  • Uses the generating model to score its own candidates unquestioned

context