skip to content

When does weighting self-consistency votes by model confidence beat a plain count?

level: middleimportance: nice to knowfreq 28%

answer

  1. one vote becomes one weight
  2. the weight usually comes from likelihood
  3. strongest on short, closed answers
  4. confidence must track correctness
  5. confidently wrong samples get amplified

basics

~20 s

Weighted voting helps mainly when the sample count is small and the answer space is short and constrained, so a per-sample likelihood score is meaningful. It hurts when confidence is miscalibrated, because it lets a few confidently wrong samples outvote a correct plurality.

solid answer

~50 s

In confidence-weighted voting each sample contributes a weight rather than one vote, and the bucket with the highest summed weight wins. The usual weight is the model's own likelihood for the answer span — the mean log-probability of the answer tokens, exponentiated — which works best on short, closed-form answers such as a four-option compliance question where the answer is a single token. The claimed benefit is sample efficiency: with few samples, a plain count is noisy, and weighting can recover the right answer when the count ties. The caveats are real. Token likelihoods are not always exposed by hosted models. Mean log-probability is length-biased, so longer answers score lower for reasons unrelated to correctness. And model confidence is imperfectly calibrated, especially where the model is systematically wrong — exactly the case where you needed the vote to save you. Reported gains over plain plurality are modest and workload-dependent, so treat it as a tuning option, not a default.

code

python · 11 lines
python
import math
from collections import defaultdict

samples = [("B", -1.6), ("B", -1.5), ("B", -1.7), ("A", -0.05), ("A", -0.10)]

plain, weighted = defaultdict(int), defaultdict(float)
for answer, mean_logprob in samples:
    plain[answer] += 1
    weighted[answer] += math.exp(mean_logprob)

print(max(plain, key=plain.get), max(weighted, key=weighted.get))

go deeper

for a junior

Know that the basic rule is one sample, one vote, and that a variant weights each sample by how confident the model was. Be able to say that this only helps if confidence actually tracks correctness.

for a middle

Explain where the weight comes from — mean log-probability of the answer tokens — and name the three failure modes: miscalibration, length bias, and unavailability of log-probabilities from the serving stack.

for a senior

Show how you would decide: replay both rules over stored samples, compare accuracy and calibration by score bucket, slice by difficulty, and prefer weight-as-tie-break so a confidently wrong minority cannot overturn a correct plurality.

for a principal

Own the explainability tradeoff. A plain count is defensible to a reviewer or auditor in one sentence; a weighted score is not, and in regulated decisions that cost can outweigh a small accuracy gain.

## What changes Plain aggregation gives every sample one vote and returns the largest bucket. Confidence-weighted aggregation gives sample *i* a weight *w_i* and returns the bucket with the largest summed weight. Everything else — extraction, canonicalisation, tie rules — is unchanged. The only design question is where *w_i* comes from and whether it carries real information. ## Sources of a weight - **Token likelihood of the answer span.** Take the mean log-probability of the tokens that form the final answer and exponentiate it. This is the most defensible weight because it is an internal quantity, not something the model narrates. It requires the serving stack to expose per-token log-probabilities, which is not universal, and for models that hide their internal reasoning the answer span may be the only part you can score. - **Likelihood of the whole chain.** Weaker: it mixes the reasoning style's fluency into the score and is heavily length-dependent. - **Verbalised confidence.** Asking the model to state a percentage is easy and popular, but self-reported confidence clusters at round, high values and shifts with phrasing, so it is a poor weight and easy to inadvertently game with prompt wording. - **An external scorer.** A separate model that rates each candidate is a different technique — reranking — and is best kept conceptually apart from confidence weighting, because its failure modes differ. ## When weighting actually pays 1. **Small sample counts.** With a handful of samples, the count is a coarse, high-variance statistic and ties are common. A continuous weight breaks ties with information rather than an arbitrary rule. 2. **Short, closed answer spaces.** A single-token answer such as an option label on a multiple-choice compliance question has a clean, comparable likelihood. Length bias effectively disappears when every candidate answer is one token. 3. **Skewed difficulty.** When most questions are easy and a minority are hard, weighting concentrates the decision on the samples the model was actually sure about. ## When it backfires - **Miscalibration.** Weighting assumes confidence tracks correctness. Models are often confidently wrong on questions where they hold a wrong prior, and precisely there the weight amplifies the mistake: three high-confidence wrong samples can outweigh five low-confidence correct ones, converting a correct plurality into a wrong weighted winner. - **Length bias.** Mean log-probability falls with answer length for reasons of surface form, not truth, so on free-form answers weighting systematically favours the terser candidate. - **Availability.** Many hosted interfaces do not return log-probabilities, and where the visible output is a summary of the model's reasoning, the scores you can obtain may not correspond to the tokens that mattered. - **Opacity.** A plain count is trivially explainable to a reviewer — "twelve of twenty said this". A weighted score is not, which matters in regulated settings where the aggregation rule has to be defended. ## How to decide Treat it empirically. Hold a labelled evaluation set, run both rules over the *same* stored samples — this is a pure post-processing change, so no new generation is needed — and compare accuracy, and also compare calibration: bucket predictions by their winning score and check whether accuracy rises with the score. If the weighted rule is not better calibrated as well as more accurate, its extra complexity is buying nothing. A middle path is worth knowing: instead of replacing counts with weights, use the weight only to break ties, keeping the plurality rule intact. That captures the small-sample benefit without exposing the whole decision to a miscalibrated score. ## Relationship to abstention Whichever rule you use, keep reporting the plain agreement fraction as well. It is the number humans and downstream policies understand, and it remains the right input to an abstain-or-escalate decision even when the winner was chosen by weight.

  • Why is asking the model to state a confidence percentage a weak source of weights?
    Verbalised confidence is generated text, not an internal quantity. It clusters at round, high values, shifts with prompt wording and answer length, and is largely insensitive to whether the model is actually right. Using it as a weight imports that bias into the aggregation rule while looking principled. Where you need a self-reported signal, treat it as a coarse flag, not a multiplier.
  • How would you evaluate whether weighted voting is worth adopting?
    Store the samples once and replay both aggregation rules over them offline — it is post-processing, so no extra generation is required. Compare accuracy on a labelled set, but also compare calibration: bucket runs by winning score and check that accuracy rises monotonically with it. Adopt the weighted rule only if it wins on both, and slice by difficulty, since gains often come only from the easy tail.
  • Is there a version of this that keeps the safety of plain counting?
    Yes — use the weight only as the tie-break, leaving plurality as the primary rule. That captures the benefit where counting is genuinely uninformative, namely small sample counts and exact ties, without letting a miscalibrated score overturn a clear plurality. It also keeps the explanation simple, which matters when the decision has to be justified to a reviewer.

saying these in an interview costs you the question

  • Treats model confidence as a reliable probability of correctness
  • Uses verbalised percentages as weights without validation
  • Ignores that mean log-probability falls with answer length
  • Assumes log-probabilities are always available from the model
  • Adopts weighting without measuring it against plain counting

context