In self-consistency decoding, how is one final answer chosen from many sampled chains?
answer
- only the final answers are compared
- reasoning text is discarded
- biggest group wins
- no fifty-percent threshold required
- ties need a documented deterministic rule
basics
~20 sSelf-consistency throws away the reasoning text and votes only on final answers: each sampled chain contributes its extracted answer, identical answers are grouped into one bucket, and the biggest bucket wins. That is a plurality, not a required majority.
solid answer
~40 sAggregation runs as a small pipeline. First **extract** the final answer from each sampled chain — usually by asking the model to end with a fixed marker so a parser can find it. Second **canonicalise** each extracted answer so equivalent forms land in the same bucket. Third **count** the buckets and return the largest one. The winner only needs more votes than any other single answer, so with an open answer space a 6-of-20 bucket can legitimately win. Ties are broken by an explicit deterministic rule — first-seen, or highest per-sample score — never left to chance. The size of the winning bucket divided by the number of samples is a useful by-product: it is the agreement fraction, a rough signal of how contested the question was.
code
python · 8 linesfrom collections import Counter
def aggregate(answers):
counts = Counter(answers)
top, votes = counts.most_common(1)[0]
return top, votes / len(answers)
print(aggregate(["42", "42", "37", "42", "37"]))go deeper
Be able to say plainly that self-consistency extracts each sample's final answer, groups identical answers, and returns the largest group — and that this is a plurality, so the winner need not have more than half the votes.
Explain why the vote marginalises over reasoning paths: different valid derivations converge on one answer while wrong ones scatter. Walk through extract, canonicalise, count, select, and name the tie rule you would use.
Show the operational side: instrument the parse-failure rate, log the full answer distribution rather than just the winner, and define an abstain threshold on the agreement fraction so weakly-supported answers never reach a user unflagged.
Own the policy question — what agreement level is good enough for which decision class, what happens on abstain, and whether the extra samples buy enough reliability on your workload to justify running this path at all rather than a single chain plus a check.
## The idea in one line Self-consistency samples several independent reasoning chains for the same question and then keeps only one thing from each: the final answer. Aggregation is the step that turns that bag of answers into a single output. ## Why the vote is over answers, not over reasoning There are many valid ways to reach the same correct conclusion. Two chains can compute an interest figure by different routes, or split a bill in a different order, and still land on the same number. If you compared the reasoning text, those two chains would look like disagreement. By comparing only the final answers, you treat the reasoning as a latent variable — an internal working-out that is marginalised away — and you let structurally different derivations reinforce each other. That is the whole statistical point: correct paths tend to converge on one answer, while wrong paths tend to scatter across many different wrong answers, so the correct answer accumulates the largest bucket even when no single chain is trustworthy on its own. It also means the aggregation step never judges quality of prose, length, or how confident a chain sounds. A beautifully written chain and a terse one count exactly one vote each. ## The pipeline 1. **Extraction.** Each sample must yield a machine-comparable answer. In practice you instruct the model to end with a fixed marker ("Answer: <value>") or to emit a small structured object, then parse that. Free-text tails with no marker are the most common source of silent aggregation bugs, because a failed parse quietly drops a vote. 2. **Canonicalisation.** Equivalent answers must map to the same key before counting — otherwise the vote splits across formatting variants. 3. **Counting.** Group identical keys, count each group. 4. **Selection.** Return the answer of the largest group. 5. **Reporting.** Return the agreement fraction alongside the answer, so downstream code can act on weak consensus. ## Plurality, not majority A majority means strictly more than half the votes. Self-consistency does not require that, and demanding it would be wrong. On a free-form arithmetic or planning question, wrong answers scatter: you might see the correct value eight times out of twenty and eleven different wrong values once or twice each. Eight of twenty is 40 percent — no majority, but a clear plurality, and it is the right answer to return. Only on a small closed answer space (a four-option question, a yes/no) does the winner routinely clear 50 percent. If your product genuinely needs a majority-strength guarantee, you do not change the vote rule — you keep the plurality rule and add an abstain threshold: if the top bucket holds fewer than some fraction of the samples, refuse to answer or escalate rather than returning a weakly-supported result. ## Ties and other edge cases Ties happen, especially with a small number of samples and a binary or four-way answer space. Leave nothing to iteration order: pick a documented rule and apply it everywhere. Common rules are first-occurrence wins (stable and reproducible), highest-scoring member wins (needs a per-sample score), or escalate the tie to a second stage. Other edge cases worth handling explicitly: samples that produce no parseable answer (count them separately and report the parse-failure rate rather than silently discarding), samples that refuse or error out, and answers that are correct but out of the expected type (a string where a number was required). ## What you get for free Because counting produces a distribution and not just a winner, you also get the shape of the disagreement: how many distinct answers appeared, how big the runner-up was, and how large a share the winner took. A run where the top two buckets are 9 and 8 is a very different situation from one where they are 17 and 1, even though both return an answer. Logging that distribution is cheap and turns self-consistency from a black-box accuracy trick into something you can monitor and route on. ## What aggregation cannot do The vote is only as good as the samples. It cannot detect that every chain misread the same premise, and it cannot invent an answer that no chain produced — the winner is always one of the sampled answers. Aggregation is selection, not synthesis.
- Your parser fails to find an answer in three of twenty samples. What should the aggregator do with them?Exclude them from the vote but count and report them separately as parse failures, and compute the agreement fraction over the samples that actually voted. Silently dropping them hides a prompt or format bug that could be biasing which chains survive — for instance if long, hedged chains are the ones most likely to omit the marker. A rising parse-failure rate is an alert, not noise.
- Would it ever make sense to vote on the reasoning steps rather than the final answer?Not in self-consistency, whose whole premise is that many different valid derivations reach the same conclusion; comparing steps would treat that diversity as disagreement. Step-level scoring belongs to a different setting — searching over a branching tree of partial states, where you must decide which branch to expand before any final answer exists.
- With twenty samples, the top answer takes six votes and the runner-up five. Do you return it?Mechanically yes — it is the plurality winner — but 30 percent agreement with a near-tie is a weak result and should not be returned as if it were confident. Treat the agreement fraction as a routing signal: below a threshold you calibrate on your own data, abstain, escalate to a stronger method, or send the case to a human.
saying these in an interview costs you the question
- Says the winning answer must exceed fifty percent of votes
- Thinks the best-written or longest chain is selected
- Believes the vote compares reasoning steps across chains
- Assumes ties cannot happen and needs no rule
- Thinks aggregation can produce an answer no sample gave