skip to content

What is self-consistency sampling, and when does it beat an iterative critique pass?

level: middleimportance: should knowfreq 45%

answer

  1. Vote, do not revise
  2. Errors scatter, truth concentrates
  3. Needs comparable answers to count
  4. Temperature zero gives no diversity
  5. Vote spread is a confidence signal

basics

~20 s

Self-consistency samples the same prompt several times at non-zero temperature and returns the answer the samples agree on most often. It beats iterative critique when the task has one comparable answer, because agreement is a real signal that needs no critic and the samples run in parallel.

solid answer

~50 s

Self-consistency replaces sequential revision with parallel voting: run the prompt N times independently, extract the final answer from each, and take the mode. It works because independent reasoning paths that reach the same conclusion are more likely to be right, while errors scatter — a genuine signal that costs no critique step. Two conditions must hold. The answers must be **comparable**, so you can tell that two of them are the same: a number, a label, a chosen option, a normalized value. And the samples must be **independent**, which means non-zero temperature; at temperature zero you get the identical answer N times and learn nothing. It suits a multi-step unit conversion or a classification far better than a two-page report, where no two samples are ever "the same answer". Cost is N full generations, but they are parallel, so latency is roughly one call rather than N sequential rounds. Vote spread also gives you a usable confidence estimate.

code

python · 8 lines
python
from collections import Counter

def self_consistency(sample_answer, prompt, n=5, temperature=0.8):
    votes = Counter(
        sample_answer(prompt, temperature=temperature) for _ in range(n)
    )
    answer, count = votes.most_common(1)[0]
    return answer, count / n

go deeper

for a junior

Know that self-consistency means asking the same question several times and keeping the answer that comes back most often, and that it needs an answer you can compare.

for a middle

Explain why independent errors scatter while correct answers concentrate, why temperature must be above zero, and why prose output has no equivalence test to vote over.

for a senior

Weigh it against a critique loop on cost and latency — N parallel generations versus sequential rounds — and use the vote distribution as a calibrated confidence signal rather than discarding it.

for a principal

Decide where in a pipeline the cost multiple is worth paying, and recognize that voting hardens against stochastic error only, leaving systematic model bias to be caught by an external check.

## The mechanism Self-consistency is a decoding strategy rather than a loop. Instead of one greedy generation, you sample the same prompt several times with temperature above zero so each run explores a different reasoning path, extract the final answer from each run, and return whichever answer appears most often. The reasoning paths themselves are discarded; only the answers are compared. The intuition is straightforward. There are many ways to reason correctly to the same conclusion and many more ways to reason incorrectly, but incorrect paths tend to land on *different* wrong answers. So the correct answer accumulates votes while errors spread thin. This is aggregation over independent samples, and it is the same reason an average of noisy measurements beats a single measurement. ## The two hard preconditions **Comparability.** You must be able to decide that two samples produced the same answer. A number, after normalizing units and precision. A label from a fixed set. A yes or no. A chosen option. Free-form prose has no equivalence test — two summaries of the same document are never literally identical, so there is nothing to count. Applying self-consistency to open-ended output either does nothing or requires a model to judge equivalence, which reintroduces the cost and the unreliability you were trying to avoid. **Independence.** Sampling at temperature zero returns the same output every time, so the votes are not evidence. You need enough temperature for genuine path diversity, but not so much that quality collapses. Note also that samples from one model are only conditionally independent — a shared misconception in the weights will show up in every sample, and the vote will be confidently wrong. Self-consistency corrects *stochastic* errors, not *systematic* ones. ## Against iterative critique | | Self-consistency | Iterative critique | |---|---|---| | Structure | N independent samples, parallel | Sequential draft, critique, revise | | Signal | Agreement between samples | A critic's findings | | Suits | One comparable answer | Rich artifacts, structured output | | Latency | About one call | N rounds, additive | | Cost | N generations | 2 or more calls per round | | Extra output | Vote distribution as confidence | Findings list, audit trail | They are not rivals so much as tools for different shapes of problem, and they compose: sample several candidate patches in parallel, verify each with the test suite, keep the ones that pass, and revise only from there. ## Choosing N and reading the votes Small odd values are typical — five is a common working point, with gains flattening well before twenty. The right N depends on how spread the answers are: a task where five samples agree unanimously needs no more, and a task where five samples give five different answers will not be rescued by fifty. That spread is itself useful output. Unanimity is a reasonable confidence signal; a three-two split is a flag to escalate, ask for clarification, or fall back to a deterministic path. Many teams get more value from the distribution than from the winning answer, because it turns an unhelpfully confident single answer into a calibrated one. ## Where it fails - **Systematic bias**: every sample makes the same wrong assumption, and the vote ratifies it with high apparent confidence. This is the dangerous failure, because the vote spread looks reassuring. - **Cost blindness**: N generations of a long-context prompt is N times the input tokens too, which on large prompts dominates and can make the technique unaffordable. Prompt caching mitigates the shared prefix but not the output cost. - **Fake comparability**: normalizing sloppily — treating "5 kg" and "5000 g" as different answers, or two differently-worded labels as distinct — splits votes across what is really one answer and produces a meaningless mode. - **Ties**: with an even N or a genuinely split distribution you need a defined tie-break, and defaulting to the first sample quietly discards the whole benefit. ## Practical placement Use it on the step where correctness is brittle and the answer is small, not on the whole pipeline. In a longer workflow, that often means sampling one extraction or one calculation step several times while the rest of the chain runs once, which keeps the cost multiple confined to where it buys something.

  • Why does self-consistency need a temperature above zero?
    Because the signal is agreement between independent reasoning paths. At temperature zero decoding is deterministic, so N samples are one sample repeated N times and the vote carries no information. You need enough temperature for genuine path diversity, while staying below the point where per-sample quality degrades enough to poison the vote.
  • Can self-consistency correct an error the model always makes?
    No. It aggregates away stochastic variation, not systematic bias. If a misconception is baked into the weights, every sample reproduces it and the vote ratifies the wrong answer unanimously — which is the dangerous case, because the spread looks like high confidence. Only an external check catches that class of error.
  • How can self-consistency and a verifier loop be combined?
    Sample several candidates in parallel, run the verifier on each, and discard the ones that fail. If several pass, take the modal answer or the cheapest; if none pass, revise from the closest candidate using its failure output. This uses parallelism to widen the search and the verifier to score it, rather than relying on votes alone.
  • What does a split vote tell you?
    That the model has no stable answer for this input. Treat it as a confidence signal: escalate to a human, ask a clarifying question, fall back to a deterministic path, or return the answer with its uncertainty attached. Many teams find the distribution more valuable than the winning answer, because it calibrates output that would otherwise be uniformly confident.

Rather than asking one accountant to double-check their own arithmetic, hand the same figures to five and see which total most of them reach independently.

saying these in an interview costs you the question

  • Runs self-consistency at temperature zero and expects diversity
  • Applies majority voting to free-form prose output
  • Treats a unanimous vote as proof of correctness
  • Ignores that N samples multiply input tokens as well as output
  • Confuses voting over samples with a critique-and-revise loop

context