Why does self-consistency break down if you sample the chains at temperature 0?
answer
- ensembles need uncorrelated members
- look at where chains branch
- wording variance is not path variance
- greedy decoding yields n identical traces
basics
~20 sTemperature 0 is greedy decoding: every sample follows the same highest-probability path, so the chains come back near-identical and the vote counts one opinion n times. Self-consistency needs stochastic sampling so chains take genuinely different reasoning routes.
solid answer
~50 sSelf-consistency is an ensemble, and an ensemble only helps when its members can make **different** mistakes. At temperature 0 the decoder takes the argmax token at every step, so all n chains retrace the same computation — I have seen forty traces on an orbital-mechanics word problem come back token-for-token identical apart from infrastructure jitter. You pay 40x and vote on one opinion. So you sample: temperature above 0 moves probability mass onto alternatives at branch points, and nucleus (top-p) truncation keeps the tail from producing junk. In practice a temperature band around 0.6–1.0 with top-p near 0.95 is a sane starting point, tuned on a labelled slice. The subtlety is that the diversity has to be **path** diversity, not paraphrase: chains that reword the same equations vote as one. Push temperature too high and per-chain quality collapses, so the curve is an inverted U.
code
python · 9 linesdef path_diversity(final_answers, step_signatures):
answer_spread = len(set(final_answers)) / len(final_answers)
step_spread = len(set(step_signatures)) / len(step_signatures)
return answer_spread, step_spread
answers = ["42", "42", "42", "38"]
steps = ["v2_over_r;42", "v2_over_r;42", "v2_over_r;42", "GM_over_r2;38"]
print(path_diversity(answers, steps))go deeper
Know that temperature 0 means greedy decoding — the model picks its top token every time — so repeated calls return the same chain and there is nothing to vote on. Be able to say self-consistency requires sampling.
Explain the mechanics: temperature rescales logits, top-p truncates the tail, and together they decide how often the chain takes a non-obvious branch. Be ready to distinguish paraphrase diversity from genuinely different reasoning routes.
Show you would measure it: distinct normalized answers and intermediate steps at fixed n, a temperature sweep on a labelled slice, and vote accuracy compared against plain greedy as the acceptance test before shipping the extra spend.
Own the tradeoff shape. Diversity and per-chain quality pull against each other, so the operating point is an empirical inverted U per model and task, and it must be re-measured on every model swap rather than inherited as a constant in a config file.
## What self-consistency asks of the sampler Self-consistency replaces a single greedy chain-of-thought with n independently sampled chains and keeps the final answer most of them agree on. The statistical bet is that a correct line of reasoning can be reached by several different routes, while wrong routes scatter — errors are only partially correlated, so the modal answer is more reliable than any one chain. Everything in that bet rests on the samples being **quasi-independent**. If the n chains are the same computation repeated, the vote carries exactly one bit of information and the extra spend buys nothing. ## Why temperature 0 degenerates Temperature rescales the logits before the softmax. As it goes to 0 the distribution collapses onto the argmax, which is greedy decoding: at every step the same token wins, so the whole trajectory is fixed by the prompt. Sample forty chains on the same orbital-mechanics word problem at temperature 0 and you get forty near-identical traces — same setup, same intermediate quantities, same arithmetic slip if there is one. The answer histogram is a single spike, the vote is trivially that spike, and the accuracy equals plain greedy chain-of-thought at forty times the token bill. A useful smoke test before trusting any self-consistency deployment: count distinct final answers at n=20. If it is almost always 1, the sampler is not doing its job. ## The knobs that actually create paths - **Temperature** decides how much probability mass reaches non-argmax tokens. It is the primary diversity control, because it operates at the decision points where the model chooses a strategy, sets up an equation, or picks a unit conversion. - **Nucleus (top-p) truncation** keeps only the smallest set of tokens whose cumulative probability reaches p and renormalizes over them. It bounds how bad a sampled token can be, which matters because raising temperature alone eventually admits incoherent continuations. Related truncation schemes such as typical sampling and min-p trade off the same way. The two interact: temperature says how flat the distribution gets, truncation says how far into the tail you are allowed to reach. Reported results for self-consistency are fairly robust across a band of temperatures rather than requiring a precisely tuned value, but the band is model- and task-specific — measure it. ## Surface diversity versus path diversity This is the distinction candidates most often miss. Two chains can differ in every sentence and still perform identical reasoning: same decomposition, same equation, same numbers. That is paraphrase, and it votes as a single chain. On an internal arithmetic word-problem set, comparing top-p 0.95 against typical sampling looked equally diverse when scored by wording variance, and clearly different when scored by distinct intermediate quantities and solution strategies. Wording varies on filler tokens, which are the cheapest ones to randomize; the tokens you want varied are the ones that commit the chain to a route. ## How to measure it - Number of **distinct normalized final answers**, or the entropy of the answer distribution, across n samples. - Number of **distinct intermediate-step signatures** — the sequence of equations, chosen operations, or retrieved facts, normalized before comparison. - Surface metrics such as edit distance or self-BLEU between rationales, treated only as a weak proxy; they cannot distinguish paraphrase from a different route. - The decisive check is downstream: vote accuracy on a labelled slice versus greedy accuracy. Diversity that does not move that number is not worth paying for. ## The quality–diversity tradeoff Raising temperature raises the chance that some chain finds the right route, and lowers the chance that any given chain is coherent. Vote accuracy is a product of the two, so it traces an inverted U: it improves over greedy up to some temperature, then falls as chains become sloppy. Sweep temperature at fixed n on a held-out labelled slice and take the peak, rather than importing a number from a paper written about a different model. ## Levers beyond the decoding knobs Prompt-level variation often produces more genuine path diversity than temperature does: shuffling few-shot exemplars, instructing different solution strategies per sample, or varying the output format pushes the chains apart at the decision points rather than at the adjectives. This also matters for reasoning-mode endpoints, where providers may restrict or ignore sampling parameters and instead expose a thinking budget — there you get diversity from prompt variation and from the model's own stochastic reasoning, not from a temperature dial you control. Check what the mode you are calling actually honours before assuming the knob exists. ## Common failure modes Sampling at temperature 0 or near it "for stability"; declaring diversity from eyeballing prose; cranking temperature past the coherence cliff and reading the resulting accuracy drop as evidence that self-consistency does not work; and never checking the distinct-answer count, which would have revealed the problem in one line of code.
- If the endpoint fixes temperature for you, what else can you vary across samples to get different reasoning routes?Vary the prompt rather than the decoder: shuffle few-shot exemplar order, instruct a different solution strategy or persona per sample, change the requested output format, or reorder the given facts. Prompt-level variation moves the chain at its decision points, which is exactly where you want divergence, and it works on reasoning-mode endpoints that restrict or ignore sampling parameters.
- How would you tell that a higher temperature is hurting rather than helping?Sweep temperature at fixed n on a labelled held-out slice and track two numbers: per-chain accuracy and vote accuracy. Diversity rises monotonically with temperature but per-chain quality falls, so vote accuracy traces an inverted U. The point where per-chain accuracy starts dropping faster than diversity buys coverage is your ceiling; beyond it you are voting on noise.
- Why is edit distance between rationales a weak diversity metric here?Because it measures wording, and wording varies most on filler tokens that do not commit the chain to anything. Two chains can be far apart in edit distance and still set up the same equation and produce the same answer, which votes as one member. Score distinct normalized intermediate steps and distinct final answers instead.
saying these in an interview costs you the question
- Claiming temperature 0 is fine because the prompt is re-processed each call
- Treating differently worded rationales as independent reasoning paths
- Assuming higher temperature always raises self-consistency accuracy
- Never counting distinct final answers before trusting the vote
- Copying a temperature from a paper without measuring on the target model