skip to content

In self-consistency, 19 of 20 chains agree on a wrong answer — what went wrong?

level: seniorimportance: must knowfreq 48%

answer

  1. variance falls, bias does not
  2. same weights, same prompt, same blind spot
  3. errors were never independent
  4. agreement is one-sided evidence
  5. decorrelate, or check from outside

basics

~20 s

Voting cancels independent errors, not shared ones. Samples from one model with one prompt share priors and a reading of the question, so a misread premise or a wrong prior propagates into every chain. Ensembling reduces variance; it cannot remove bias.

solid answer

~50 s

High agreement is evidence of low variance, not of correctness. Every sample came from the same weights, the same prompt and the same framing of the problem, so their errors are correlated by construction: if the model reads a quarterly figure as annual, all twenty chains inherit that reading and converge confidently on a wrong number. That is bias, and averaging does not touch bias. The practical consequences are two. First, the agreement fraction is an overconfident confidence signal — useful for detecting *contested* cases, useless for certifying agreed ones. Second, the fix is diversity along a different axis than sampling noise: reformulate the question, vary the prompt framing, ask for the problem to be restated before it is solved, or sample from a different model family, so that the shared assumption is not shared by every voter. Where correctness can be checked independently — recomputing, testing a constraint — do that instead of trusting consensus.

go deeper

for a junior

Know that all the samples come from the same model and prompt, so they can make the same mistake together. Agreement means the model was consistent, not that it was right.

for a middle

Explain it in ensemble terms: voting removes variance but not bias, and a misread premise or wrong prior is bias shared by every chain. Give a concrete shared-error source such as confusing a per-period figure with a total.

for a senior

Show what you would do: calibrate accuracy against agreement on your own data, escalate on low agreement, add an independent check whose errors are uncorrelated with the generator's, and decorrelate by reformulating the prompt or switching model family rather than raising the sample count.

for a principal

Own the risk framing — the dangerous output is the unanimous wrong one, since it reaches production unflagged. Decide which decisions may rest on model consensus at all, and where an independent verification path is a requirement rather than an optimisation.

## The statistical claim, stated precisely Self-consistency is an ensemble, and ensembles reduce the variance component of error. Formally, the benefit of averaging depends on the errors being at least partly independent. When they are perfectly correlated, averaging changes nothing at all. Sampling the same model with the same prompt at a non-zero temperature perturbs only the surface path the reasoning takes — it does not perturb what the model believes, how it parses the question, or which facts it has wrong. Whatever error lives in those shared components is bias, and it survives any number of samples. This is why the 19-of-20 case is not a bug in the aggregation code. The aggregation did exactly what it was designed to do, over a sample set whose errors were not independent. ## Where shared error actually comes from - **Misread premise.** The most common and most damaging. Units confused, a per-period figure taken as a total, a negation missed, an exclusion clause skipped. Every chain then solves a different problem than the one asked, correctly. - **A wrong prior or memorised fact.** The model believes something false about a rule, a rate or an API. No amount of resampling contradicts it. - **A prompt that steers.** A leading example, an anchor number in the instructions, or an option order that biases selection nudges every sample the same way. - **A surface-similar template.** The question resembles a common textbook form, and the model applies that template's solution rather than reading the details that make this instance different. Sampling explores wordings of the same template. - **Ambiguity resolved silently.** Where a question admits two readings, the model consistently picks one and never signals that a choice was made. ## What agreement can and cannot tell you Agreement is **asymmetric evidence**. Low agreement is a genuinely useful alarm: it means the model's own paths do not converge, and those cases correlate strongly with errors, so routing them to a stronger method, a different prompt or a human pays off. High agreement is much weaker evidence. It rules out the noisy-reasoning failure mode and rules out nothing else. The practical move is to calibrate rather than assume. On a labelled set from your own workload, bucket runs by agreement fraction and measure accuracy per bucket. You will typically find accuracy rises with agreement but plateaus well below 100 percent, and that plateau is the honest ceiling of consensus as a confidence signal. Publishing that number internally stops downstream teams treating unanimity as certainty. ## Breaking the correlation The cure is diversity in a dimension other than sampling randomness: - **Reformulate the question.** Sample across two or three paraphrases of the prompt. If the disagreement appears across paraphrases but not within them, you have found a framing-sensitive case. - **Force explicit restatement.** Requiring each chain to restate the given quantities and their units before solving surfaces the misread premise into the output, where a comparison across chains can catch it. - **Vary the approach.** Ask some chains to solve forward and others to verify a candidate answer backwards or to estimate an order of magnitude first. Different solution strategies fail differently. - **Vary the model.** Sampling from a different model family is the strongest decorrelation available, since the weights and training data differ. It is also the most operationally expensive and complicates deployment. - **Verify externally.** Any check that does not come from the same model — a constraint, a recomputation in code, a lookup against an authoritative source — breaks the correlation completely, because its errors are unrelated to the generator's. ## Designing around it Treat consensus as one input, not the decision. A defensible pipeline reports the winning answer, the agreement fraction, and the result of whatever independent check exists, and it escalates on low agreement *or* failed check. For consequential outputs, the failure mode you must guard against is not the visible disagreement — it is the confident, unanimous, wrong answer, because that is the one that reaches production without a flag. Sampling more paths makes it more expensive and no less likely. ## The honest interview answer Say plainly that self-consistency buys reliability against reasoning noise and nothing against systematic error, that agreement is a one-sided signal, and that the mitigation is decorrelation and independent verification rather than a larger sample count.

  • If more samples do not help here, is the agreement fraction useless as a confidence signal?
    No, but it is one-sided. Low agreement is a strong alarm — those cases correlate with errors and are worth escalating. High agreement only rules out noisy reasoning. Calibrate it on your own labelled data by bucketing runs by agreement and measuring accuracy per bucket; the plateau you find is the honest ceiling, and publishing it stops teams reading unanimity as certainty.
  • Which decorrelation technique gives the most benefit per unit of engineering effort?
    Usually forcing each chain to restate the given quantities, their units and any assumption it made before solving. It is a prompt change, and it converts a silent shared misreading into visible text that differs across chains or that a human or checker can inspect. Sampling paraphrases of the question is the next step up; switching model families is the strongest and the most operationally costly.
  • How does this change what you log for a self-consistency run in production?
    Log the full answer distribution rather than the winner alone, the agreement fraction, the parse-failure count, and the outcome of any independent check. That combination lets you distinguish a contested run from a unanimous one after the fact and lets you build the accuracy-versus-agreement calibration you need. A pipeline that records only the returned answer cannot tell a confident correct result from a confident wrong one.

Twenty people handed the same misprinted map will confidently agree on the same wrong turn; asking more of them does not fix the map.

saying these in an interview costs you the question

  • Says high agreement means the answer is probably right
  • Proposes more samples as the fix for a systematic error
  • Assumes sampled chains make independent mistakes
  • Thinks temperature alone produces genuinely diverse reasoning
  • Treats the agreement fraction as a calibrated probability

context