How do you apply self-consistency when the sampled answers are free-form text?
answer
- no two paragraphs are ever equal
- counting gives up, judging takes over
- all candidates in one prompt
- the winner is one of the candidates
- blandness can win a consistency contest
basics
~20 sExact-match voting is impossible on prose, so universal self-consistency puts all the sampled responses into one prompt and asks a model to pick the most consistent one. Selection replaces counting — the winner is still one of the candidates, never a new synthesis.
solid answer
~50 sWhen the output is a paragraph rather than a value, no canonicalisation makes two answers equal, so the counting step has to be replaced. **Universal self-consistency** concatenates the candidate responses into a single prompt and asks the model to select the one most consistent with the others — it is voting where the model, rather than a string comparison, judges agreement. It needs no answer extractor and works on summaries, explanations and code. A cheaper structural alternative is to embed or entail-compare the candidates, cluster them, and return the medoid of the largest cluster; a more invasive one is to decompose each response into atomic claims, vote on claims, and re-synthesise — which risks producing text no sample wrote. All of these share the ceiling of the basic method: selection cannot add a fact that every candidate missed. And the selector inherits judge-model biases — position and verbosity effects — plus a context limit, since every candidate must fit in one prompt.
go deeper
Know that voting needs answers you can compare, so plain counting fails on paragraphs. Be able to say that the model itself can be asked to pick the response most consistent with the others.
Explain universal self-consistency concretely — candidates in one prompt, the model returns an index, no extractor needed — and contrast it with clustering by similarity and returning the largest cluster's representative.
Show the operational gaps: the lost agreement number that your abstain policy depended on, the context limit when candidates are long, judge position bias, and the pull toward bland consensus answers on exactly the hard cases.
Own the framing decision — reshape the task so part of the output is extractable and can be voted on exactly, and confine judged selection to the narrative remainder, rather than accepting a weaker aggregation rule for the whole response.
## Why the standard recipe stops working Aggregation by counting rests on an equality test. For a number or an option label, canonicalisation makes that test meaningful. For eight free-text incident postmortem summaries, it does not: no two will be byte-identical, no normalizer will make them so, and every bucket has size one. Agreement between prose answers is semantic, and semantics needs a judge. ## Universal self-consistency The direct adaptation is to keep the *idea* of "pick the answer the others support" and change the *mechanism* from counting to model judgement. All candidate responses are placed in one prompt, labelled, and the model is asked which single response is most consistent with the rest. Its output is an index, not new text. The properties worth stating in an interview: - **No extractor.** Nothing has to be parsed out of the response, which is what makes it applicable to summaries, explanations, code and plans alike. - **Selection, not synthesis.** The winner is one of the sampled responses verbatim, so it is internally coherent and reads naturally. Nothing is stitched together. - **A second call.** You pay one extra inference over a prompt that holds every candidate, and long candidates multiply quickly against the context limit. On eight substantial postmortems this is the binding constraint, and it forces either truncation, a two-stage tournament, or fewer candidates. ## Alternatives worth knowing **Cluster-and-pick.** Represent each candidate as an embedding or compare them pairwise with an entailment check, cluster, and return the member of the largest cluster closest to its centre. This recovers something much closer to the original vote — the largest cluster *is* the bucket — and it gives you a numeric agreement measure, which the model-judge version does not. It is weaker when the candidates differ in ways the representation does not capture, for example two summaries that agree on wording but disagree on a root cause. **Claim-level voting.** Decompose each response into atomic factual claims, count claims across responses, keep the claims that most responses assert, and regenerate a response from the surviving claims. This is the only variant that can beat every individual candidate, because it can combine a fact from one with a fact from another. It is also the riskiest: decomposition is lossy, the re-synthesis step can assemble a combination no candidate endorsed, and consistency between claims is not checked by the vote. **Hybrid.** For structured free-form output — a postmortem with a fixed set of sections — you can vote exactly on the extractable fields (the identified root cause, the affected component, a numeric duration) and use universal self-consistency only on the narrative. This gives you a hard agreement number where one exists and a judged selection where it does not. ## Failure modes to name - **Shared blind spot.** If none of the eight postmortems identifies the real trigger, no selection method will produce it. Selection is bounded by the sample set. - **Judge bias.** A model choosing among candidates shows position effects and a preference for longer, more assertive text. Shuffling candidate order across repeated selections and checking whether the choice is stable is a cheap sanity check. - **Consensus toward blandness.** "Most consistent with the others" favours the response that commits to the least. A candidate that alone identifies a subtle cause is by construction the least consistent one, so this method can systematically discard the best answer on exactly the cases you cared about. - **No agreement number.** Unlike counting, the model-judge version returns a choice with no natural measure of how contested it was, which removes the abstain signal that makes the exact-match version safe to deploy. Recovering it — by clustering as well, or by asking the selector to rate agreement — is worth the extra work. ## The judgement to show The honest position is that free-form aggregation is weaker than exact-match aggregation, not merely different. Where you can force the task to yield an extractable answer — a score, a category, a decision — do that and vote on it, and treat prose selection as the fallback for the part of the output that genuinely resists it.
- What agreement signal do you lose by switching from counting to model-based selection, and how would you get it back?You lose the agreement fraction — the share of samples backing the winner — which is the input to any abstain or escalate policy. Recover it by also clustering the candidates semantically and reporting the largest cluster's share, or by asking the selector to report how many candidates it judged consistent with its pick. Without some such number, a contested run and a unanimous one look identical downstream.
- Why can selecting the most consistent free-text response sometimes discard the best one?Consistency with the others rewards the response that overlaps most with the consensus. A candidate that alone identifies a subtle root cause is, by definition, the least consistent — so the method systematically drops the insight it should have surfaced. This bias is strongest on hard cases where only a minority of samples got there, which is precisely where you were hoping for help.
- When would you decompose responses into claims and vote on those instead?When completeness matters more than fluency and no single candidate is expected to be complete — for example assembling a fact list where different samples each catch part of the picture. It is the only variant that can beat every candidate. Accept the costs: lossy decomposition, an unverified re-synthesis step, and the risk of assembling a combination no sample actually endorsed.
saying these in an interview costs you the question
- Thinks string normalization can make two paragraphs match
- Believes the selector merges the candidates into a better answer
- Ignores that all candidates must fit in one context window
- Forgets the selector shows position and verbosity bias
- Assumes selection can recover a fact no candidate contained