skip to content

Why normalize sampled answers before counting votes in self-consistency decoding?

level: middleimportance: must knowfreq 52%

answer

  1. string equality, not meaning
  2. same value, four different renderings
  3. the true answer splits across buckets
  4. canonical key before counting
  5. rounding too coarsely merges real disagreement

basics

~20 s

Without normalization, one correct answer written four different ways lands in four separate buckets and loses to a wrong answer that happened to be formatted consistently. Canonicalising each answer into a comparison key is what makes the vote count meaning rather than string formatting.

solid answer

~50 s

Voting compares answers by string equality, and language models express the same value many ways. On an insurance payout question, twenty samples might return `$4,200`, `4200.00`, `4,200 dollars` and "about four thousand two hundred" — all the same answer, split across four buckets of five, while a wrong `$3,800` appears six times and wins. Normalization maps each raw answer to a canonical key first: strip currency symbols, separators and units, parse to a number and round to an agreed precision, lower-case and trim text, map multiple-choice text back to its option letter, and use symbolic equivalence where 1/2 and 0.5 must count as one answer. The counter-risk is over-normalization — round too coarsely and genuinely different answers merge, which is a silent correctness bug rather than a formatting one. The cheapest half of the fix is upstream: constrain the output format so extraction is reliable in the first place.

code

python · 8 lines
python
import re

def canonical(raw):
    text = raw.strip().lower().replace(",", "")
    match = re.search(r"-?\d+(\.\d+)?", text)
    return f"{float(match.group()):.2f}" if match else text

print([canonical(x) for x in ["$4,200", "4200.00", "  4200 dollars "]])

go deeper

for a junior

Know that the vote compares answers as strings, so $4,200 and 4200.00 do not match unless you convert them first. Be able to name the fix: turn every answer into one canonical form before counting.

for a middle

Walk through the extract-then-canonicalise pipeline and give concrete rules — strip separators and units, parse and round numbers, map option text to its label, treat 1/2 and 0.5 as equal — and name over-normalization as the opposite failure.

for a senior

Demonstrate that you would instrument this: log raw versus canonical bucket distributions, track the parse-failure rate as an alert, unit-test the normalizer, and choose rounding precision from the decision the answer feeds rather than from convenience.

for a principal

Own the contract question — decide what counts as the same answer for your domain, since that threshold silently sets both the accuracy and the confidence signal, and push the fix upstream into a constrained output format rather than growing an ever-larger parser.

## The failure this prevents Aggregation counts buckets of *identical* answers. Identical means identical after whatever transformation you apply — and if you apply none, it means byte-identical strings. Free-running language model output is never byte-consistent, so the raw vote measures formatting agreement, not semantic agreement. Make it concrete. An insurance claim question is sampled twenty times. Fourteen chains reason correctly to a payout of four thousand two hundred, but they render it as `$4,200`, `4200.00`, `4,200 dollars`, and "about four thousand two hundred". Six chains make the same arithmetic slip and all render it plainly as `3800`. Counted raw, the buckets are 5, 4, 3, 2 and 6 — the wrong answer wins with the smallest true support in the run. Accuracy has been destroyed by string formatting, and nothing in the logs looks broken: the pipeline reports a clean winner with 30 percent agreement. This is the single most common implementation bug in self-consistency, and it is invisible in aggregate metrics unless you look at the bucket distribution. ## Extraction comes before normalization You cannot canonicalise what you did not find. Each sample must end in something parseable — a fixed marker line, a delimiter, or a small structured object. Two practical rules: - **Constrain generation.** Asking for the answer on a final line in a fixed shape is far cheaper and more reliable than a clever parser over free prose. - **Measure parse failures.** A sample whose answer cannot be extracted is a lost vote. Count those explicitly; a rising failure rate usually means the prompt drifted or a class of hard inputs makes the model hedge instead of committing. ## What normalization actually does Think of it as a function from raw answer to comparison key. Typical components: - **Numbers**: strip currency symbols, thousands separators, trailing units and words like "approximately"; parse to a number; round to a fixed, domain-chosen precision; normalise sign and leading zeros. - **Text**: trim, lower-case, collapse whitespace and punctuation, strip filler like "the answer is". - **Enumerations**: map an answer written as option text back to its option label, so "Non-compliant" and "B" are one bucket. - **Mathematical forms**: use symbolic equivalence where 1/2, 0.5 and 50% are the same answer, or where two algebraic expressions are equal. A symbolic engine or a canonical simplification step does this properly; regular expressions do not. - **Sets and lists**: sort, deduplicate, and decide whether order carries meaning before comparing. - **Dates and units**: convert to one representation (ISO dates, SI units) rather than matching surface forms. ## The opposite failure: over-normalization Every merge you perform is an assertion that two answers mean the same thing. Round a payout to the nearest thousand and 4,200 merges with 4,499 — you have now hidden a real disagreement and inflated your agreement fraction. Lower-casing merges case-sensitive identifiers. Stripping units merges 40 kilometres with 40 miles. Over-normalization is worse than under-normalization because under-normalization shows up as suspiciously fragmented buckets, while over-normalization looks like healthy consensus. The discipline is to choose precision from the domain: currency to cents, distances to the tolerance the task allows, and never coarser than the difference that would change a decision. ## Practical shape of a good implementation - Normalize into a key, but keep the raw answer of the winning bucket to return to the user, so you do not hand back a stripped, reformatted string. - Make normalization pure and unit-tested — it is ordinary code, and it is where correctness leaks. - Log both the raw and canonical distributions for a sample of traffic. Comparing them is how you discover you were splitting or merging. - When answers are structured (several fields), normalise per field and vote per field only if the fields are genuinely independent; otherwise vote on the whole tuple, since field-wise voting can assemble a combination no sample ever produced. ## Where this stops working Normalization assumes answers are comparable at all. Once the output is a paragraph, a summary or a block of code, no canonicalisation makes two of them equal, and exact-match voting has to be replaced by a selection method that judges semantic agreement instead of string identity.

  • How would you tell from production logs that your normalization is splitting votes?
    Log the bucket distribution, not just the winner. Split votes look like many small buckets with a low agreement fraction and, on inspection, several buckets that a human reads as the same answer. Comparing the raw-string distribution against the canonical-key distribution on sampled traffic makes it obvious: if canonicalisation barely reduces the bucket count on numeric tasks, the normalizer is not doing its job.
  • When the answer is a structured object with several fields, do you vote per field or on the whole object?
    Vote on the whole tuple by default. Field-wise voting can produce a combination that no single sample generated and that is internally inconsistent — for example a payout amount from one chain and a policy clause from another that contradicts it. Vote per field only when the fields are genuinely independent quantities, and validate the assembled result afterwards.
  • What is the cheapest way to reduce normalization work altogether?
    Constrain the output at generation time. Instructing the model to end with a fixed marker and a bare value — or to emit a small structured object with a typed numeric field — removes most formatting variance before it is ever produced, leaves the normalizer with only genuine equivalence cases such as 1/2 versus 0.5, and cuts parse failures at the same time.

saying these in an interview costs you the question

  • Assumes the model returns a consistent answer format
  • Compares raw strings and calls formatting variance disagreement
  • Rounds aggressively without checking what it merges
  • Treats a parse failure as just a lost sample worth ignoring
  • Thinks normalization is only needed for numbers

context