Why normalize sampled answers before counting votes in self-consistency decoding?
answer
- string equality, not meaning
- same value, four different renderings
- the true answer splits across buckets
- canonical key before counting
- rounding too coarsely merges real disagreement
basics
~20 sWithout normalization, one correct answer written four different ways lands in four separate buckets and loses to a wrong answer that happened to be formatted consistently. Canonicalising each answer into a comparison key is what makes the vote count meaning rather than string formatting.
solid answer
~50 sVoting compares answers by string equality, and language models express the same value many ways. On an insurance payout question, twenty samples might return `$4,200`, `4200.00`, `4,200 dollars` and "about four thousand two hundred" — all the same answer, split across four buckets of five, while a wrong `$3,800` appears six times and wins. Normalization maps each raw answer to a canonical key first: strip currency symbols, separators and units, parse to a number and round to an agreed precision, lower-case and trim text, map multiple-choice text back to its option letter, and use symbolic equivalence where 1/2 and 0.5 must count as one answer. The counter-risk is over-normalization — round too coarsely and genuinely different answers merge, which is a silent correctness bug rather than a formatting one. The cheapest half of the fix is upstream: constrain the output format so extraction is reliable in the first place.
code
python · 8 linesimport re
def canonical(raw):
text = raw.strip().lower().replace(",", "")
match = re.search(r"-?\d+(\.\d+)?", text)
return f"{float(match.group()):.2f}" if match else text
print([canonical(x) for x in ["$4,200", "4200.00", " 4200 dollars "]])go deeper
Know that the vote compares answers as strings, so $4,200 and 4200.00 do not match unless you convert them first. Be able to name the fix: turn every answer into one canonical form before counting.
Walk through the extract-then-canonicalise pipeline and give concrete rules — strip separators and units, parse and round numbers, map option text to its label, treat 1/2 and 0.5 as equal — and name over-normalization as the opposite failure.
Demonstrate that you would instrument this: log raw versus canonical bucket distributions, track the parse-failure rate as an alert, unit-test the normalizer, and choose rounding precision from the decision the answer feeds rather than from convenience.
Own the contract question — decide what counts as the same answer for your domain, since that threshold silently sets both the accuracy and the confidence signal, and push the fix upstream into a constrained output format rather than growing an ever-larger parser.
## The failure this prevents Aggregation counts buckets of *identical* answers. Identical means identical after whatever transformation you apply — and if you apply none, it means byte-identical strings. Free-running language model output is never byte-consistent, so the raw vote measures formatting agreement, not semantic agreement. Make it concrete. An insurance claim question is sampled twenty times. Fourteen chains reason correctly to a payout of four thousand two hundred, but they render it as `$4,200`, `4200.00`, `4,200 dollars`, and "about four thousand two hundred". Six chains make the same arithmetic slip and all render it plainly as `3800`. Counted raw, the buckets are 5, 4, 3, 2 and 6 — the wrong answer wins with the smallest true support in the run. Accuracy has been destroyed by string formatting, and nothing in the logs looks broken: the pipeline reports a clean winner with 30 percent agreement. This is the single most common implementation bug in self-consistency, and it is invisible in aggregate metrics unless you look at the bucket distribution. ## Extraction comes before normalization You cannot canonicalise what you did not find. Each sample must end in something parseable — a fixed marker line, a delimiter, or a small structured object. Two practical rules: - **Constrain generation.** Asking for the answer on a final line in a fixed shape is far cheaper and more reliable than a clever parser over free prose. - **Measure parse failures.** A sample whose answer cannot be extracted is a lost vote. Count those explicitly; a rising failure rate usually means the prompt drifted or a class of hard inputs makes the model hedge instead of committing. ## What normalization actually does Think of it as a function from raw answer to comparison key. Typical components: - **Numbers**: strip currency symbols, thousands separators, trailing units and words like "approximately"; parse to a number; round to a fixed, domain-chosen precision; normalise sign and leading zeros. - **Text**: trim, lower-case, collapse whitespace and punctuation, strip filler like "the answer is". - **Enumerations**: map an answer written as option text back to its option label, so "Non-compliant" and "B" are one bucket. - **Mathematical forms**: use symbolic equivalence where 1/2, 0.5 and 50% are the same answer, or where two algebraic expressions are equal. A symbolic engine or a canonical simplification step does this properly; regular expressions do not. - **Sets and lists**: sort, deduplicate, and decide whether order carries meaning before comparing. - **Dates and units**: convert to one representation (ISO dates, SI units) rather than matching surface forms. ## The opposite failure: over-normalization Every merge you perform is an assertion that two answers mean the same thing. Round a payout to the nearest thousand and 4,200 merges with 4,499 — you have now hidden a real disagreement and inflated your agreement fraction. Lower-casing merges case-sensitive identifiers. Stripping units merges 40 kilometres with 40 miles. Over-normalization is worse than under-normalization because under-normalization shows up as suspiciously fragmented buckets, while over-normalization looks like healthy consensus. The discipline is to choose precision from the domain: currency to cents, distances to the tolerance the task allows, and never coarser than the difference that would change a decision. ## Practical shape of a good implementation - Normalize into a key, but keep the raw answer of the winning bucket to return to the user, so you do not hand back a stripped, reformatted string. - Make normalization pure and unit-tested — it is ordinary code, and it is where correctness leaks. - Log both the raw and canonical distributions for a sample of traffic. Comparing them is how you discover you were splitting or merging. - When answers are structured (several fields), normalise per field and vote per field only if the fields are genuinely independent; otherwise vote on the whole tuple, since field-wise voting can assemble a combination no sample ever produced. ## Where this stops working Normalization assumes answers are comparable at all. Once the output is a paragraph, a summary or a block of code, no canonicalisation makes two of them equal, and exact-match voting has to be replaced by a selection method that judges semantic agreement instead of string identity.
- How would you tell from production logs that your normalization is splitting votes?Log the bucket distribution, not just the winner. Split votes look like many small buckets with a low agreement fraction and, on inspection, several buckets that a human reads as the same answer. Comparing the raw-string distribution against the canonical-key distribution on sampled traffic makes it obvious: if canonicalisation barely reduces the bucket count on numeric tasks, the normalizer is not doing its job.
- When the answer is a structured object with several fields, do you vote per field or on the whole object?Vote on the whole tuple by default. Field-wise voting can produce a combination that no single sample generated and that is internally inconsistent — for example a payout amount from one chain and a policy clause from another that contradicts it. Vote per field only when the fields are genuinely independent quantities, and validate the assembled result afterwards.
- What is the cheapest way to reduce normalization work altogether?Constrain the output at generation time. Instructing the model to end with a fixed marker and a bare value — or to emit a small structured object with a typed numeric field — removes most formatting variance before it is ever produced, leaves the normalizer with only genuine equivalence cases such as 1/2 versus 0.5, and cuts parse failures at the same time.
saying these in an interview costs you the question
- Assumes the model returns a consistent answer format
- Compares raw strings and calls formatting variance disagreement
- Rounds aggressively without checking what it merges
- Treats a parse failure as just a lost sample worth ignoring
- Thinks normalization is only needed for numbers