skip to content

CoT Evaluation Methods

How reasoning quality is actually measured: step-level correctness rather than only the final answer, GSM8K/MATH/BBH-style benchmarks, faithfulness checks, and the process-versus-outcome reward distinction from RLHF. Interviewers ask because a right answer reached through a wrong chain is still a bug.

on this pageshow

questions

5

Why isn't final-answer accuracy enough to evaluate chain-of-thought reasoning?

level: middleimportance: must knowfreq 62%

answer

  1. scores the last line only
  2. right for the wrong reason
  3. cancels out, still counts correct
  4. first invalid step index
  5. gap between correct and fully valid

basics

~20 s

Final-answer accuracy scores only the last line, so a model that reaches the right number through a broken chain still counts as correct. Step-level scoring catches these right-for-the-wrong-reason solutions, which generalize worst to new problems.

solid answer

~50 s

Final-answer accuracy (exact match on the extracted answer) is one bit per problem, and it is blind to how that bit was produced. Grading an actuarial exam free-response makes the gap obvious: if step 3 uses the wrong denominator but a later factor cancels it out, the final number is right and exact match awards full credit — a human grader would not. Step-level scoring instead labels each intermediate step and reports things like the fraction of valid steps, the position of the first invalid step, and the share of correct-answer solutions that contain at least one invalid step. That last number is the interesting one: it estimates how much of your headline accuracy is luck. The error runs both ways too — a sound derivation that trips on the final arithmetic or fails answer extraction is scored as a total failure. Report both, because they answer different questions.

code

json · 11 lines
json
{
  "problem_id": "act-2019-q4",
  "final_answer_correct": true,
  "steps": [
    { "index": 0, "text": "Total exposure = 4,200 policy-years", "label": "correct" },
    { "index": 1, "text": "Deaths observed = 63", "label": "correct" },
    { "index": 2, "text": "Rate = 63 / 4,200 * 1.0 using in-force count", "label": "incorrect" }
  ],
  "first_error_index": 2,
  "fully_valid": false
}

go deeper

for a junior

Be able to say plainly that exact match compares only the final answer, so a model can be right for the wrong reason. Naming that failure and suggesting you also check the steps is enough at this level.

for a middle

Explain the two-sided error concretely: an invalid step that cancels out scores correct, and a sound derivation with a slip at the end scores zero. Name at least two step-level metrics, such as first-error position and fully-valid solution rate.

for a senior

Show you would quantify the gap — among correct-answer solutions, what share are fully valid — and use it to decide whether a benchmark gain is real. Talk about where the labels come from and how you would validate an automated labeller.

for a principal

Own the reporting policy: which metric goes on the dashboard, what you accept as evidence for a model or prompt change, and what annotation budget you are willing to spend for that evidence. Be clear that step-level scoring buys diagnosis, not comparability, and say when cheap exact match is the right call.

## The metric being critiqued The standard way to score a reasoning benchmark is exact match on the final answer: run the model, extract the answer (usually the text after a marker such as "the answer is", or a boxed expression), normalise it, and compare it to the gold value. This gives one binary outcome per problem, and accuracy is the mean over the set. It is cheap, objective, and reproducible, which is why nearly every public reasoning suite is reported this way. Its weakness is that it compresses an entire multi-step derivation into a single comparison at the end. Two solutions that reach the same number are indistinguishable to the grader, even if one is a clean derivation and the other is a mess that accidentally cancels out. ## Failure mode one: right answer, wrong chain This is the false positive. On an actuarial exam free-response, a candidate computes a mortality rate but divides by the wrong exposure base in step 3; two steps later a factor drops out and the final figure matches the answer key. Exact match scores it correct. So does a majority vote over several samples, if the samples share the flaw. A human grader deducts most of the marks, because the derivation does not generalise: change the numbers and the same chain produces a wrong answer. The practical consequence is that headline accuracy overstates capability by however often this happens, and the overstatement is not uniform — it is largest exactly where the problem space is small enough for a wrong method to land on a right answer (multiple-choice, small integer answers, problems with symmetric numbers). ## Failure mode two: sound chain, wrong final answer The false negative. The derivation is entirely valid and the model slips on the last multiplication, or writes the answer in a format the extraction regex misses (a fraction where the key has a decimal, units attached, an answer given as a sentence). Exact match scores zero, identical to a model that had no idea how to start. For diagnosis these are completely different situations: one calls for a better answer-extraction step or a calculator tool, the other calls for a different approach to the task. ## What step-level scoring actually reports Segment the trace into steps (usually lines or sentences) and attach a label to each — typically correct, neutral, or incorrect. From those labels you derive: - **First-error position**: the index of the earliest step marked invalid. Useful because everything after the first error is conditioned on a broken state and is not independently meaningful. - **Valid-step fraction**: how much of the derivation survives grading, giving partial credit like a human marker. - **Fully-valid solution rate**: the share of solutions with no invalid step. Compare this to final-answer accuracy; the gap is your luck estimate. - **Conditional soundness**: among solutions with the correct final answer, what fraction are fully valid. This is the number that tells you whether a benchmark gain is real reasoning improvement or better guessing. ## Where the labels come from Humans, a model acting as a step grader, or an automated procedure that samples continuations from each prefix and labels a step by how often it still leads to the correct answer. Each is a different cost and noise profile, and any automated labeller should be validated against a human-labelled sample before you trust its numbers. ## When it is worth the cost Step-level scoring costs far more than exact match — you are grading every line, not one token. It earns its keep when the reasoning is the product (tutoring, audit trails, regulated derivations where a reviewer must follow the work), when you are choosing between prompts or models that look tied on final accuracy, and when you are training or selecting on the reasoning signal itself. For a quick regression check on a stable pipeline, final-answer accuracy on a fixed set is fine. ## What to report Both, side by side, on the same items. Final-answer accuracy is the comparable, cheap number that other people's results can be read against. Step-level metrics are what you use internally to decide whether the model is reasoning or the benchmark is being gamed. Reporting only the first is how a team ships a prompt that improves the score and degrades the work.

  • How would you score a solution whose steps are all valid but whose final answer is wrong?
    Report it separately from a solution that goes wrong early — it is a near miss, not a failure to reason. Check first whether the answer is actually wrong or merely unextractable, since format mismatches masquerade as reasoning failures. If the derivation is sound and only the last arithmetic slips, the fix is a calculator tool or a verification pass, not a different prompting strategy.
  • Does step-level scoring require a human to label every step?
    No. You can have a grader model label steps, or use an automated procedure that samples continuations from each prefix and scores a step by how often it still reaches the correct answer. Both are noisy, so validate them against a human-labelled sample and report the agreement rate; otherwise you are measuring the labeller, not the model.
  • Why is it standard to stop labelling a solution at its first invalid step?
    Everything after the first error is generated conditioned on a broken state, so labelling those steps mixes two things: whether the model reasons well, and whether it reasons well given a wrong premise. Stopping at the first error keeps the label set interpretable and cuts annotation cost, at the price of losing information about recovery behaviour.

saying these in an interview costs you the question

  • Assumes a correct final answer implies a correct chain
  • Says exact-match accuracy is the only metric worth reporting
  • Treats a formatting or extraction failure as a reasoning failure
  • Claims longer chains are automatically better reasoning
  • Thinks step-level scoring just means checking the trace is verbose

context

open as a page

What do process reward models reward that outcome reward models don't?

level: middleimportance: should knowfreq 45%

basics

~20 s

An outcome reward model scores only the final answer, so every chain that reaches it gets reinforced, including lucky ones. A process reward model scores each reasoning step, penalising an invalid derivation even when the final answer is right.

open as a page

How would you test whether a model's stated reasoning actually drove its answer?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Faithfulness is measured by intervening on the trace and watching the answer. Truncate it early, corrupt a step, paraphrase it, or replace it with filler, then check whether the final answer moves. A trace you can break without changing the answer was not doing the work.

open as a page

How do you obtain step-level labels for reasoning traces without hand-annotating every step?

level: seniorimportance: should knowfreq 28%

basics

~20 s

The cheap route samples several continuations from each prefix and labels a step by how often it still reaches the correct final answer, needing no annotator. Human labelling stays for a stratified audit sample used to validate the automated labels.

open as a page

A new CoT prompt lifts GSM8K by 4 points — what would you check before believing it?

level: principalimportance: should knowfreq 33%

basics

~20 s

Check four things before accepting the gain: run-to-run variance with confidence intervals on paired items, harness differences such as shot count and answer extraction, contamination and saturation on a decade-old public set, and whether the gain transfers to the actual product task.

open as a page