skip to content

Does a confident chain-of-thought trace mean its intermediate steps are true?

level: juniorimportance: must knowfreq 60%

answer

  1. every step is predicted, not verified
  2. fluency is not correctness
  3. structure buys undeserved trust
  4. a bad step propagates to the answer
  5. check the arithmetic and the citation

basics

~20 s

No. Every step is predicted text, so a chain can contain invented facts — a factor pair that does not multiply out, a contract subsection that does not exist — while reading as careful and rigorous.

solid answer

~50 s

Each line of a reasoning chain is generated the same way the final answer is, so a step can be fabricated just as easily as an answer can. Two familiar shapes: ask whether 17077 is prime and a model may assert "17077 = 113 x 151, so it is composite" — those factors multiply to 17063, and 17077 is in fact prime; or ask a yes/no question about a contract and the trace cites a subsection that is not in the document at all. What makes this dangerous is presentation. The fabrication arrives inside a numbered, methodical-looking chain, so a reader who would have challenged a bare wrong answer accepts it. The mitigation is mechanical, not rhetorical: identify which steps are checkable — arithmetic, citations, lookups, unit conversions — and check them outside the model rather than asking the model whether it is sure.

go deeper

for a junior

Know that each reasoning step is generated text, so it can be invented. Say concretely what you would check: recompute the arithmetic, and confirm that any cited section actually exists in the source.

for a middle

Explain propagation — once a fabricated value or clause is written, the rest of the chain conditions on it — and why self-reported confidence is not a check. Name the deterministic verifications you would automate.

for a senior

Show how you would design the fabrication out: delegate computation to code, require verbatim quoted spans with a source match, and reject outputs whose citations fail validation before a human ever sees them.

for a principal

Frame the exposure. Decide which answer classes may never rest on generated intermediate steps, and require a verifiable evidence path for those, accepting the extra cost as the price of a defensible decision.

## The claim being tested Step-by-step output feels like a proof. It is not. A chain of thought is a sequence of generated sentences, produced by the same next-token process that produces the final answer, under the same pressure to be fluent and plausible. Nothing in that process checks a step against the world before writing it down. So an invented step is not an exotic failure — it is the ordinary failure mode of generation, wearing the costume of a derivation. ## Two concrete shapes **Invented computation.** Ask a model whether 17077 is prime. A plausible-looking answer is "17077 = 113 x 151, so it is composite." Multiply it out: 113 x 151 = 17063, not 17077. And 17077 has no divisor up to its square root, so it is prime. The chain is confident, it is formatted as arithmetic, and it is simply wrong at the step that carries all the weight. **Invented sources.** Ask a yes/no question about a supplier contract — may the buyer terminate for convenience? — and a trace may reason "under section 4.3(b), the buyer may terminate on 30 days' notice, therefore yes," where the document contains no 4.3(b) at all. The citation looks like grounding. It is decoration. Other members of the family: a lookup that returns a fact the source never stated, a unit conversion done with a made-up factor, a definition attributed to a standard that defines nothing of the sort, an intermediate quantity carried forward that was never derived. ## Why fabricated steps are more dangerous than bare wrong answers Three reasons compound. First, **structure buys trust**. A numbered chain reads as work shown. Reviewers who would push back on a naked assertion nod along to a derivation, because the derivation implies a process was followed. Second, **errors propagate**. Once a bad intermediate value or a bogus clause is written down, the rest of the chain conditions on it. The model does not go back; it builds. So a single fabricated step usually determines the final answer rather than getting averaged away. Third, **the answer can be right anyway**, which is worse for your evaluation loop than being wrong. If a model claims a nonexistent clause and still concludes "yes, termination is allowed" because some other clause happens to permit it, outcome-level testing scores a pass and the defect stays invisible until a case where the same habit lands on the wrong side. ## How to catch them The unhelpful move is to ask the model to double-check itself, or to ask for confidence per step. Self-reported confidence is generated text too, and it correlates with fluency more than with truth. The useful moves are external and mechanical: - **Recompute anything arithmetic.** If a step multiplies, divides, sums or converts, run it. Delegating the computation to a calculator or code path removes the failure entirely rather than detecting it. - **String-match every citation against the source.** A trace that cites section 4.3(b) should be rejected automatically when "4.3(b)" does not appear in the document. This is a cheap, deterministic check that catches an entire class of fabrication. - **Require quoted spans, not paraphrase.** Ask for the verbatim text the conclusion rests on, then verify the quote appears in the source. A model that must quote has far less room to invent. - **Verify the direction, not just the value.** If the chain says "therefore composite", the check is not only "are these factors real" but "does their product equal the input". ## The right mental model Read a chain of thought as a draft argument written by someone knowledgeable, fast, and occasionally willing to make something up rather than say "I don't know". You would not sign off on that person's memo without spot-checking the numbers and pulling up the cited section. Apply exactly that habit, and prefer to design the system so the checkable parts never depend on the model's word in the first place.

  • Why does an invented step often persuade a reviewer more than a bare wrong answer would?
    Because a numbered chain signals that work was done. Reviewers treat structure as evidence of process and lower their scrutiny, so a fabricated clause or factor slips through where a naked assertion would have been challenged. The presentation, not the content, is doing the persuading.
  • Which steps are worth verifying automatically, and how?
    Anything mechanically checkable: arithmetic and unit conversions (recompute, or delegate to a calculator or code path so it is never generated), citations and section references (string-match against the source document), and named lookups (re-run against the system of record). These checks are deterministic, cheap, and catch whole classes of fabrication.
  • What if the final answer is correct but one intermediate step was invented?
    Treat it as a defect, not a pass. The system got the right answer for a reason it did not actually have, so nothing about that success predicts the next case. Outcome-only scoring hides this, which is why spot-checking steps on a sample of traces matters even when accuracy looks fine.

saying these in an interview costs you the question

  • Assumes showing work means the work was verified
  • Trusts a cited section number without checking the document
  • Asks the model to rate its own confidence as the check
  • Thinks a correct final answer proves the steps were sound
  • Believes fabricated steps only appear in obscure or hard questions

context