skip to content

Why can a chain-of-thought trace be unfaithful to how the model actually reached its answer?

level: middleimportance: must knowfreq 70%

answer

  1. explanation is not the same as cause
  2. the text is output, not instrumentation
  3. the real driver goes unstated
  4. post-hoc rationalisation of a prompt cue

basics

~10 s

A reasoning trace is generated text, not a log of the computation behind the answer. The model can write a plausible justification while the cue that actually drove its output goes unmentioned.

solid answer

~50 s

Chain-of-thought text is sampled the same way as any other output — nothing binds it to whatever internally produced the answer, and there is no mechanism that forces a model to report its own causes. So a trace can be **post-hoc rationalisation**: a fluent, well-structured argument for a conclusion that was really driven by something the trace never names. The standard demonstration uses a biased few-shot prompt — for example a resume-screening task where the answer in every exemplar happened to be option (A). On a fresh candidate the model answers (A) and its trace argues confidently from years of experience and keyword matches, never mentioning the pattern that actually moved it. The practical rule: treat a trace as a hypothesis about the answer, not as evidence for it. Verify the checkable claims independently, and never hand a trace to a user or an auditor as "the reason" for a decision.

go deeper

for a junior

Be ready to say plainly that the model writes the reasoning as text, so a neat step-by-step explanation is not proof the answer is right or that those steps are what produced it.

for a middle

Explain the mechanism: the trace is sampled output with nothing tying it to the internal computation, so an unmentioned prompt cue can drive the answer while the trace rationalises it. Be able to describe the biased-exemplar demonstration.

for a senior

Show the production consequence. Say how you would detect a cue-driven answer, why debugging from the trace can send a team after the wrong fix, and how you keep verifiable evidence rather than narration as the thing the system stands on.

for a principal

Own the governance angle: decide, for your organisation, what may be shown as an explanation of an automated decision. Traces are not audit-grade evidence, and committing to them as such creates a liability you cannot later substantiate.

## What "faithful" means here A chain-of-thought (CoT) trace is the intermediate reasoning text a model emits before its final answer. A trace is **faithful** if the steps it states are the actual reasons for the answer — if the reasoning had gone differently, the answer would have gone differently. A trace is **unfaithful** when the answer is driven by something the trace omits, misstates, or contradicts. Note that faithfulness is not the same as correctness: a trace can be entirely accurate as prose and still be an unfaithful account of what produced the output, and a wrong trace can sit above a right answer. ## Why the trace can come apart from the computation The trace is output, not instrumentation. It is a sequence of tokens sampled from the same distribution as the answer, conditioned on the prompt. There is no channel that reads out the model's internal computation and renders it as English, and no part of training guarantees that a model's self-report matches its own causes. Training pressure runs the other way: text that *looks* like good reasoning is rewarded, and reasoning-shaped prose is cheap to produce whether or not it describes anything real. Humans do a version of this. Ask someone why they picked one of four identical items on a shelf and they will give you a confident reason — quality, packaging, familiarity — even when the real driver was position on the shelf. The reason is generated after the fact, by a process that has no access to the one that made the choice. ## The canonical demonstration Build a few-shot prompt for a screening task — say, ranking candidates from a resume — where by construction the correct answer in every worked example is option (A). Present a new candidate. Models tend to answer (A) at a rate far above chance, and their traces argue from the candidate's experience and skills. The traces essentially never say "every example I was shown was (A)". Strip the bias out of the exemplars and the answer changes, which shows the cue was doing the work. The trace narrated a different story the whole time. ## The recurring shapes - **Post-hoc rationalisation.** An unmentioned cue — option position, exemplar bias, the order of choices, a preference the user expressed — decides the answer, and the trace supplies respectable-sounding support for it. - **Decorative reasoning.** The model has effectively settled the answer immediately; the steps are filler that neither constrains nor changes the conclusion. - **Non-entailing traces.** The stated steps do not actually imply the final answer, yet the answer is stated as though they did. This one is detectable by careful reading, which makes it the least dangerous kind. - **Unacknowledged influence.** A retrieved snippet, a system instruction, or a formatting hint moves the answer, and the trace presents the conclusion as derived from first principles. ## Why it matters in a real system Well-structured prose buys trust it has not earned. Human reviewers approve outputs faster when a tidy trace is attached, which is exactly backwards if the trace is decoration. Debugging suffers too: if you read a trace to find out why a classifier is wrong, you may spend a sprint fixing the reason the model wrote down instead of the cue that actually drove it — a prompt artefact, a template quirk, an ordering effect. And when a trace is shown to an end user, a regulator, or an internal risk committee as the explanation for an automated decision, unfaithfulness turns from a technical nuance into a governance problem: you are attesting to a causal story you cannot support. ## What to do instead Design so that trust rests on things you can check, not on narration: - **Verify the checkable steps.** Where the trace does arithmetic, run the arithmetic. Where it cites a document, match the citation against the source. Where it claims a lookup, do the lookup. - **Push decision-relevant facts into structured, verifiable form.** Quoted spans with offsets, retrieved evidence with identifiers, tool results — all of these can be validated; a paragraph of reasoning cannot. - **Test at the outcome level, including on adversarial variants.** Construct inputs where a cue is present and inputs where it is not, and watch whether the answer moves. That the trace stayed the same while the answer flipped is the tell. - **Label traces honestly in the product.** "Model-generated reasoning, not a record of how the decision was made" is an accurate caption; "why we decided this" is not. None of this means CoT is a bad technique. It means the trace is a useful artefact for reading, sampling and debugging, and a poor artefact for proving anything.

  • How would you notice unfaithfulness in a system already running in production?
    Work at the outcome level. Build paired inputs that differ only by a suspected cue — option order, an exemplar pattern, a formatting quirk — and check whether the answer moves while the trace stays essentially the same. Independently verify the trace's checkable claims (arithmetic, citations, lookups) and track how often they fail. Correlations between answers and prompt artefacts that no trace ever mentions are the signal.
  • Does a longer, more detailed trace tend to be a more faithful one?
    No. Verbosity is not fidelity. A longer chain gives more surface for confident-sounding justification, and in practice it makes an unfaithful answer more persuasive to a human reviewer rather than less. Length correlates with cost and latency far more reliably than it correlates with whether the stated steps caused the answer.
  • If traces can be unfaithful, is it still reasonable to show them to users?
    Showing them is fine; framing them as the decision's justification is not. Label the trace as model-generated reasoning, keep the verifiable evidence — quoted sources, tool outputs, computed values — as the thing the product actually stands behind, and never let a trace be the only artefact backing a consequential decision.

Ask someone why they chose one of four identical boxes on a shelf and they will confidently cite quality or packaging, when the real driver was shelf position. The explanation is produced after the choice, by a process with no access to the one that made it.

saying these in an interview costs you the question

  • Treats the trace as a log of what the model internally computed
  • Assumes a confident, well-structured trace proves the answer is correct
  • Believes asking the model to explain itself reveals the true cause
  • Thinks unfaithfulness only affects small or weak models
  • Concludes chain-of-thought is worthless because traces can be unfaithful

context