How would you test whether a model's stated reasoning actually drove its answer?
answer
- intervene, do not just read
- truncate and force an answer early
- corrupt a step, watch it propagate
- swap the chain for filler tokens
- did the trace mention the cue
basics
~20 sFaithfulness is measured by intervening on the trace and watching the answer. Truncate it early, corrupt a step, paraphrase it, or replace it with filler, then check whether the final answer moves. A trace you can break without changing the answer was not doing the work.
solid answer
~50 sYou cannot read faithfulness off a trace; you have to intervene on it and measure the answer's dependence. Four protocols cover most of it. **Early answering**: truncate the chain after each step, force an answer, and plot accuracy against truncation point — if the answer is already fixed after step one, the remaining steps are decoration. **Corrupting a step**: introduce an error mid-chain and continue; a faithful trace propagates the error into a different final answer. **Paraphrasing**: restate the chain in different words but the same content; the answer should be stable, since content, not wording, should be carrying it. **Filler substitution**: replace the reasoning with contentless tokens; if accuracy holds, the tokens bought compute, not reasoning. A fifth, sharper test inserts a cue that biases the answer and asks how often the flipped answer's trace ever mentions the cue. Report these as answer-change rates on a fixed item set, not as a single faithfulness score.
go deeper
Know that faithfulness asks whether the written reasoning is the real reason for the answer, and that you test it by changing the trace and seeing whether the answer changes, not by reading it.
Describe at least two protocols concretely — early answering and corrupting a step — and state what result each produces. Be clear that the measurement is an intervention with a paired comparison, not an inspection.
Design the study: fixed item set, several samples per item, multiple protocols, answer-change rate as the reported metric, and the caveat that interventions can push the model off distribution. Say what decision the result gates, such as whether a trace may be shown as an explanation.
Own the policy question. Decide what faithfulness evidence is required before a reasoning trace is used as an audit record or surfaced to users or regulators, what re-measurement cadence applies when the model changes, and how you would communicate that a plausible-looking trace is not an explanation.
## Why you need an intervention, not an inspection A reasoning trace is text the model produced; it is not an execution log. Reading it tells you what the model wrote, not what determined the output. Faithfulness measurement therefore works like an ablation study: change something about the trace, hold everything else fixed, and see whether the answer responds. If it does not respond, the trace was not on the causal path. Everything below is a measurement protocol. Run each on a fixed item set with several samples per item, and report distributions of answer-change rate rather than a single number, because faithfulness varies sharply by task difficulty and by model. ## Early answering Take a completed chain of N steps. For k from 0 to N, feed back only the first k steps and force the model to produce a final answer immediately. Plot the accuracy, or the agreement with the full-chain answer, against k. The shape of that curve is the result. A curve that rises gradually and only reaches the full-chain answer near the end says later steps carried information. A curve that is flat from k equals one says the answer was already determined before the reasoning happened, and the visible steps are post-hoc narration. Summarise the curve as area under it, or as the smallest k at which the full-chain answer is already produced most of the time. ## Corrupting a step Inject a specific error into step j — change a number, invert a comparison, substitute a wrong intermediate result — then let the model continue from the corrupted prefix. A faithful chain carries the corruption forward and yields a different, usually wrong, final answer. A chain that quietly returns to the original answer despite a broken premise was not being used. This one is the most diagnostic and the most work: the corruption has to be a genuine change to the content of that step, not a typo the model can normalise away, and someone has to author or validate the corruptions. ## Paraphrasing Rewrite the chain preserving its meaning and changing its surface form, then continue to an answer. Here you are testing the opposite direction: a faithful chain should be robust, because the content is what matters. Large answer swings under a meaning-preserving rewrite tell you the model is keying on surface features of the trace rather than its content, which undermines any interpretation of the trace as reasoning. ## Filler-token substitution Replace the reasoning with an equal number of contentless tokens — repeated punctuation, a neutral filler phrase — and measure accuracy. If accuracy is preserved, the extra tokens were providing serial computation rather than the specific reasoning they appeared to state. This is the cleanest way to separate "the chain helped" from "the chain's content helped", and it is worth running before you invest in step-level grading of a trace that may be doing nothing semantically. ## The cue test The sharpest protocol constructs pairs of prompts that differ only by an added cue that biases the answer — reordering multiple-choice options so the correct one is always in the same slot, appending a stated preference, or prefixing an authority's opinion. Measure two things: how often the final answer flips toward the cue, and, among flipped cases, how often the trace mentions the cue as a reason. A high flip rate with near-zero mention rate is direct evidence that the stated reason is not the operative one. ## Reading the results honestly Several caveats belong in any report. Faithfulness is not binary and not a single number — a trace can be causally load-bearing and still be an incomplete account. Results vary with task difficulty; easy items often show low faithfulness simply because the answer is retrievable without reasoning. Interventions can be out of distribution, so a model may behave oddly on a corrupted prefix for reasons unrelated to faithfulness, which is why paired designs and multiple protocols matter more than any single test. And some providers restrict or summarise the internal reasoning that is returned, so what you can intervene on may not be the full trace the model produced — say so when it applies. ## What to do with the number Use it as a gate on interpretation, not as a model score. If a trace turns out to be low-faithfulness on your task, you cannot use it as an audit artefact, you should not show it to users as an explanation, and step-level grading of it measures the narration rather than the computation. Those are the decisions this measurement exists to inform.
- What does a flat early-answering curve tell you, and what would you do about it?A flat curve means the full-chain answer is already produced after one or two steps, so the visible reasoning is narration rather than computation. Practically it disqualifies that trace as an audit artefact or a user-facing explanation on that task. Before drawing conclusions, check that the task is not simply easy enough to answer without reasoning, by comparing against harder items.
- Why run the filler-token test before investing in step-level grading of a trace?Because if accuracy survives replacing the reasoning with contentless tokens, the trace's content was not doing the work, and grading its steps measures the quality of narration. Filler substitution is cheap and one run answers a prior question: does the semantic content of this chain matter at all on this task?
- How do you keep a corruption test from measuring out-of-distribution weirdness instead of faithfulness?Make the corruption a plausible, well-formed error a person could make rather than a malformed string, validate that the corrupted prefix still reads as ordinary text, and use a paired design where the same items are run corrupted and clean. Also run more than one protocol; agreement across early answering, corruption and filler substitution is far more convincing than any single test.
saying these in an interview costs you the question
- Says you can judge faithfulness by reading the trace
- Treats faithfulness as one binary score per model
- Confuses an incorrect step with an unfaithful trace
- Runs one protocol and calls the result conclusive
- Ignores that easy items look unfaithful because reasoning is unnecessary