Why does emitting intermediate steps improve an LLM's multi-step accuracy?
answer
- answers cost computation, not just knowledge
- one token, one forward pass
- fixed depth per token, unbounded token count
- earlier results become readable context
- hard prediction split into easy ones
basics
~20 sEach generated token costs one pass through a fixed-depth network, so a direct answer must fit the whole computation into a single pass. Writing steps spreads the work across many passes and stores partial results in the context for later tokens to read.
solid answer
~50 sTwo mechanisms, and interviewers want both. First, compute: a transformer applies a fixed number of layers per token, so the amount of serial work available before committing to an answer token is bounded. Emitting intermediate tokens unrolls the computation over many forward passes, giving the model more total steps of processing for the same problem. Second, conditioning: generated tokens are appended to the context, so a partial result written down is literally re-read by every later token instead of having to survive inside one pass of activations. Together they also decompose one hard next-token prediction into several easy ones - predicting 'VAT on 372.74 is 40.06' is far more tractable than predicting the final total straight from a shipment description. Theory backs this up: fixed-depth transformers are limited to a shallow parallel circuit class in one pass, and intermediate tokens let them simulate substantially longer serial computations.
go deeper
Know that a model gets one pass through the network per token, so writing more tokens gives it more chances to work, and that written steps stay visible for the rest of the answer.
Explain both mechanisms cleanly: extra forward passes for serial computation, and intermediate results stored in the context that later tokens condition on. Add that each local prediction becomes easier.
Draw the operational consequence: the technique raises a ceiling but does not guarantee correctness, so numeric or safety-critical steps still need a verifier or a tool, and step order must put constraints first.
Own the framing that reasoning length is a compute budget you are spending per request. Be ready to argue where that budget belongs - in the model's tokens, in orchestration, or in deterministic code - across a product's cost and reliability targets.
## The question behind the question Why should text the model writes for itself make it better at arithmetic or multi-hop lookup? Nothing new enters the system - no data, no tool, no extra parameters. The answer is that something else does change: how much computation the model gets to spend, and where its partial results live. ## Mechanism one - tokens are units of computation A transformer decoder has a fixed number of layers. Producing one token means running the input through those layers once. That per-token budget is constant: it does not grow because the question is hard. So if the model answers a freight quote directly, the entire chain - identify the weight tier, look up the rate, multiply, take a percentage, add, take another percentage, add again - has to be accomplished inside one fixed-depth pass, in parallel across positions, and then squeezed into a single output token. That is a strong constraint. Generating twelve tokens of working, by contrast, means twelve passes through the same network, each conditioned on more information than the last. The serial depth of the computation is no longer capped by the model's layer count; it is capped by how many tokens you let it write. This is not just intuition. Complexity-theoretic analyses place a single forward pass of a fixed-depth, limited-precision transformer in a shallow parallel circuit class (TC0 in the usual formalisation), which provably cannot express certain inherently sequential computations. Allowing the model to emit intermediate tokens and read them back lifts that ceiling: with enough intermediate steps the same architecture can simulate much longer serial computations. The practical translation is blunt - some problems are not hard for the model because it lacks knowledge, they are hard because it was not given enough steps to think in. ## Mechanism two - the context is external memory The second effect is about state. In a direct answer, every intermediate quantity (the base charge, the subtotal) exists only as activations inside one pass, competing for representational room with everything else the model is doing. Nothing forces those quantities to be preserved cleanly. Write them out and they become tokens. Tokens are cached and attended to. "Subtotal: EUR 372.74" is now a discrete, unambiguous fact the model can retrieve exactly when it computes VAT. The chain functions as a scratchpad in the plain sense: the model offloads working state into the output stream rather than holding it implicitly. This is also why order matters. A chain that states the max-dose ceiling before computing the dose lets the later step condition on the constraint; a chain that mentions it only at the end cannot influence what was already generated. Autoregressive generation is one-directional: a token can only be influenced by what precedes it. ## Mechanism three - easier local predictions There is a third, more prosaic reason, and it is worth saying out loud. Chain-of-thought turns one very hard conditional prediction into a sequence of easy ones. P(final total | raw shipment description) is a distribution the model has weak evidence about. P(next line | tier stated, rate stated, base charge stated) is close to something it has seen thousands of times in training. Each step is a short hop within the distribution the model models well, and the chain composes those hops. A related benefit: decomposition surfaces the sub-results, so an error is localised rather than diffuse. If the fuel surcharge line is wrong, that is visible and checkable - which is what makes verifier-based approaches possible at all. ## What this predicts, and what it does not The compute argument predicts the observed shape of CoT's benefit: large gains on tasks with genuine serial dependency, small or negative gains on tasks that were already a single easy prediction. It also predicts that padding with meaningless tokens should not help nearly as much as meaningful ones - and empirically, filler-token results are mixed and much weaker than real reasoning traces, because the second mechanism (useful state written down) is doing real work alongside the first. What the argument does not give you is any guarantee of correctness. More computation and better working memory raise the ceiling; they do not make the model's arithmetic exact or its steps honest. A model can execute a well-formed chain and still multiply wrongly, and it can produce a chain that is not the actual cause of its answer. ## Answering it in an interview A strong answer names both mechanisms in one breath: more forward passes before committing to an answer (compute), and partial results written into the context where later tokens can read them (state). Mention that the decomposition also makes each individual prediction easier. Mention the ceiling: the benefit is bounded by the model's ability to produce valid steps, not by the technique.
- If the gain came purely from extra forward passes, would nonsense filler tokens work just as well?They should help somewhat if compute were the only mechanism, and there is research showing filler or pause tokens can give measurable gains on some tasks. But the effect is far weaker than real reasoning traces, because meaningful steps also externalise partial results and make each subsequent prediction easier. Both mechanisms are doing work, so nonsense captures only one of them.
- Does chain-of-thought remove the need for a calculator tool on numeric work?No. The steps give the model more computation and a place to keep intermediates, but each arithmetic operation is still a token prediction and can be wrong. For anything where an exact number matters - billing, dosing, financial totals - the right pattern is to use the chain for decomposition and hand the actual arithmetic to a tool, then reason over the tool's result.
- Why does the order of steps in the chain matter so much?Generation is autoregressive: a token can only be conditioned on tokens before it. A constraint or a retrieved fact stated after the computation cannot influence that computation - it can only comment on it. So constraints, units and retrieved facts belong early in the chain, before the steps that must respect them.
Mental arithmetic versus arithmetic on paper. The brain is the same either way, but paper lets you take more sequential steps and lets you look back at line two while writing line three.
saying these in an interview costs you the question
- Saying the model 'thinks harder' without naming a mechanism
- Claiming extra tokens increase the model's depth or parameters
- Believing intermediate steps make arithmetic exact
- Assuming any extra tokens help as much as meaningful ones
- Missing that written steps are re-read as context by later tokens