skip to content

Chain of Thought

Making the model show its work: explicit intermediate steps, the zero-shot and few-shot ways to elicit them, and sampling several chains to vote on an answer. Interviewers ask about it constantly, including the part people skip — when a reasoning trace is unfaithful or actively hurts.

on this pageshow

explore

questions

page 1 of 2

In LLM prompting, what is chain-of-thought and how does it change the output?

level: juniorimportance: must knowfreq 82%

answer

  1. the model writes before it answers
  2. steps live in the output text
  3. worked solution, then the result
  4. generated tokens, no calculator involved
  5. helps when step three needs step two

basics

~20 s

Chain-of-thought means having the model write out intermediate reasoning steps before its final answer instead of answering immediately. The output becomes a short worked solution followed by the result, which raises accuracy on problems needing several dependent steps.

solid answer

~50 s

Chain-of-thought prompting asks the model to produce the intermediate steps of a solution in its output, before stating the final answer. Instead of jumping straight to a freight quote, the model writes that the shipment falls in the 500-1000 kg weight tier, that the tier's base rate is X, that the fuel surcharge is a percentage of that base, that VAT applies to the sum, and only then gives the total. The steps are ordinary generated tokens - nothing symbolic is executed - but because each step is written into the context, the tokens that follow can read and build on earlier results. It pays off on problems whose answer depends on several dependent sub-results; on a single-fact lookup it mostly adds tokens. Most current frontier models are post-trained to reason this way by default, so it is now as much a model behaviour as a prompt technique.

code

markdown · 12 lines
markdown
Direct answer:
Q: A 640 kg shipment. Quote it.
A: EUR 412.80

Chain-of-thought answer:
Q: A 640 kg shipment. Quote it.
A: 640 kg falls in the 500-1000 kg tier, rate EUR 0.52/kg.
   Base charge: 640 x 0.52 = EUR 332.80
   Fuel surcharge 12% of base: EUR 39.94
   Subtotal: EUR 372.74
   VAT 10.75%: EUR 40.06
   Total: EUR 412.80

go deeper

for a junior

Be able to say plainly that chain-of-thought means the model writes intermediate steps in its output before the final answer, and that this helps on problems with several dependent steps. Give a concrete before-and-after example.

for a middle

Explain that the steps are generated tokens that later tokens condition on, so the model both decomposes the problem and keeps partial results in context. Be clear that no external computation happens.

for a senior

Show judgement about when the extra tokens are worth their cost, what you do with the trace once you have it (log, review, discard), and why a readable chain is not evidence of a faithful one.

for a principal

Frame it as a design choice about where reasoning lives in your system: in the model's output, in orchestration code, or in a tool. Own the cost, latency, exposure and auditing consequences of that choice across a product.

## The idea in one line Chain-of-thought (CoT) is the practice of getting a language model to emit the intermediate steps of a solution as part of its output, ahead of the final answer, rather than producing the answer directly. A language model generates text one token at a time, each token conditioned on everything already in its context - the prompt plus whatever it has written so far. "Direct answering" means the tokens right after the question are the answer itself. Chain-of-thought means the tokens right after the question are a worked solution, and the answer comes at the end of it. ## What it looks like Take a freight quote. A direct answer is a single number: "EUR 412.80." A chain-of-thought answer is a short sequence of stated sub-results: - The shipment weighs 640 kg, so it falls in the 500-1000 kg tier. - That tier's base rate is EUR 0.52 per kg, so the base charge is 640 x 0.52 = EUR 332.80. - The fuel surcharge is 12% of the base charge: EUR 39.94. - Subtotal is EUR 372.74. - VAT at 10.75% on the subtotal is EUR 40.06. - Total: EUR 412.80. Every one of those lines is generated text. No calculator ran. The model is still predicting the next token; it is simply predicting tokens that happen to spell out a decomposition of the problem. ## Why writing the steps matters Two things change once the steps are on the page. First, the problem is decomposed. Predicting "the base charge is 332.80" given the tier and the rate is a much easier next-token prediction than predicting the final total given only the raw shipment description. Each written step turns one hard prediction into several easy ones. Second, the steps become context. Generated tokens are appended to the sequence, so every later token can attend back to them. The intermediate results are stored in the output itself rather than having to be held implicitly inside one pass of the network. If the model needs the subtotal to compute VAT, it can literally read the subtotal it just wrote. ## Where it helps and where it does not CoT is a tool for problems with dependent sub-results: arithmetic with several operations, multi-hop questions (find the coach, find the coach's team, find that team's stadium), constraint checks, unit conversions, date and timezone arithmetic. The common shape is that you cannot compute step three without the output of step two. It does far less for problems that are a single retrieval or a single classification - naming a capital city, extracting a phone number from a paragraph, choosing between two synonyms. There is no chain to unroll, so the extra tokens buy little except cost and latency. ## Elicited versus native The technique first became well known as something you asked for in the prompt: show the model worked examples, or instruct it to reason before answering. Since roughly 2024-2025, models have increasingly been post-trained (via supervised traces and reinforcement learning on verifiable tasks) to produce intermediate reasoning on their own for hard inputs. In practice today you often are not switching CoT on so much as shaping it: deciding how much reasoning a task warrants, what structure the steps should have, and whether the steps belong in the user-visible answer at all. The underlying mechanism is unchanged - tokens generated before the answer, conditioned on by the answer. ## What CoT is not It is not a proof, and it is not a guaranteed window into how the model actually arrived at the answer; a chain can be readable and still not be the true cause of the final token. It is also not tool use: no arithmetic is offloaded, so a model can still make a multiplication error inside an otherwise well-structured chain. And it is not free - it multiplies output tokens, which costs money and time. ## The interview framing If you are asked "what is chain-of-thought", the answer an interviewer wants is: intermediate reasoning steps emitted in the output before the answer; the steps are generated text, not computation offloaded elsewhere; they decompose the problem and act as externalised working state; and they pay for themselves on multi-step problems, not on lookups.

  • Does the model actually compute anything differently, or is it just formatting?
    It genuinely computes differently. The intermediate tokens are inputs to every later token, so the model conditions on its own partial results rather than having to hold them implicitly. It also spends more forward passes on the problem. The formatting is a side effect; the extra conditioning and the extra computation are the substance.
  • If the reasoning is wrong but the final answer is right, was the chain useful?
    Sometimes, but you should not trust it. A readable chain is not automatically the true cause of the answer, so a right answer with broken steps may be pattern matching that happened to land. Treat that combination as a warning sign in evaluation rather than a success, and check whether the correct answers survive on similar inputs.
  • Would you show the reasoning steps to end users?
    Usually not verbatim. Steps are verbose, occasionally wrong in embarrassing ways, and in regulated domains they can read as advice you did not intend to give. A common pattern is to generate the reasoning, use it internally or log it for review, and present the reader with the conclusion plus a short justification derived from it.

It is the difference between a student writing only the final number on an exam and writing the working out. The working is not magic - but a student forced to write each line makes fewer silent mistakes, and can look back at line two while computing line three.

saying these in an interview costs you the question

  • Claiming the steps are executed by a calculator or solver
  • Saying chain-of-thought guarantees a correct answer
  • Treating the trace as proof of how the model really decided
  • Assuming it helps equally on single-fact lookups
  • Describing it as a formatting instruction with no effect on accuracy

context

open as a page

Does a confident chain-of-thought trace mean its intermediate steps are true?

level: juniorimportance: must knowfreq 60%

basics

~20 s

No. Every step is predicted text, so a chain can contain invented facts — a factor pair that does not multiply out, a contract subsection that does not exist — while reading as careful and rigorous.

open as a page

What are thinking tokens in an extended-thinking LLM, and are you billed for them?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Thinking tokens are reasoning a model generates before its final answer, on a separate channel. They are generated tokens like any others, so you pay for them and wait for them even when the API returns only a summary, or nothing at all.

open as a page

In self-consistency decoding, how is one final answer chosen from many sampled chains?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Self-consistency throws away the reasoning text and votes only on final answers: each sampled chain contributes its extracted answer, identical answers are grouped into one bucket, and the biggest bucket wins. That is a plurality, not a required majority.

open as a page

In prompting, how does zero-shot CoT differ from few-shot CoT?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Zero-shot CoT adds only an instruction to reason step by step, and the model invents its own reasoning shape. Few-shot CoT puts solved examples in the prompt that demonstrate both how to reason and how the answer should look.

open as a page

Why isn't final-answer accuracy enough to evaluate chain-of-thought reasoning?

level: middleimportance: must knowfreq 62%

basics

~20 s

Final-answer accuracy scores only the last line, so a model that reaches the right number through a broken chain still counts as correct. Step-level scoring catches these right-for-the-wrong-reason solutions, which generalize worst to new problems.

open as a page

Why does emitting intermediate steps improve an LLM's multi-step accuracy?

level: middleimportance: must knowfreq 72%

basics

~20 s

Each generated token costs one pass through a fixed-depth network, so a direct answer must fit the whole computation into a single pass. Writing steps spreads the work across many passes and stores partial results in the context for later tokens to read.

open as a page

Why can a chain-of-thought trace be unfaithful to how the model actually reached its answer?

level: middleimportance: must knowfreq 70%

basics

~10 s

A reasoning trace is generated text, not a log of the computation behind the answer. The model can write a plausible justification while the cue that actually drove its output goes unmentioned.

open as a page

Why normalize sampled answers before counting votes in self-consistency decoding?

level: middleimportance: must knowfreq 52%

basics

~20 s

Without normalization, one correct answer written four different ways lands in four separate buckets and loses to a wrong answer that happened to be formatted consistently. Canonicalising each answer into a comparison key is what makes the vote count meaning rather than string formatting.

open as a page

Why does self-consistency break down if you sample the chains at temperature 0?

level: middleimportance: must knowfreq 70%

basics

~20 s

Temperature 0 is greedy decoding: every sample follows the same highest-probability path, so the chains come back near-identical and the vote counts one opinion n times. Self-consistency needs stochastic sampling so chains take genuinely different reasoning routes.

open as a page

How do you choose an extended-thinking budget when doubling it doubles cost per document?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Treat the budget as an empirical dial, not a preference. Sweep several budgets over a labelled eval set, plot accuracy against spend, and pick the knee where extra reasoning stops buying correctness. Remember the budget is a ceiling the model often underspends.

open as a page

In self-consistency, 19 of 20 chains agree on a wrong answer — what went wrong?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Voting cancels independent errors, not shared ones. Samples from one model with one prompt share priors and a reading of the question, so a misread premise or a wrong prior propagates into every chain. Ensembling reduces variance; it cannot remove bias.

open as a page

What do process reward models reward that outcome reward models don't?

level: middleimportance: should knowfreq 45%

basics

~20 s

An outcome reward model scores only the final answer, so every chain that reaches it gets reinforced, including lucky ones. A process reward model scores each reasoning step, penalising an invalid derivation even when the final answer is right.

open as a page

How fine-grained should chain-of-thought steps be for a multi-step calculation?

level: middleimportance: should knowfreq 47%

basics

~20 s

Aim for one step per operation whose result you would want to check on its own. Steps coarse enough to hide several operations invite a guessed number; steps so fine they narrate digits inflate tokens and add more places for the chain to drift.

open as a page

When does forcing chain-of-thought make an LLM's answer worse rather than better?

level: middleimportance: should knowfreq 52%

basics

~20 s

On atomic tasks the model already handles in one shot — a one-word sentiment tag, a lookup, a simple classification — forced reasoning can talk it out of a correct first answer, drift the output format, and add cost and latency for nothing.

open as a page

With extended thinking enabled, does adding "think step by step" to the prompt still help?

level: middleimportance: should knowfreq 44%

basics

~20 s

Usually not. A reasoning model already runs its own scratchpad pass, so generic step-by-step instructions mostly duplicate it, burn tokens, and can push reasoning-style prose into the visible reply. Task-specific guidance about what to check and how to format still helps.

open as a page

How do you apply self-consistency when the sampled answers are free-form text?

level: middleimportance: should knowfreq 34%

basics

~20 s

Exact-match voting is impossible on prose, so universal self-consistency puts all the sampled responses into one prompt and asks a model to pick the most consistent one. Selection replaces counting — the winner is still one of the candidates, never a new synthesis.

open as a page

Can you reproduce a sampled self-consistency run in CI by fixing the seed?

level: middleimportance: should knowfreq 34%

basics

~20 s

No. A seed pins the sampling draw only if everything else is identical, and on hosted inference it is not: batch composition, kernel and hardware changes, and silent model updates all shift the output. Assert on aggregate metrics with tolerance, not on exact chains.

open as a page

Why does zero-shot CoT often need a second call to extract the answer?

level: middleimportance: should knowfreq 46%

basics

~20 s

A bare step-by-step instruction returns free-form prose with no fixed answer slot. The original zero-shot CoT recipe therefore uses two stages: generate the reasoning, then re-prompt with that text plus an extraction cue that yields a short, parseable answer.

open as a page

How would you test whether a model's stated reasoning actually drove its answer?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Faithfulness is measured by intervening on the trace and watching the answer. Truncate it early, corrupt a step, paraphrase it, or replace it with filler, then check whether the final answer moves. A trace you can break without changing the answer was not doing the work.

open as a page

How do you obtain step-level labels for reasoning traces without hand-annotating every step?

level: seniorimportance: should knowfreq 28%

basics

~20 s

The cheap route samples several continuations from each prefix and labels a step by how often it still reaches the correct final answer, needing no annotator. Human labelling stays for a stratified audit sample used to validate the automated labels.

open as a page

Why can verbalized reasoning steps help a large model but hurt a small one?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A model that cannot produce valid steps still produces fluent ones, then answers consistently with its own broken chain - so errors compound instead of cancelling. Early work saw this below a capability threshold; by 2026 the deciding factor is reasoning-specific training, not raw size.

open as a page

Why can a cosmetic reword of a chain-of-thought prompt shift accuracy on the same task?

level: seniorimportance: should knowfreq 35%

basics

~20 s

A reasoning prompt is not a specification, it is an input that conditions generation. Wording, exemplar order, formatting and where the instruction sits all move the output distribution, so a harmless-looking edit can shift measured accuracy by several points.

open as a page

Why does a chain-of-thought answer flip to a wrong one when the user pushes back?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Preference training rewards agreeable, helpful-sounding replies, so mild pushback like "are you sure?" reads as a signal that the previous answer was unwanted. The model then writes a fresh chain that argues its way to the reversal.

open as a page

How should you handle hidden reasoning traces in logs, retention and multi-turn calls?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Treat a reasoning trace as sensitive conversation data: same access controls, retention window and redaction as the transcript, never surfaced to end users. Across turns, follow the provider's rule on returning thinking blocks unmodified, since editing or dropping them can break the call.

open as a page

When does verifier-based best-of-n reranking beat majority voting over samples?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Reranking wins when a scorer can recognise a correct answer that only one sample produced — hard problems where the modal answer is wrong, or answer spaces too open for counting. Voting wins when no trustworthy verifier exists, because it needs none.

open as a page

What sets end-to-end latency when self-consistency samples n chains in parallel?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The slowest chain, not the average one — a parallel fan-out finishes when its last call returns. Add queueing whenever n exceeds your concurrency or rate limit, plus any retry. Tail latency therefore worsens as n grows even if per-call latency is unchanged.

open as a page

When can few-shot CoT exemplars hurt accuracy on unusual inputs?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Worked examples teach format as well as method. When a new input does not resemble them, the model tends to force it into the demonstrated shape — mapping an unusual case onto the nearest example rather than reasoning about what is actually in front of it.

open as a page

A new CoT prompt lifts GSM8K by 4 points — what would you check before believing it?

level: principalimportance: should knowfreq 33%

basics

~20 s

Check four things before accepting the gain: run-to-run variance with confidence intervals on paired items, harness differences such as shot count and answer extraction, contamination and saturation on a decade-old public set, and whether the gain transfers to the actual product task.

open as a page

How do you choose n for self-consistency on a nightly batch of 3,000 tickets under a fixed budget?

level: principalimportance: should knowfreq 48%

basics

~20 s

Measure vote accuracy against n on a labelled slice: cost grows linearly with samples while accuracy saturates. Pick the n where the marginal accuracy per dollar stops being worth the cost of the errors it prevents, and reserve large n for hard items.

open as a page

showing 1–30 of 33