skip to content

In LLM prompting, what is chain-of-thought and how does it change the output?

level: juniorimportance: must knowfreq 82%

answer

  1. the model writes before it answers
  2. steps live in the output text
  3. worked solution, then the result
  4. generated tokens, no calculator involved
  5. helps when step three needs step two

basics

~20 s

Chain-of-thought means having the model write out intermediate reasoning steps before its final answer instead of answering immediately. The output becomes a short worked solution followed by the result, which raises accuracy on problems needing several dependent steps.

solid answer

~50 s

Chain-of-thought prompting asks the model to produce the intermediate steps of a solution in its output, before stating the final answer. Instead of jumping straight to a freight quote, the model writes that the shipment falls in the 500-1000 kg weight tier, that the tier's base rate is X, that the fuel surcharge is a percentage of that base, that VAT applies to the sum, and only then gives the total. The steps are ordinary generated tokens - nothing symbolic is executed - but because each step is written into the context, the tokens that follow can read and build on earlier results. It pays off on problems whose answer depends on several dependent sub-results; on a single-fact lookup it mostly adds tokens. Most current frontier models are post-trained to reason this way by default, so it is now as much a model behaviour as a prompt technique.

code

markdown · 12 lines
markdown
Direct answer:
Q: A 640 kg shipment. Quote it.
A: EUR 412.80

Chain-of-thought answer:
Q: A 640 kg shipment. Quote it.
A: 640 kg falls in the 500-1000 kg tier, rate EUR 0.52/kg.
   Base charge: 640 x 0.52 = EUR 332.80
   Fuel surcharge 12% of base: EUR 39.94
   Subtotal: EUR 372.74
   VAT 10.75%: EUR 40.06
   Total: EUR 412.80

go deeper

for a junior

Be able to say plainly that chain-of-thought means the model writes intermediate steps in its output before the final answer, and that this helps on problems with several dependent steps. Give a concrete before-and-after example.

for a middle

Explain that the steps are generated tokens that later tokens condition on, so the model both decomposes the problem and keeps partial results in context. Be clear that no external computation happens.

for a senior

Show judgement about when the extra tokens are worth their cost, what you do with the trace once you have it (log, review, discard), and why a readable chain is not evidence of a faithful one.

for a principal

Frame it as a design choice about where reasoning lives in your system: in the model's output, in orchestration code, or in a tool. Own the cost, latency, exposure and auditing consequences of that choice across a product.

## The idea in one line Chain-of-thought (CoT) is the practice of getting a language model to emit the intermediate steps of a solution as part of its output, ahead of the final answer, rather than producing the answer directly. A language model generates text one token at a time, each token conditioned on everything already in its context - the prompt plus whatever it has written so far. "Direct answering" means the tokens right after the question are the answer itself. Chain-of-thought means the tokens right after the question are a worked solution, and the answer comes at the end of it. ## What it looks like Take a freight quote. A direct answer is a single number: "EUR 412.80." A chain-of-thought answer is a short sequence of stated sub-results: - The shipment weighs 640 kg, so it falls in the 500-1000 kg tier. - That tier's base rate is EUR 0.52 per kg, so the base charge is 640 x 0.52 = EUR 332.80. - The fuel surcharge is 12% of the base charge: EUR 39.94. - Subtotal is EUR 372.74. - VAT at 10.75% on the subtotal is EUR 40.06. - Total: EUR 412.80. Every one of those lines is generated text. No calculator ran. The model is still predicting the next token; it is simply predicting tokens that happen to spell out a decomposition of the problem. ## Why writing the steps matters Two things change once the steps are on the page. First, the problem is decomposed. Predicting "the base charge is 332.80" given the tier and the rate is a much easier next-token prediction than predicting the final total given only the raw shipment description. Each written step turns one hard prediction into several easy ones. Second, the steps become context. Generated tokens are appended to the sequence, so every later token can attend back to them. The intermediate results are stored in the output itself rather than having to be held implicitly inside one pass of the network. If the model needs the subtotal to compute VAT, it can literally read the subtotal it just wrote. ## Where it helps and where it does not CoT is a tool for problems with dependent sub-results: arithmetic with several operations, multi-hop questions (find the coach, find the coach's team, find that team's stadium), constraint checks, unit conversions, date and timezone arithmetic. The common shape is that you cannot compute step three without the output of step two. It does far less for problems that are a single retrieval or a single classification - naming a capital city, extracting a phone number from a paragraph, choosing between two synonyms. There is no chain to unroll, so the extra tokens buy little except cost and latency. ## Elicited versus native The technique first became well known as something you asked for in the prompt: show the model worked examples, or instruct it to reason before answering. Since roughly 2024-2025, models have increasingly been post-trained (via supervised traces and reinforcement learning on verifiable tasks) to produce intermediate reasoning on their own for hard inputs. In practice today you often are not switching CoT on so much as shaping it: deciding how much reasoning a task warrants, what structure the steps should have, and whether the steps belong in the user-visible answer at all. The underlying mechanism is unchanged - tokens generated before the answer, conditioned on by the answer. ## What CoT is not It is not a proof, and it is not a guaranteed window into how the model actually arrived at the answer; a chain can be readable and still not be the true cause of the final token. It is also not tool use: no arithmetic is offloaded, so a model can still make a multiplication error inside an otherwise well-structured chain. And it is not free - it multiplies output tokens, which costs money and time. ## The interview framing If you are asked "what is chain-of-thought", the answer an interviewer wants is: intermediate reasoning steps emitted in the output before the answer; the steps are generated text, not computation offloaded elsewhere; they decompose the problem and act as externalised working state; and they pay for themselves on multi-step problems, not on lookups.

  • Does the model actually compute anything differently, or is it just formatting?
    It genuinely computes differently. The intermediate tokens are inputs to every later token, so the model conditions on its own partial results rather than having to hold them implicitly. It also spends more forward passes on the problem. The formatting is a side effect; the extra conditioning and the extra computation are the substance.
  • If the reasoning is wrong but the final answer is right, was the chain useful?
    Sometimes, but you should not trust it. A readable chain is not automatically the true cause of the answer, so a right answer with broken steps may be pattern matching that happened to land. Treat that combination as a warning sign in evaluation rather than a success, and check whether the correct answers survive on similar inputs.
  • Would you show the reasoning steps to end users?
    Usually not verbatim. Steps are verbose, occasionally wrong in embarrassing ways, and in regulated domains they can read as advice you did not intend to give. A common pattern is to generate the reasoning, use it internally or log it for review, and present the reader with the conclusion plus a short justification derived from it.

It is the difference between a student writing only the final number on an exam and writing the working out. The working is not magic - but a student forced to write each line makes fewer silent mistakes, and can look back at line two while computing line three.

saying these in an interview costs you the question

  • Claiming the steps are executed by a calculator or solver
  • Saying chain-of-thought guarantees a correct answer
  • Treating the trace as proof of how the model really decided
  • Assuming it helps equally on single-fact lookups
  • Describing it as a formatting instruction with no effect on accuracy

context