skip to content

How fine-grained should chain-of-thought steps be for a multi-step calculation?

level: middleimportance: should knowfreq 47%

answer

  1. granularity is a dial, not a style
  2. one step per checkable value
  3. too coarse hides a guessed number
  4. too fine buys tokens, not accuracy
  5. constraints get their own early step

basics

~20 s

Aim for one step per operation whose result you would want to check on its own. Steps coarse enough to hide several operations invite a guessed number; steps so fine they narrate digits inflate tokens and add more places for the chain to drift.

solid answer

~50 s

Granularity is a real dial, and the useful rule is: one step per independently checkable state change. Take a paediatric dose - convert the weight from pounds to kilograms, multiply by the mg/kg rate, compare the result against the maximum single-dose ceiling, then round to a dispensable increment. Four steps, each with a value a pharmacist could verify in isolation. Collapse that to 'work out the dose' and the model must produce the answer in one prediction, which is exactly the situation chain-of-thought exists to avoid. Go the other way - narrating each digit of the multiplication - and you pay many more tokens, give the model more opportunities to lose the thread, and gain nothing, because arithmetic errors are not usually fixed by verbosity. Constraints deserve their own explicit step, stated before the value they constrain, since later tokens can only condition on earlier ones.

code

markdown · 13 lines
markdown
Too coarse (one prediction, no checkable values):
  Work out the dose: 200 mg.

Well grained (one checkable value per step):
  Weight: 34 lb / 2.205 = 15.4 kg
  Dose at rate: 15.4 x 15 mg/kg = 231 mg
  Ceiling: max single dose is 200 mg; 231 > 200, so cap at 200 mg
  Rounding: 200 mg is a multiple of 2.5 mg -> 200 mg

Too fine (tokens without new checkable state):
  15.4 times 15: 15 times 15 is 225, then 0.4 times 15 is 6,
  then 225 plus 6 is 231, and 231 is the product, and now
  I will consider the ceiling...

go deeper

for a junior

Know that a reasoning step should carry one clear result, and that lumping several operations into one line makes the answer harder to trust or to check.

for a middle

Explain the two-sided tradeoff: too coarse forfeits the decomposition, too fine costs tokens and invites drift. Give the rule of one step per independently checkable value, with units written down.

for a senior

Show that you tune grain from failure traces rather than taste, promote repeatedly-missed constraints to explicit early steps, and hand exact arithmetic to a tool instead of to more prose.

for a principal

Own reasoning shape as a budgeted, measured property of a system: accuracy per token, reviewability for regulated domains, and a standard step schema that downstream verification and logging can rely on.

## Granularity is a design decision, not an accident Once you have decided the model should reason before answering, a second question follows immediately: at what resolution? One line per arithmetic operation? One paragraph per stage of the problem? A single sentence of justification? People treat this as style. It is not - it changes accuracy, cost and reviewability. ## The working rule One step per state change you would want to verify independently. A paediatric dosing calculation makes this concrete. The task: a child weighs 34 lb, the drug is dosed at 15 mg/kg, single doses are capped at 200 mg, and the syrup is dispensed in 2.5 mg increments. A well-grained chain looks like: 1. Weight: 34 lb / 2.205 = 15.4 kg. 2. Rate applied: 15.4 kg x 15 mg/kg = 231 mg. 3. Ceiling check: 231 mg exceeds the 200 mg single-dose maximum, so the dose is capped at 200 mg. 4. Rounding: 200 mg is already a multiple of 2.5 mg, so the dose stands. Each line has a value a reviewer - or an automated checker - can confirm alone. That is the test. ## Too coarse "Work out the appropriate dose for this child" gives the model one prediction to make. Everything the technique was supposed to buy - serial computation, partial results held in context, easy local predictions - is forfeited. Worse, coarse steps hide which sub-result failed. When the answer comes back as 231 mg, you cannot see whether the model forgot the ceiling or never computed it, because it never wrote either down. Coarse chains also lose constraints silently. A cap, a jurisdictional rule, a maximum weight tier - if it is not a step of its own, it competes for attention with everything else in a single generation and is frequently dropped. Constraints that must be honoured should be stated as steps, and stated before the value they constrain, because generation is one-directional: a ceiling mentioned after the dose cannot change the dose already written. ## Too fine The opposite failure is less discussed and just as real. Narrating a multiplication digit by digit, or restating the problem before every line, has costs: - Token cost and latency scale with chain length, and reasoning tokens are often the dominant cost in a request. - Every additional generated step is another opportunity to sample something wrong or to drift off the task. - Very long chains push earlier content further from the answer and, on long inputs, can degrade what the model attends to. - Verbosity does not fix arithmetic. If the model multiplies badly, writing more prose about the multiplication rarely rescues it; a calculator tool does. There is also a subtle harm: over-fine chains encourage the model to elaborate rather than to decide. On tasks with a clear answer, a chain that keeps qualifying itself often ends up hedging in the final answer too. ## How to choose in practice A few heuristics that hold up: - Grain to the unit of verification. If your evaluation checks the converted weight, make the converted weight a step. - Grain to the unit of failure. Look at real failures: if the model keeps skipping the surcharge, promote the surcharge to its own step. - Keep constraints, units and assumptions explicit and early. Units in particular: a chain that writes '15.4 kg' rather than '15.4' is measurably harder to misuse two lines later. - Prefer structure over prose for anything mechanical. A short labelled line per step is easier for both the model and a downstream parser than a flowing paragraph. - Where a stage is genuinely a judgement rather than a calculation - 'assess whether this shipment needs hazmat handling' - a paragraph is the right grain, because there is no discrete value to check. - Do not fix granularity by instruction alone if you can demonstrate it. Showing the shape you want is more reliable than describing it, and the shape you show is the shape you get back. ## The tuning loop Treat granularity as something you measure. Run your evaluation set, look at the failed traces, and ask of each failure: was the step that went wrong too coarse to be checked, or was the chain so long it lost the thread? Coarse failures get split; long-chain failures get compressed or handed to a tool. That loop converges quickly, and it is the answer an interviewer is really probing for - whether you tune reasoning shape empirically or by taste.

  • Where in the chain should a hard constraint like a maximum dose appear?
    Before the value it constrains, as its own step. Generation is autoregressive, so a ceiling stated after the computed dose can only comment on it, not change it. Stating the ceiling first, then computing, then explicitly comparing, gives the comparison something concrete to condition on and leaves a line a reviewer or automated checker can verify on its own.
  • How would you decide the right grain empirically rather than by intuition?
    Run a fixed evaluation set at two or three grains and read the failing traces. Failures where one line hides several operations mean the grain is too coarse - split that step. Failures where the model drifts, contradicts an earlier line, or blows the latency budget mean it is too fine - compress or offload. Track accuracy against tokens per request, not accuracy alone.
  • Does finer granularity help arithmetic errors specifically?
    Only up to a point. Splitting a compound expression into separate operations genuinely helps, because each becomes an easier prediction. Narrating the internals of a single multiplication mostly does not - the model is still predicting digits. For exact arithmetic the reliable fix is a calculator tool, with the chain deciding what to compute and how to interpret the result.

saying these in an interview costs you the question

  • Treating step granularity as cosmetic formatting
  • Assuming more steps always means more accuracy
  • Leaving a hard constraint implicit rather than as a step
  • Stating a limit after the value it is supposed to bound
  • Splitting arithmetic into digits instead of using a tool

context