When do you escalate from a direct answer to chain-of-thought or self-consistency?
answer
- cheapest rung that passes
- each rung multiplies output tokens
- voting needs a short discrete answer
- random error votes away, systematic does not
- climb one rung, on evidence
basics
~20 sEscalate only when the cheaper rung measurably fails. Direct answers suit lookup and formatting; chain-of-thought pays on multi-step reasoning; sampling several chains and voting pays only when the answer is a short discrete value worth the multiplied cost.
solid answer
~60 sTreat these as a ladder of increasing cost. Rung one is a direct answer — cheapest, right for lookup, extraction and formatting. Rung two adds reasoning before the answer, which pays when the errors are genuinely multi-step: arithmetic over several quantities, rule application with interacting conditions, constraint checking. Rung three samples the reasoning several times and takes the answer that appears most often, which multiplies output cost by the number of samples and only works when the answer is a short comparable value — a label, a number, a verdict — because prose drafts have nothing to count. Rung four stops asking one call to do everything and splits the task into separate steps with checked handoffs. The rule is to climb one rung at a time, on evidence: for a moderation classifier, "reason briefly about which policy clauses apply, then output only the verdict" often cuts false negatives at maybe triple the output tokens, and that is a good trade — blanket sampling on the same traffic usually is not.
go deeper
Know the order — answer directly, then reason before answering, then sample and vote — and that each step costs more. Say that you would try the cheap option first.
Explain the multiplier of each rung and the task shape it requires, especially why voting needs a short discrete answer and cannot aggregate prose.
Demonstrate the diagnostic: decide whether the remaining errors are random or systematic before climbing, because sampling only fixes the random kind, and weigh latency alongside token cost.
Frame it as unit economics — cost per correct answer against the cost of an error, per traffic segment — and be willing to leave the ladder entirely when the real fix is the task boundary or the input data.
## The ladder Most prompting decisions reduce to choosing how much computation to spend per request. Framing the options as rungs makes the tradeoff explicit and stops the common failure of jumping straight to the most expensive technique. **Rung 1 — direct answer.** The model answers immediately. Cheapest and lowest latency. Correct choice for retrieval-shaped work, extraction, reformatting, translation, short classification where the label is obvious from surface features. Any technique that adds tokens here is pure cost. **Rung 2 — reason, then answer.** The model produces intermediate work before committing. Costs roughly the length of the reasoning in extra generated tokens and the matching latency. Pays where the answer depends on several interacting facts: multi-step arithmetic, applying a rule set with exceptions, checking a plan against constraints, deciding between two close categories where the distinguishing evidence has to be assembled. **Rung 3 — reason several times, then aggregate.** Run the reasoning repeatedly with enough variation that the chains differ, then take the answer that appears most often. Cost is multiplied by the number of samples — five samples means roughly five times the output tokens, and latency is either five times worse or you pay for parallel capacity. The gain comes from the fact that wrong chains tend to fail in different ways while correct chains converge on the same result. **Rung 4 — decompose.** Stop asking one call to do everything. Split the task into separate steps, check each handoff, and let cheap deterministic code do the parts that are not language problems. This is the rung that changes engineering cost, not just token cost. ## What determines when you climb Two numbers govern it: **the cost of an error** and **the multiplier of the next rung**. If a wrong answer costs a cent of wasted attention, a five-times multiplier for a few points of accuracy is a bad trade. If a wrong answer means a harmful post stays up or a wrong figure reaches a customer, it can be an easy one. The third input is whether the technique *fits the task shape*: - Rung 2 needs errors that are actually reasoning errors. If the model is failing because it lacks a fact, or because the instruction is ambiguous, reasoning will produce a confident wrong chain and change nothing. - Rung 3 needs a **short, comparable answer**. Voting works over labels, numbers and verdicts. It does not work over a three-paragraph email, because no two drafts are the same string and there is nothing to count. Aggregating free-form output requires a judge or a rubric instead, which is a different technique with a different cost profile and its own biases. - Rung 4 needs the task to decompose cleanly. Forcing a split on a task whose steps are mutually dependent trades one hard call for several fragile ones. ## A worked case A content-moderation classifier is emitting too many false negatives — borderline items slipping through as ALLOW. Rung 1 is where it starts. Moving to rung 2 with a *brief* reasoning instruction ("name the policy clauses that could apply, judge each, then output only the verdict") tends to help here because the failure is a rule-application failure: the model was pattern-matching on tone instead of checking clauses. Output tokens perhaps triple; the verdict itself is unchanged in shape, so nothing downstream moves. Does rung 3 follow? Only if measurement says the remaining errors are unstable — the same borderline item classified differently across runs. Voting suppresses that instability. If the remaining errors are *consistent* — the model reliably reads one clause the wrong way — sampling five times gives you the same wrong answer five times and simply costs five times more. That is the single most useful diagnostic before climbing: is the error random or systematic? Sampling fixes randomness. Better instructions, better definitions, or decomposition fix systematic error. ## Common mistakes - **Starting at the top.** Building sampling into the first version guarantees you cannot tell whether it was needed, and it bakes a multiplier into your unit economics. - **Climbing on anecdote.** Two examples that improved is not evidence; you need the same fixed set of cases scored at each rung. - **Stacking rungs on a rung that did not help.** If reasoning produced no measured gain, sampling that reasoning multiplies a null result. - **Ignoring latency.** Cost is often the number people quote, but the perceived tradeoff for an interactive product is usually latency, and rung 3 hurts it hardest. - **Forgetting the model already climbed.** On models that deliberate internally, part of rung 2 is already applied by default, which changes where the ladder starts. ## What good looks like A strong answer names the rungs, gives each one's multiplier, states the task shapes each requires, and insists on climbing one rung at a time against a fixed set of cases. It also names the exit: sometimes the right move is to leave the ladder entirely and fix the instruction, the input data or the task boundary instead.
- What breaks if you apply majority voting to a free-form writing task?There is nothing to count. Voting assumes several samples produce the same comparable answer, which holds for a label, a number or a verdict and never for prose — every draft differs in wording. To aggregate free-form output you need a judge model or a rubric to score candidates, which is a different technique with its own cost and its own biases toward length and style.
- Reasoning helped, so you add sampling and accuracy barely moves. What does that tell you?That the remaining errors are systematic, not random. Sampling suppresses variance — it only helps when wrong chains disagree with each other and correct ones converge. If the model reliably misreads the same rule, five samples reproduce the same wrong answer at five times the cost. The fix is better instructions, clearer definitions or decomposition, not more samples.
- Why is "just use a stronger model" not simply the top rung?It is a different axis. A stronger model raises the per-token price of all traffic, including the easy majority; a technique raises token count only where you apply it. They also compose — a stronger model can make rung 2 unnecessary or make rung 3 redundant. The honest answer is to measure both against the same cases and compare cost per correct answer, not accuracy alone.
saying these in an interview costs you the question
- Always use self-consistency, it is strictly better
- Chain-of-thought is free because the input is cached
- Majority voting works fine on free-form prose answers
- Climbing to a more expensive rung on two eyeballed examples
- Adding sampling on top of reasoning that showed no measured gain