skip to content

Chain-of-thought on every request costs tokens — how do you decide which ones need it?

level: seniorimportance: should knowfreq 48%

answer

  1. traffic is not uniformly hard
  2. recalled answers versus assembled answers
  3. gate on evidence, not on a guess
  4. validator failure is the best trigger
  5. same output contract on both branches

basics

~20 s

Segment the traffic. Reason only on the slice whose errors are genuinely multi-step; answer directly on lookup and formatting work, where forced deliberation adds latency and can talk the model out of a correct answer. Gate on a cheap difficulty signal and measure each segment separately.

solid answer

~50 s

Blanket reasoning is a flat tax on a workload that is rarely uniformly hard, and on the easy slice it can actively hurt: a spreadsheet-formula generator that was reliable as a direct pattern completion got worse once every request was forced through step-by-step deliberation, because the extra room let the model second-guess a form it already had right. So gate it. Useful gate signals are structural (does the request involve several quantities, interacting conditions, or a multi-part constraint?), a cheap classifier over the input, or — best of all — a validator on the direct answer: if an arithmetic check, schema check or unit test fails, retry that request on the reasoning branch. Keep the output contract identical on both branches so downstream parsing does not fork. Then measure per segment, not in aggregate: accuracy with and without reasoning on each slice, the gate's own error rate, and the cost of the gate relative to what it saves.

go deeper

for a junior

Know that reasoning costs extra tokens and time on every request that uses it, and that not every request needs it. Say you would apply it where the task has multiple steps.

for a middle

Explain the mechanism of the backfire — deliberation adds variance to tasks the model already gets right in one shot — and name concrete gate signals.

for a senior

Show the production design: validator-driven escalation, one shared output contract, per-segment measurement, and asymmetric handling of the two gate error directions.

for a principal

Own the economics and the drift: cost per correct answer across the traffic mix, an operating point justified by the cost of a wrong answer, shadow sampling to catch mix shift, and re-validation of the whole gate on every model upgrade.

## Why blanket reasoning is the wrong default at scale A production workload is almost never uniformly difficult. A typical mix has a long tail of genuinely multi-step requests and a large head of easy ones — restatements, lookups, format conversions, obvious classifications. Applying reasoning to everything charges the whole distribution for the tail's problem. The cost shows up on three axes: generated tokens on every request, latency before the answer exists on every request, and — least expected — accuracy loss on the easy head. ## Where reasoning backfires The accuracy loss is the part that surprises people, so be ready to explain the mechanism. When a task is essentially a reliable direct mapping — the model has seen the pattern thousands of times and completes it correctly in one shot — inserting deliberation adds a stretch of generated text in which the model can raise objections, consider alternatives, and drift from the form it would otherwise have produced. A spreadsheet-formula generator is the canonical case: asked plainly, it emits the right formula; asked to reason first, it sometimes reasons its way into a more "thorough" but wrong construction, or into prose that no longer matches the expected output shape. The reasoning did not make it dumber; it created room for variance on a task that had none. The general rule: reasoning helps when the answer must be *assembled* and hurts when the answer is *recalled*. ## Building the gate A gate decides, per request, which branch runs. Options, roughly in order of cost: - **Structural signals.** Cheap, deterministic features of the request: does it contain multiple numeric quantities, several constraints, a comparison, a conditional, a reference to more than one document? These are crude but free and often surprisingly effective as a first cut. - **A cheap classifier.** A small model or a trained classifier scores difficulty and only the high scores take the reasoning branch. This costs a small fixed amount per request; it is worth it only if the saved reasoning tokens exceed that amount across the whole mix, which means it pays best when the easy head is large. - **Model self-assessment.** Ask for a quick confidence or difficulty tag alongside the direct answer. Convenient, but confidence self-reports are weakly calibrated, so treat this as a soft signal. - **Validator-driven escalation.** Answer directly, run a real check — recompute the arithmetic, validate against a schema, run the generated code, verify the cited passage exists — and only escalate to the reasoning branch when the check fails. This is usually the strongest option because the trigger is evidence of an actual error rather than a guess about difficulty. Its cost is a second call on the failing slice, so it is best where failures are a minority and the check is cheap and trustworthy. ## Keep one contract Whatever the gate, both branches must return the same output shape. The moment the reasoning branch returns a differently structured response, every downstream consumer — parser, logger, evaluator, UI — grows a fork, and the two paths drift apart in ways that only show up in production. Constrain the reasoning branch to hide its deliberation and emit the identical answer block. ## Measuring it honestly Aggregate accuracy hides everything here, because the gate deliberately changes what happens per segment. Report: - Accuracy **per segment**, with and without reasoning. This is what tells you the gate is pointing at the right slice; if the "hard" segment shows no benefit from reasoning, the gate is selecting on the wrong signal. - **Gate errors, split by direction.** A hard request routed to the cheap branch produces a wrong answer; an easy request routed to the expensive branch merely costs money and time. Those costs are asymmetric, and the asymmetry sets the operating point — most systems should lean toward over-escalating. - **Cost and latency percentiles**, not means. Escalation makes the tail heavier; if p95 matters to the product, a gate that escalates 15% of traffic changes p95 far more than it changes the average. - **Gate overhead** as a share of what it saves. A gate that costs nearly as much as the reasoning it avoids is not worth its complexity; simplify to always-reason on that route. ## Operating it over time Traffic mix drifts. A gate tuned when 10% of requests were hard silently misroutes when a new customer segment pushes that to 40%, and the symptom is a quiet accuracy decline rather than an alert. Sample the cheap branch periodically — run a slice through the reasoning branch as a shadow and compare — so drift shows up as a measured divergence rather than a support ticket. And re-check the whole gate on every model upgrade: a newer model may already handle the segment that previously needed escalation, at which point the gate is spending money to solve a problem that no longer exists.

  • How do you choose the gate's operating point?
    By the asymmetry of the two errors. Routing a hard request to the cheap branch produces a wrong answer; routing an easy one to the expensive branch costs tokens and latency. In almost every product the first is worse, so you bias toward over-escalating and let cost, not accuracy, absorb the imprecision. Then track how much traffic that actually escalates.
  • What if the gate itself costs about as much as the reasoning it avoids?
    Then delete the gate and always reason on that route. A gate only pays when the saved tokens across the whole mix clearly exceed its per-request cost, which requires a large easy head. If the mix is mostly hard, the gate is complexity that buys nothing and adds a second failure mode — a misrouting bug on top of the model's own errors.
  • How would you notice that a gate tuned six months ago has stopped fitting the traffic?
    Shadow sampling. Periodically run a slice of the cheap-branch traffic through the reasoning branch as well and compare the answers. A rising divergence rate means the gate is now sending genuinely hard requests down the cheap path. Aggregate accuracy will not show this in time, because the mix shift and the accuracy drop move together and look like noise.

It is triage, not a policy of admitting everyone: most arrivals need a five-minute look, and the expensive workup is reserved for the cases whose symptoms actually warrant it.

saying these in an interview costs you the question

  • Turning reasoning on globally because it helped on hard cases
  • Assuming more reasoning can never reduce accuracy
  • Gating on message length as a difficulty proxy
  • Letting the reasoning branch return a different output shape
  • Judging the gate on aggregate accuracy instead of per segment

context