When does spending more test-time compute stop paying off for a task?
answer
- Roughly logarithmic, then flat
- Ask what the binding constraint is
- Computation helps; missing information does not
- Flat from the origin on easy traffic
- Accuracy per dollar, not accuracy
basics
~20 sWhen the bottleneck stops being computation. Accuracy rises roughly with the logarithm of thinking spend and then flattens; past that knee, and on any task limited by missing information rather than missing reasoning, extra thinking buys latency and cost only.
solid answer
~50 sTwo separate curves matter and conflating them is the usual mistake. On genuinely hard, structured problems — an ARC-AGI-style puzzle, a multi-constraint schedule — accuracy climbs with thinking spend, but roughly logarithmically: each doubling buys less, and the curve flattens at a knee that is task-specific. On easy or under-specified requests the curve is flat from the start, and a trivial lookup that burns thousands of thinking tokens is a pure cost and latency regression. The deciding factor is what limits the answer. If it is computation the model can do — enumerate, check, backtrack — spending pays until the knee. If it is information the model does not have, an ambiguous specification, or taste no checker can score, more deliberation produces more confident wrong answers, not better ones. So measure the curve on your own eval per request class, set the default at the knee rather than at the ceiling, and escalate only where an error is expensive **and** checkable.
go deeper
Know that thinking helps on hard multi-step problems and wastes money on simple ones, and that the amount of help gets smaller as you spend more.
Explain the shape: accuracy rises roughly with the logarithm of compute and then flattens, and it never rises at all when the limit is missing information rather than missing reasoning.
Diagnose over-reasoning in production from thinking-token distributions per route, and justify a default depth from your own eval rather than from published curves or a vendor's recommendation.
Own the allocation policy: buy depth only where errors are expensive and checkable, compare it against retrieval, clearer specification and verifiers at equal cost, and mandate re-measurement after every model change since effort labels are not portable units.
## The two curves When people say "more thinking helps", they are describing one regime and generalizing it to all traffic. Separate them. **The hard-problem curve.** On problems with genuine sequential structure — a puzzle requiring induction over examples, a schedule with interacting constraints, a proof — accuracy really does rise as you spend more inference-time compute on identical weights. It rises **roughly logarithmically**: the first big increment in thinking buys a lot, the next buys less, and eventually the line goes flat. There is a knee, and it is task-specific. **The easy-problem curve.** On lookups, classifications and one-line answers, the curve is flat from the origin. Spending here is not a smaller gain; it is no gain. A route that answers "what timezone is Denver in" and burns three thousand thinking tokens doing it has produced a latency and cost regression with zero accuracy benefit — and occasionally a negative one, when the model talks itself out of a correct first instinct. Real traffic is a mixture of both. A single global effort setting is therefore always wrong for part of your traffic; the only question is which part you have chosen to be wrong about. ## What actually determines the payoff The useful test is: **what is the binding constraint on this answer?** Spending pays when the constraint is *computation the model can perform*. If the model has everything it needs and simply has to enumerate, evaluate, check and backtrack, more tokens are more of that work. Spending does not pay when the constraint is anything else: - **Missing information.** No amount of deliberation recovers a fact outside the model's knowledge or a document not in context. The fix is retrieval or a tool call, not depth. - **Ambiguous specification.** If the request admits several readings, thinking elaborates one reading confidently. The fix is a clarification or a better prompt. - **Unverifiable quality.** If nobody — not the model, not a checker — can tell a better answer from a worse one, the model has no gradient to climb during deliberation. Extra tokens produce polish, not correctness. A closely related asymmetry: deliberation pays most where the model can **check itself**. Arithmetic can be re-derived; a constraint can be re-tested; code can be traced. Where a self-check is possible, extra thinking finds errors. Where the answer is a judgement call, extra thinking mostly generates justification for whatever it first believed. ## Over-reasoning as a production defect Treat this as a real regression class, not a curiosity. Symptoms: p95 latency multiplying after thinking is enabled by default; a bill dominated by output tokens on routes whose answers are short; users abandoning a surface that used to feel instant. The cause is a depth policy applied uniformly to a mixed workload. Diagnosis is straightforward if you instrumented for it: log thinking tokens per response, bucket by route and by request class, and look at the distribution rather than the mean. The pathology is visible as a fat tail on routes whose answers are trivially short — high spend, low output, no accuracy difference against a lower-effort baseline. The remedies are ordinary engineering: drop the depth setting or disable thinking on shallow routes; keep depth for the routes where your eval shows it moves accuracy; make depth a per-request-class decision rather than a deployment-wide one. ## How to find the knee, honestly Published curves are about published benchmarks. Yours will differ, so measure: 1. Build an eval set per request class from real traffic, including the hard tail rather than a convenient sample. 2. Run each class at each available depth setting, several times per item, since output varies run to run. 3. Plot accuracy against actual measured cost and latency — not against the setting's name, which is not a unit and does not transfer across model families. 4. Find where the accuracy line flattens. Set the default there. 5. Re-run this after any model change. A new model can move both the accuracy and the cost at every level, and settings tuned on the old one silently become wrong. The discipline that separates a senior answer from a principal one is step 3: reasoning about accuracy **per dollar and per second**, not accuracy alone. Any depth setting improves accuracy somewhere; the question is what you paid and whether that beat spending the same budget on retrieval, better context, or a verifier. ## Where to spend the budget you have With a fixed budget, put depth where two conditions hold together: **errors are expensive**, and **the task is checkable**. A root-cause analysis on a live incident qualifies on both counts — a wrong conclusion costs an outage extension, and the reasoning can be validated against the evidence. A tag-suggestion feature qualifies on neither. And hold the alternatives in view. More thinking is one lever among several. Giving the model the right document, tightening an ambiguous instruction, or adding a real verifier that catches errors after the fact will often beat a depth increase at the same cost — and unlike depth, those do not have a flat region waiting a doubling away.
- Your eval shows accuracy still rising at the top effort setting. Does that justify shipping at maximum?Only if the marginal accuracy is worth its marginal cost and latency for that request class. A rising line says the knee is beyond your range, not that spending is free. Quantify what the last increment bought — points of accuracy per unit cost and per second of added wait — and compare it against spending the same budget on retrieval, clearer instructions, or a post-hoc verifier.
- How do you catch over-reasoning before it shows up as a bill?Instrument thinking tokens per response and alert on the distribution per route, not the mean. The signature is a fat tail of high thinking spend on routes whose answers are short. Pair it with a periodic shadow run at a lower depth setting on the same traffic: if accuracy is indistinguishable, the extra spend was buying nothing and the default should move down.
- A task is hard but has no automatic checker. Does extra thinking still help?Sometimes, but the evidence is weaker and you must verify it rather than assume it. Deliberation pays most where the model can test a candidate against something. With no check available it tends to elaborate and justify its first direction, which reads as higher quality without being more correct. Measure against human-labelled outcomes before committing budget.
saying these in an interview costs you the question
- More thinking always improves the answer
- One global effort setting fits all traffic
- Benchmark scaling curves transfer to your workload
- Thinking can compensate for missing context
- Over-reasoning is a cost issue only, never accuracy