skip to content

How do you choose an extended-thinking budget when doubling it doubles cost per document?

level: seniorimportance: must knowfreq 52%

answer

  1. a ceiling, not a spend
  2. measure, do not guess
  3. accuracy against spend, plotted
  4. find the knee, sit just past it
  5. tiny budgets can be worse than none

basics

~20 s

Treat the budget as an empirical dial, not a preference. Sweep several budgets over a labelled eval set, plot accuracy against spend, and pick the knee where extra reasoning stops buying correctness. Remember the budget is a ceiling the model often underspends.

solid answer

~50 s

The budget is a cap on reasoning tokens, so it sets a worst case, not a per-request charge — on easy inputs the model stops early and you pay less than the ceiling. That means you cannot reason about it from the number alone; you have to measure. Build a labelled eval set from real documents, run it at several budgets (say off, small, medium, large), and plot task accuracy and p95 latency against measured reasoning tokens. Almost always there is a knee: a range where accuracy climbs steeply, then a long flat stretch where you are paying for tokens that change nothing. Pick just past the knee and leave headroom in the overall output cap so the answer itself is never truncated. Then keep watching the *distribution* of reasoning tokens in production, because a shift in document mix moves your real spend even with the ceiling unchanged.

go deeper

for a junior

Know that the budget is an upper bound on reasoning, that it drives both cost and delay, and that the honest way to pick one is to try several and compare results on real examples.

for a middle

Be able to describe the sweep concretely: fixed prompt and model, four budgets including off, accuracy plus token count plus p95 latency recorded per run, knee read off the curve.

for a senior

Show judgment past the chart — error asymmetry, latency SLOs, answer headroom in the output cap, and re-running the sweep when the model or document mix changes. Mention that very small budgets can underperform none.

for a principal

Frame it as a quality-cost frontier the business chooses a point on, and own the governance: who re-runs the sweep, what alert fires when the reasoning-token distribution drifts, and which document classes are worth buying accuracy for.

## Why this is a measurement problem, not a taste problem A contract-summarisation service is a good lens. Each document produces a short structured summary, and the accuracy that matters is whether obligations, dates and termination clauses came out right. Turning extended thinking up genuinely helps on gnarly documents. It also, in the case that makes this an interview question, doubles the per-document bill. There is no principled way to pick the number from first principles — the only defensible answer is that you measured the curve. ## The budget is a ceiling, not a quota The first thing to get right, because candidates routinely get it wrong: a thinking budget is an upper bound. A model handed a large allowance on a two-page NDA will spend a fraction of it and move on. So the average cost of a fleet running at a high ceiling is usually well below the ceiling, and the ceiling mostly determines your *tail*. This has two consequences. Raising the budget does not multiply the average bill by the same factor it multiplies the cap. And the number you should be watching is the measured reasoning-token distribution, not the configured maximum. The same is true of coarse effort levels where a provider exposes low/medium/high instead of a token count: higher effort raises how much reasoning the model is willing to do, not how much it must. ## Building the curve The procedure is unglamorous and it is what a senior interview wants to hear: 1. **Assemble a labelled eval set** from real production documents, deliberately stratified — short and long, clean and scanned, standard and unusual clause structures. A hundred well-chosen items beats a thousand random ones. 2. **Fix everything else.** Same prompt, same model, same temperature. The only variable is the budget. 3. **Sweep at least four settings**, including thinking off. Off is a real candidate and it is the cheapest, so it has to be on the chart. 4. **Record three axes per run**: task accuracy on your grader, measured reasoning tokens (mean and p95), and end-to-end latency (mean and p95). Cost follows from tokens. 5. **Plot accuracy against measured spend.** Read the knee. What you nearly always see is a steep climb from off to a modest budget, then diminishing returns, then flat. The flat region is money you are lighting on fire. Occasionally you see something more interesting — accuracy dipping at very small budgets, because a truncated reasoning pass is worse than none, having started an analysis it could not finish. That is a real effect and a good thing to name. ## Choosing the operating point Pick just past the knee, then apply the constraints the chart cannot see: - **Latency ceiling.** If a human is waiting on the summary, p95 latency may bind before cost does. In a batch pipeline it does not bind at all. - **Error asymmetry.** A missed termination clause may cost far more than a few cents of tokens. Where the downside of a wrong answer is large and the volume is small, buy past the knee deliberately. Where volume is enormous and errors are cheaply caught downstream, sit before it. - **Answer headroom.** The overall output cap must exceed the thinking budget plus the longest expected summary, or you will trade accuracy for truncation. - **Floor.** Consider whether some document classes should run with thinking off entirely; the sweep tells you which ones lose nothing. ## Living with it A budget chosen once and never revisited decays. Three things move it: the document mix shifts (a new customer sends 90-page master agreements), the prompt changes (a better prompt often flattens the curve, so the old budget is now overspend), and the model changes (a new version may reach the same accuracy with less reasoning, or ignore the old ceiling entirely). Re-run the sweep on model upgrades — it is a half-day job and it routinely finds a cheaper operating point. In production, alert on the *reasoning-token distribution*, not just total spend. A rising p95 with flat volume means inputs got harder, which is usually the earliest signal that something upstream changed. ## The honest caveat There is no published constant here. How much reasoning a task needs is a property of the task, the model and the prompt together, and it is genuinely contested how much of the gain from a large budget is reasoning versus simply more chances to notice something. Saying that plainly, and then describing the sweep, is a stronger answer than quoting a number.

  • Why can a very small thinking budget produce worse answers than turning thinking off entirely?
    Because the reasoning pass gets cut mid-analysis. The model commits to a decomposition, explores part of it, and is then forced to answer from an unfinished state — which can be worse than answering directly from the prompt, where it never split the problem in the first place. If your sweep shows a dip at the low end, do not interpolate through it; either buy past it or run with thinking off.
  • Your average reasoning spend is well under the configured ceiling. Is the ceiling therefore harmless?
    Not harmless — it governs the tail. The ceiling sets your worst-case latency and worst-case per-request cost, which is what capacity planning and timeouts are sized against. A generous ceiling with a low average is fine for a batch pipeline and dangerous behind a synchronous endpoint with a hard SLO, because the slow tail is exactly where the ceiling binds.
  • How would you split the budget across easy and hard documents rather than using one setting?
    Segment by a cheap upfront signal — page count, clause density, whether OCR was involved, or a small classifier over the first page — and attach a budget per segment. Validate that the segmentation actually predicts the accuracy gain on your eval set; if the curves for the segments look the same, the split is complexity for nothing. Keep the number of tiers small so the behaviour stays explainable.

saying these in an interview costs you the question

  • Set the maximum budget so quality is never limited
  • Cost scales exactly with the configured ceiling
  • More reasoning tokens always mean better answers
  • Pick the budget once and never revisit it
  • Ignore latency because only cost is being tuned

context