skip to content

How does a reasoning effort level differ from setting a fixed thinking-token budget?

level: middleimportance: must knowfreq 55%

answer

  1. A policy, not a number you pick
  2. The model sees the difficulty; you do not
  3. Advisory shaping, not a hard cap
  4. Budget the distribution, not the call
  5. Hard limits come from elsewhere

basics

~20 s

An effort level sets a policy — think shallowly, moderately or deeply — and lets the model decide per request how far to go. A fixed token budget makes you guess a number in advance for a difficulty you cannot see yet.

solid answer

~50 s

A fixed budget puts the depth decision on the caller: you pick a thinking-token count before you know whether the request is trivial or hard, so you either starve difficult requests or overspend on easy ones. An effort level — `reasoning_effort`, `thinking_level`, an `effort` field, depending on the provider — inverts that. You declare an intent band and the model adapts within it, stopping early when the problem is easy and going deeper when it is not. The practical consequences matter. Effort is **advisory, not a cap**: it shapes the distribution of thinking tokens, it does not guarantee any particular request's spend, so hard limits still come from output-token limits and your own budget enforcement. Budgeting becomes statistical — you reason about p50 and p95 thinking tokens for a route, not a per-call number. And on many reasoning models the sampling knobs are locked, so effort is the main depth control you have.

go deeper

for a junior

Know that the effort setting tells the model how hard to think, that it is set per request, and that higher effort means more tokens, more money and more waiting.

for a middle

Explain the inversion: a fixed budget forces the caller to guess depth before seeing difficulty, while an effort level lets the model adapt within a band. Note that effort is advisory and that thinking still counts against the output limit.

for a senior

Show operational judgment — instrumenting the thinking-token distribution per route, setting effort per request class rather than globally, and enforcing spend with real limits because the dial is not one.

for a principal

Own the portability problem: effort labels are provider-internal policy, not units, so they must be re-measured on your own evals after any model change, and your architecture should not assume a stable mapping from level to cost or quality.

## Two ways to ask for depth When thinking models first shipped, the control was a number: you passed a thinking-token budget and the model was allowed roughly that much room to deliberate. It was a reasonable first design and it had one structural flaw — **you have to choose the number before you know the difficulty of the request**. A support endpoint receiving both "what timezone is Denver in" and "reconcile these three conflicting incident timelines" gets one budget for both. Set it low and the hard request is cut off mid-deliberation. Set it high and the trivial one is billed for deliberation it never needed. As of mid-2026 the frontier providers have converged on the other design: a small set of discrete **effort levels** — low through high or max, named `reasoning_effort`, `thinking_level` or an `effort` field depending on whose API you are calling. You state how much you care; the model decides per request how much to actually spend inside that band. ## Why moving the decision to the model helps The model is the only party in the loop that has read the request. It can tell after a few hundred tokens whether the problem has real structure or whether the answer is a lookup, and it can stop. That is a strictly better place for the decision than your call site, which sees a route name and a payload. The effect is that the same effort setting produces wildly different actual spends across a realistic traffic mix — which is the point. Low effort on an easy question may cost almost nothing; high effort on a genuinely hard one may run long. You have expressed a preference over the *trade*, not a quantity. ## What effort is not This is where interviews probe, so be precise: - **It is not a hard cap.** Effort shapes a distribution. It does not promise that any individual request stays under a number. If you need a ceiling, that comes from the response's output-token limit and from your own accounting — separate mechanisms with different semantics. - **It is not a model switch.** Same weights, same endpoint; you are buying inference-time compute, not a bigger model. - **It is not a quality dial with a monotone payoff.** Higher effort raises accuracy on problems with sequential structure and does approximately nothing — at real cost — on lookups and classifications. On some easy inputs it can actively hurt, as the model second-guesses a correct first instinct. - **It is not free of interaction effects.** Deep thinking consumes the output-token allowance. A high effort setting under a tight output limit is a good way to get truncated responses that never reach the answer. ## How this changes your operational thinking **Budget statistically.** With a fixed budget you could multiply a number by a request count. With effort levels you cannot, and pretending otherwise is how a bill surprises you. Instrument thinking tokens per response and reason about the distribution per route: median, p95, and the tail. The tail is where the money is. **Set effort per request class, not globally.** A one-word intent classifier and a root-cause analysis should not share a setting even when they share an endpoint. The classifier wants the floor — or thinking off entirely, if the provider allows it; the analysis wants depth. This is a per-call parameter precisely so it can vary with what you are asking. **Expect the sampling knobs to be unavailable.** On several current reasoning models, temperature, top-p and similar parameters are rejected outright rather than ignored. That surprises people migrating code from older models, and it means effort is not merely the primary depth control — it may be the only one exposed. **Re-tune when the model changes.** Effort levels are labels over a provider's internal policy, not a portable unit. "High" on one model family is not "high" on another, and a model upgrade can shift both the accuracy and the cost at every level. Treat effort as a setting to re-measure on your own eval after any model change, the same way you would re-measure a prompt. ## The compact answer A token budget asks *how many tokens may you think for*, decided by the caller in ignorance. An effort level asks *how hard should you try*, decided by the model with the request in front of it. The gain is adaptivity; the price is that you lose a per-call number to plan against and must budget over distributions with enforcement elsewhere.

  • If effort is advisory, how do you actually stop a runaway request from costing too much?
    Through mechanisms that are genuinely enforced rather than suggested: the response's output-token limit, which thinking counts against, and your own accounting around the call — per-request and per-session spend tracking with a hard stop. Effort shapes the typical case; enforcement handles the tail. Treating a soft dial as a spend control is the mistake that produces surprise bills.
  • You raise effort from low to high across an endpoint and quality does not improve. What does that tell you?
    That the bottleneck is not deliberation. Usually it means the task lacks sequential structure, or the model is missing information no amount of thinking supplies — absent context, an ambiguous specification, stale knowledge. The fix is grounding, better instructions or a cleaner task decomposition, not more compute. Confirm it on an eval set before paying for the higher level in production.
  • Why do some reasoning models reject temperature and top-p instead of ignoring them?
    Because the training that produced the thinking behaviour assumes a particular decoding regime, and altering it degrades reasoning quality in ways that are hard for a caller to detect. Rejecting with an error is the honest failure: it surfaces the incompatibility at integration time rather than silently returning worse answers. Practically, it means depth is controlled by effort, not by sampling.

saying these in an interview costs you the question

  • Effort is a hard cap on thinking tokens
  • Higher effort always gives a better answer
  • Effort switches you to a larger model
  • High effort works fine under a small output limit
  • "High" means the same across model families

context