How does routing reasoning effort per request class differ from model routing?
answer
- Same model, different thinking budget
- Hidden tokens, billed and unbounded
- One prompt, one eval baseline
- Class it before the call, not during
- Effort cannot add missing capability
basics
~20 sEffort routing keeps one model and varies how much hidden reasoning it may spend per request class; model routing swaps the model itself. Effort routing preserves the prompt, output format and eval baseline, so it is the cheaper dial to move first.
solid answer
~50 sFrontier providers expose an effort dial alongside the model choice — OpenAI's `reasoning_effort`, Google's `thinking_budget`, Anthropic's discrete effort levels (as of mid-2026). Turning it up lets the model spend far more hidden reasoning tokens before answering; those tokens are billed and can grow several-fold, which is why effort is now a first-class cost and latency dial rather than a quality knob. The operational difference is that effort routing changes one parameter on a fixed model: the prompt, the output schema, the tool definitions and the eval baseline all stay put, so the tiers cannot drift apart behaviourally. Model routing changes all of that at once. Route effort by *request class* using attributes you know before the call — amended returns get high effort, routine ones get low — rather than by per-request model judgment, so the decision is deterministic, auditable and free.
go deeper
Know that modern models take an effort or thinking-budget setting, that higher settings spend more hidden reasoning tokens, and that those tokens cost money and time.
Explain why effort routing keeps the prompt, schema and eval baseline fixed while model routing changes all of them, and why that makes effort the first dial to move.
Show that you route effort by pre-known request class rather than per request, meter thinking tokens as their own dashboard line, and set each class's effort against that class's latency budget.
Own the composition of the two dials across a portfolio of features: which classes deserve deliberation spend at all, how effort settings are re-validated on every model upgrade, and who owns that cadence.
## Two different dials Cost and latency optimization for an LLM feature has two independent controls. The first is *which model* answers. The second is *how hard that model thinks* before answering. Since extended-thinking models became standard, the second dial moves cost and latency at least as much as the first, and it is usually the one to reach for first. As of mid-2026 every frontier provider exposes it, under different names and shapes: OpenAI has `reasoning_effort`, Google has `thinking_budget`, and Anthropic exposes discrete effort levels. The names differ; the semantics are the same at the level that matters here — an upper bound or a preference for how much hidden reasoning the model may spend before it emits a visible answer. ## What the effort dial actually changes Raising effort does not lengthen the answer. It lengthens the *reasoning* that precedes it. Those tokens are billed like output tokens and are invisible to the user, and they can balloon several-fold between the lowest and highest settings with no natural ceiling — the model stops when it is satisfied, not at a length the prompt implies. Two consequences follow directly: - **Cost per request becomes decoupled from prompt and answer size.** Two requests with identical prompts and identical answers can differ several-fold in billed tokens because one thought longer. If your cost model is "tokens in plus tokens out", it is wrong; thinking tokens need their own line. - **Latency moves with effort, hard.** A high-effort request can spend tens of seconds reasoning before a single visible token appears. That is a UX problem as much as a cost one, and it is why effort settings and interactive-path SLOs have to be designed together. ## Why effort routing is operationally cheaper than model routing When you route between two *models*, you take on everything that differs between them: prompt sensitivity, instruction-following quirks, tool-calling reliability, output formatting, refusal behaviour, tokenizer and context limits, and two separate upgrade cadences. Your eval suite has to cover both, and any prompt fix has to be applied and re-validated twice. The tiers drift apart over time because they are maintained separately. Routing *effort* on one model changes exactly one parameter. The prompt is identical, the schema is identical, the tools are identical, and the eval baseline is one baseline with an extra dimension swept across it. The quality difference between effort levels is monotone in a way that a model swap is not: a stronger model is not uniformly better on every request class, but more thinking on the same model rarely makes a request class worse in aggregate. That makes effort the safer first dial and the easier one to justify in review. The flip side is range. Effort cannot rescue a model that lacks the underlying capability — thinking longer does not add knowledge it never had, and past a point the marginal accuracy per extra thinking token collapses. When the gap is capability rather than deliberation, only a model change closes it. In practice the two dials compose: a cheap model at low effort for the routine class, the strong model at high effort for the hard class, and the middle handled by moving effort before moving model. ## Route by class, not per request The unit of effort routing should be a *request class* defined by attributes you already hold before the call: document type, whether a filing is amended, whether the customer is in a regulated segment, whether the input arrived from a trusted internal pipeline or an end user, whether an earlier attempt failed. Classifying by rule is free, deterministic, auditable, and testable — you can assert in CI that amended returns get high effort. Asking a model to decide how hard to think adds a call, adds nondeterminism, and adds a new failure mode where the classifier itself is wrong. Where the class boundary is genuinely fuzzy, a small learned classifier over request features can set the effort level, but it should be treated the same way as any other router: labeled data, a swept threshold, and both error directions measured. ## Where effort routing goes wrong **Over-thinking trivial inputs.** High effort on a simple request wastes tokens and time and sometimes hurts: the model talks itself out of a correct simple answer. Default to the lowest effort that clears your quality bar, and escalate the class, not the request. **Unmetered thinking tokens.** If your dashboards do not separate thinking tokens from visible output tokens, an effort change looks like a mysterious cost regression with no prompt change to blame. **Tail latency on interactive paths.** Raising effort on a user-facing class can blow a p99 SLO even while the average is fine. Decide the effort level per class *with* the latency budget of that class in hand. **Assuming portability.** Effort settings are provider-specific in name and in granularity, and their behaviour changes across model versions. A level that was right last quarter needs re-sweeping after an upgrade; treat it as a tuned parameter, not a constant in a config file nobody revisits.
- Why can raising effort blow a latency SLO even when average latency looks fine?Because thinking happens before any visible token is produced, and its length is decided by the model rather than capped by the prompt. A high-effort class can sit tens of seconds in reasoning while the mean across all classes stays low, since most traffic is on a lower level. Look at latency per effort class and at the tail, not the blended average, and set each class's effort against its own budget.
- When would you move the model rather than the effort level?When the gap is capability, not deliberation. If the small model gets the shape of the problem wrong — misses a domain rule, cannot follow the schema, lacks the knowledge — more thinking amplifies a wrong approach instead of fixing it. The tell is a flat or falling accuracy curve as effort rises on that request class. When accuracy is still climbing with effort, keep turning the cheaper dial.
- How do you keep effort settings honest across a model upgrade?Treat them as tuned parameters with an owner. After any model version change, re-sweep effort levels for each request class against the eval suite and re-plot accuracy and cost, because both the quality at each level and the token consumption at each level shift. Pin the model version in the eval run so a config that was validated is traceable to what it was validated against.
saying these in an interview costs you the question
- Thinks raising effort just makes the visible answer longer
- Assumes thinking tokens are free or not billed
- Sets one global effort level for all traffic
- Asks the model per request how hard it should think
- Believes more effort can compensate for a missing capability