How should a product adapt when its endpoint rejects temperature and top_p?
answer
- the knob is a property of the target
- never silently drop a caller's parameter
- constrained output replaces low temperature
- variety must move into the prompt
- two lanes cost more than one knob
basics
~20 sControl of output variability moves out of the sampler and into the system around it: prompt specificity, constrained output, validators, and generating several candidates and selecting. Keep an open-weight lane only where tuned sampling is a genuine requirement, and never silently drop rejected parameters.
solid answer
~50 sAs of mid-2026 several frontier reasoning endpoints no longer accept sampling parameters at all — Anthropic's current models return an error for temperature, top_p and top_k, and other vendors lock sampling on their reasoning models, exposing a coarse effort control instead. Treat that as a design constraint rather than a bug. First, make the client layer explicit: send sampling parameters only to targets that accept them, and surface a rejection as a configuration error rather than swallowing it, because silently dropping a parameter someone tuned is worse than failing. Second, move the levers you lost. Determinism becomes property-based assertions and pinned versions; conservatism becomes schema-constrained output plus validation; variety becomes prompt-level variation or generating several candidates and selecting. Third, decide honestly whether any workload still needs a real entropy dial — bulk creative generation often does — and route that lane to open-weight serving, accepting that you now run two lanes.
go deeper
Know that sampling parameters are not universal — some hosted endpoints reject temperature and top_p outright — and that your code should not assume it can always send them.
Explain what each rejected knob was buying and what replaces it: constrained output and validation for conservatism, pinned versions and property assertions for reproducibility, prompt variation for variety.
Design the client layer so unsupported parameters fail at configuration time rather than being dropped, and be able to migrate a tuned route onto a locked endpoint by naming the property it actually needed and enforcing it directly.
Own the portfolio call: whether any workload genuinely justifies running a second, sampling-controllable serving lane, priced against the capacity, evaluation and on-call cost of operating it — and insist the decision is made on measured evals, not on attachment to a knob.
## What changed For most of the API era, temperature and a truncation cutoff were assumed present on every text endpoint, and abstraction layers passed them through unconditionally. That assumption no longer holds. As of mid-2026, frontier vendors have converged on models that reason internally before answering, and several of them lock the sampler. Anthropic's current models reject temperature, top_p and top_k with an error rather than ignoring them; OpenAI's reasoning models likewise do not honour sampling parameters and expose a reasoning-effort setting; Google's latest models replaced a numeric thinking budget with a coarse level. The knobs remain fully available in open-weight serving stacks you run yourself. The practical statement is: sampling controls are a property of the target, not of the field. The rationale is that these models are tuned as a whole system — internal reasoning plus final answer — and their behaviour was validated at fixed sampling settings. Letting callers perturb the distribution can degrade the reasoning path in ways the vendor cannot support. ## The client-layer decision The first design choice is how your abstraction handles a parameter a target does not accept, and there are three options. *Pass through and let it fail* keeps the truth visible: a request configured with a temperature against a locked model errors, loudly, at the boundary. Noisy, but honest. *Silently drop* is the tempting one and the one to avoid. A team tunes temperature 0.2 for an extraction path, the router quietly discards it, and the behaviour they validated is not the behaviour in production — with no signal anywhere. *Validate at configuration time* is the mature answer: model capabilities are declared in your registry, and a route that sets a sampling parameter against a model that does not support it fails at startup or in CI, not at request time. Requests then only ever carry parameters the target accepts. Whichever you choose, the anti-pattern is uniform: never let a caller believe a knob is in effect when it is not. ## Replacing what the knobs bought The knobs bought three distinct things, and each has a different replacement. **Conservatism** — low temperature for extraction, classification and tool-argument filling — is replaced by structure and verification. A schema-constrained output path plus a validator gives you a stronger guarantee than temperature 0 ever did, because it enforces the property you actually cared about instead of approximating it by narrowing the distribution. Tighten the prompt so the space of acceptable answers is small, and reject-and-retry on validation failure. **Reproducibility** never really came from temperature 0 anyway on hosted infrastructure. Pin model versions, keep prompts byte-stable, assert on properties rather than exact strings, and run cases repeatedly. **Variety** is the one that genuinely hurts. If you were generating flavour text at temperature 1.1 to keep outputs distinct, a locked endpoint gives you a fixed level of diversity. Compensating means varying the *input* rather than the sampler: rotate constraints, seed each call with different facts, or generate a batch of items in one call and ask explicitly for mutually distinct results, which uses the model's own attention over its list rather than the sampler's entropy. When those are not enough, you have a routing decision. ## The portfolio question This is where the judgement sits. Running an open-weight lane alongside hosted endpoints is not free: capacity planning, quantization choices, engine upgrades, your own evals, your own on-call. The honest framing is that you pay that cost when a workload has a property the hosted lane cannot provide — a tuned entropy dial for bulk creative generation, hard data-residency limits, unit economics at very high volume, or a genuine need for reproducible numerics. What you should not do is stand up a second lane out of attachment to a knob. For most enterprise workloads — extraction, summarisation, classification, tool use — locked sampling costs nothing measurable, because those tasks wanted near-deterministic behaviour anyway and are better served by constrained output. Measure the gap on your own evals before deciding. ## Second-order effects Three worth naming. Vendor comparison gets harder: you can no longer equalise sampling across two providers, so A/B results reflect each vendor's chosen defaults — which is arguably the more honest comparison, since that is what production sees. Old prompt presets carrying tuned values become dead configuration and should be deleted rather than left inert. And the abstraction layer's job shifts from smoothing over provider differences to declaring them, because a capability registry that lies is worse than no registry at all.
- Why is silently dropping a rejected sampling parameter worse than returning an error?Because it breaks the link between what a team validated and what runs. Someone tunes a value, measures a result, ships it — and the router discards the parameter with no signal, so production behaviour differs from the tested behaviour and nothing in the logs says why. An error at the boundary, or better a configuration-time check, keeps the mismatch visible when it is cheap to fix.
- A team insists they cannot migrate to a locked endpoint because they need temperature 0. What do you tell them?That temperature 0 was a proxy for the guarantee they actually want, and on hosted infrastructure it never delivered it. Ask what property they are protecting — parseable structure, a stable label, a specific field — and enforce that directly with constrained output and a validator. Then measure the migration on their own eval set. Usually the gap is nil; occasionally it is real, and then routing is the answer.
- How does locked sampling change how you compare two vendors?You can no longer equalise the sampler, so any comparison reflects each vendor's own defaults rather than a controlled setting. That removes a tuning lever from the experiment but arguably improves validity, since production will run those defaults too. Hold the prompt, the eval set and the grading rubric fixed instead, and report cost and latency alongside quality rather than treating quality alone as the comparison.
saying these in an interview costs you the question
- Assumes every text endpoint accepts a temperature parameter
- Silently strips unsupported parameters in the client layer
- Says locked sampling makes an endpoint unusable for production
- Replaces a lost temperature knob by retrying until output differs
- Stands up self-hosted serving without measuring whether the knob mattered