Which sampling parameters have no effect on DeepSeek's deepseek-reasoner model?
answer
- Accepted is not the same as honoured
- Two failure modes, not one
- Silence is worse than an error here
- The decoding dials are fixed by the vendor
- Per-model parameter policy, not global config
basics
~20 sDeepSeek documents temperature, top_p, presence_penalty and frequency_penalty as unsupported on deepseek-reasoner: the request succeeds but the values do nothing. logprobs and top_logprobs are rejected with an error instead of being ignored. Verify against the docs for your model release.
solid answer
~40 sOn `deepseek-reasoner`, DeepSeek documents the usual decoding knobs — `temperature`, `top_p`, `presence_penalty`, `frequency_penalty` — as unsupported: the call still returns 200 and a normal completion, but setting them changes nothing. That silent acceptance is the trap, because a shared client that tunes `temperature` per environment appears to work and quietly has no effect. A second group behaves differently again: `logprobs` and `top_logprobs` are rejected with an error rather than ignored. The practical rule is to keep model-specific parameter sets rather than one global request builder, and to assert in tests that the reasoner path sends no dead knobs. Because DeepSeek has revised reasoner parameter support across model releases, treat the exact list as version-specific and check the docs for the release you call.
go deeper
Know that deepseek-reasoner ignores the usual sampling knobs such as temperature and top_p, so setting them in your request accomplishes nothing.
Distinguish the two behaviours — silently ignored versus rejected with an error — and explain why the silent case is the one that misleads teams into tuning a dead dial.
Demonstrate the defensive practice: per-model parameter policy, a test asserting the reasoner request omits dead knobs, and an empirical extremes check before trusting any compatible endpoint.
Own the platform rule that OpenAI-compatible does not mean OpenAI-equivalent, and design the abstraction so provider capability differences are declared and validated rather than discovered in production.
## Two different kinds of "unsupported" DeepSeek's reasoning model does not honour the full OpenAI-compatible parameter set, and it splits the gap into two behaviours that matter operationally: - **Accepted but inert.** `temperature`, `top_p`, `presence_penalty`, `frequency_penalty`. The request validates, you get a normal 200 and a normal completion, and the value had no influence on generation. - **Rejected.** `logprobs` and `top_logprobs` produce an error rather than being ignored. The first group is the dangerous one. An error teaches you immediately; a silently ignored parameter teaches you nothing, and teams have shipped elaborate temperature tuning against a model that never read it. The classic symptom is an A/B test between `temperature: 0` and `temperature: 1` that shows no behavioural difference at all — not a subtle one, none — because both requests were decoded identically. ## Why a reasoning model would fix decoding The general principle is that a reasoning model's quality depends on a decoding regime chosen during training and post-training. Handing that dial to callers invites configurations where the long thinking phase derails — a high-temperature chain of thought can wander for thousands of tokens before answering, and repetition penalties applied across a long trace interact badly with the deliberate restatement reasoning traces do. Pinning the sampling policy makes behaviour reproducible for the vendor and removes a whole class of caller-induced failure. Note what this does *not* mean: fixed sampling parameters do not make the model deterministic. Serving-side batching and floating-point non-determinism still make repeated identical requests differ. If you need reproducibility, capture outputs, do not assume them. ## The compatibility lesson generalised This is a concrete instance of the broader trap with OpenAI-compatible APIs: the request schema being accepted says nothing about the semantics being implemented. A gateway or SDK written against one vendor will happily serialise every field it knows, and a compatible-but-different backend may honour it, ignore it, or reject it. Assume nothing from a 200. The defensive shape is per-model parameter policy: - Keep a small map from model name to the set of parameters you are allowed to send, and build requests through it rather than spreading a global config into every call. - Fail loudly in your own code — a unit test that asserts the reasoner request body contains no `temperature` key is cheap and catches the regression the API will not. - Log the request body shape (not just the model name) when you are debugging quality differences between environments, so an inert knob is visible in the trace. ## When a parameter is inert, what do you tune instead? With decoding fixed, your remaining levers are the prompt, the model choice, and the output budget. If answers are too verbose or too terse, say so in the instructions rather than reaching for penalties. If you want variety across samples, request several completions or vary the prompt, rather than raising a temperature the server discards. If answers truncate, that is the output-budget lever, not a sampling one. ## Version sensitivity DeepSeek has changed what the reasoner accepts across model releases — capabilities that were absent in an early reasoning release have appeared in later ones. So the durable interview answer is the *shape* of the fact, held in this order: some sampling parameters are accepted and ignored rather than rejected; a smaller set errors; and the authoritative list lives in the API docs for the specific model version you are calling. A candidate who states the mechanism and then says "and I'd pin that against the current docs, because DeepSeek has moved it" is answering better than one who recites a list with false confidence. ## Testing for it The empirical check is simple and worth knowing: send the same prompt twice at the extremes of the parameter — once at the lowest value and once at the highest — with everything else fixed, several times each. If the distribution of outputs is indistinguishable, the parameter is not being applied. Do the same when you migrate a workload to any OpenAI-compatible endpoint; it takes minutes and prevents months of tuning a dial connected to nothing.
- How would you empirically prove a parameter is being ignored by an endpoint?Hold everything else fixed and sample the same prompt many times at the two extremes of the parameter, then compare the output distributions. A parameter that is honoured shows a visible spread difference between the extremes; an ignored one gives statistically indistinguishable sets. It is a few minutes of work and it settles the question without relying on documentation.
- If temperature has no effect, what levers remain to control output style and variety?The prompt, the model choice, and the output budget. Ask for the length and register you want in the instructions instead of reaching for penalties; get variety by sampling multiple completions or varying the prompt rather than by a discarded temperature; and fix truncation through the token budget, which is a separate mechanism from sampling.
- Does a fixed sampling policy make deepseek-reasoner deterministic?No. Even with decoding pinned by the vendor, serving-side batching and floating-point non-determinism mean identical requests can produce different text. Treat reproducibility as something you achieve by capturing and storing outputs, not something you infer from the absence of a temperature parameter.
saying these in an interview costs you the question
- Assumes a 200 response means the parameter was applied
- Tunes temperature on the reasoner and reports quality gains
- Expects an error for every unsupported parameter
- Thinks fixed decoding makes the model deterministic
- Sends one global parameter set to every model in the fleet