In OpenAI Chat Completions, what happens if you set both temperature and top_p?
answer
- Same sampling step, two knobs
- Legal but discouraged together
- One at a time, per the docs
- Zero is not a determinism promise
- Reasoning models refuse the knobs
basics
~20 sNothing errors — both are applied to the same sampling step, and OpenAI's own guidance is to change one or the other, not both, because their combined effect is hard to predict. Defaults are temperature 1 and top_p 1.
solid answer
~40 sThey are two knobs on the same sampling step and the API accepts both, so there is no error — the request just becomes hard to reason about. `temperature` (0 to 2, default 1) rescales the probability distribution before sampling; `top_p` (0 to 1, default 1) restricts sampling to the smallest set of tokens whose probabilities sum to `top_p`. OpenAI documents that you should generally alter one **or** the other. In practice teams pick temperature for creativity control and leave `top_p` at 1. Two API-level facts matter beyond that: temperature 0 is near-greedy but **not** a determinism guarantee — for best-effort reproducibility you also pass `seed` and compare the returned `system_fingerprint` — and OpenAI's o-series reasoning models do not accept non-default sampling values, rejecting the request rather than quietly ignoring it.
go deeper
Know that temperature controls randomness, that its default is 1, and that lower values give more focused, repeatable-sounding answers.
Explain the ranges and defaults of both parameters, what each does to the candidate token set, and why OpenAI advises changing only one.
Demonstrate that temperature 0 is not a determinism guarantee, reach for seed and system_fingerprint, and design tests that do not depend on exact strings.
Own sampling as configuration policy across a model fleet: per-model overrides, the reasoning models that reject the knobs, and how output-variance settings interact with evaluation and cost.
## Two knobs, one sampler After the model produces a probability for every possible next token, the API applies your sampling parameters and picks one. `temperature` and `top_p` both act at that moment, in the same pipeline, which is exactly why tuning both at once is discouraged. ## What temperature does at the API level `temperature` accepts 0 to 2 and defaults to 1. Lower values sharpen the distribution so the highest-probability tokens dominate, producing focused, repetitive, more predictable text. Higher values flatten it, giving rarer tokens a real chance — more varied, and more likely to go off the rails above roughly 1.2. Values near 0 are what you want for extraction, classification and code-shaped output. ## What top_p does at the API level `top_p` accepts 0 to 1 and defaults to 1, meaning no restriction. It keeps only the smallest set of candidate tokens whose probabilities add up to `top_p` and samples from that set. At 0.1, only the tokens making up the top ten percent of probability mass are eligible. Unlike temperature it does not reshape the distribution, it truncates it. ## Setting both The API will happily take `temperature: 0.2` together with `top_p: 0.5`. Nothing warns you. The problem is interpretive: the truncation and the rescaling interact, so you can no longer attribute a behaviour change to either knob, and a value that seems conservative on one axis can be swamped by the other. OpenAI's parameter documentation says explicitly to alter this or temperature but not both. The practical convention is: leave `top_p` at its default of 1 and move `temperature`, or, if you prefer nucleus control, pin `temperature` at 1 and move `top_p`. Write the choice down in your config so nobody later helpfully tunes the other one. ## Determinism is not what temperature 0 gives you A very common wrong answer is that temperature 0 makes the API deterministic. It makes sampling effectively greedy, which removes one source of variation, but identical requests can still return different text: floating-point non-determinism, batching, load balancing across hardware, and silent model updates all contribute. The API offers `seed` for best-effort reproducibility and returns `system_fingerprint` on the response, which identifies the backend configuration; if the fingerprint changes between calls, OpenAI is telling you the setup changed and reproducibility is not promised. Treat determinism as best-effort, and assert on structure or a grader in tests, never on exact strings. ## Models that reject these parameters This is where cross-provider habits get people. OpenAI's o-series reasoning models do not support arbitrary sampling values; sending a non-default `temperature` or `top_p` produces a 400 rather than being ignored. So code that blindly sets `temperature: 0` for every model breaks the moment someone points it at a reasoning model. Make the sampling parameters part of your per-model config, not a global constant, and be ready for the request to be rejected rather than silently degraded. ## Related knobs you may be asked to distinguish `frequency_penalty` and `presence_penalty` (both roughly -2 to 2, default 0) discourage repetition; they are not substitutes for temperature. `n` asks for more than one completion in a single response and multiplies your output-token cost. `stop` gives a small list of strings that end generation, and when one fires the response comes back with `finish_reason` `"stop"`. ## What to say in an interview Name the ranges and defaults, say that setting both is legal but discouraged and cite the one-or-the-other guidance, refuse the determinism trap, and mention that reasoning models reject the parameters outright. That combination is the answer that sounds like someone who has run production traffic rather than read a quickstart.
- You need repeatable output for a regression test. What do you actually set?Set `temperature` to 0 and pass a fixed `seed`, then check the `system_fingerprint` on the response — if it changes between runs, the backend configuration changed and identical output is not promised. Because reproducibility is best-effort, the test should assert on structure, schema validity or a grader's verdict rather than on an exact string match.
- What are the valid ranges, and what does temperature 2 typically produce?`temperature` runs 0 to 2 with default 1; `top_p` runs 0 to 1 with default 1. At 2 the distribution is flattened so far that low-probability tokens are picked often — output drifts into incoherence, switches language mid-sentence, or breaks required formats. Anything above roughly 1.2 is rarely useful outside deliberate brainstorming.
- Your service supports several OpenAI models behind one interface. How do you handle sampling parameters safely?Keep sampling settings in per-model configuration rather than a global default, and omit parameters a given model does not accept — the o-series reasoning models reject non-default `temperature` and `top_p` with a 400 instead of ignoring them. Fail loudly in tests when a model gains or loses support, since a silently dropped parameter changes output quality without any error.
saying these in an interview costs you the question
- Says temperature 0 makes the API deterministic
- Thinks setting both temperature and top_p returns an error
- Claims temperature ranges 0 to 1 in this API
- Believes top_p rescales probabilities the way temperature does
- Assumes every OpenAI model accepts temperature