Why doesn't temperature 0 make an LLM eval suite reproducible in CI?
answer
- greedy decoding is not the whole story
- the sampler is not the only variance
- floating-point order, batch composition
- aliases move without a commit
- measure the residue, then design around it
basics
~20 sTemperature 0 only removes sampling randomness. Serving-side variance (batched floating-point arithmetic, expert routing), silently repointed model aliases, and unpinned fixtures all still move scores, so a suite needs a full pin set and repeated runs rather than one greedy pass.
solid answer
~50 sTemperature 0 makes decoding greedy — it always takes the highest-probability token — but that is only one of several sources of variation. The provider's own serving stack is not bitwise deterministic: results depend on batch composition and kernel scheduling, floating-point addition is not associative, and mixture-of-experts routing can shift with the batch, so the top-two tokens can swap on a near-tie and the outputs diverge from there. A seed, where offered, is best-effort and does not cover that. The larger practical hazard is version drift: pointing at a floating model alias means the provider can move you to a new snapshot overnight and your suite drops several points with no commit to blame. So you pin the exact dated model snapshot, decoding parameters, prompt template hash, retrieval index snapshot, fixtures, and the judge model and rubric version — then accept the residual noise and design the gate's decision rule around it.
code
json · 11 lines{
"model_snapshot": "provider-model-2026-03-11",
"temperature": 0,
"top_p": 1,
"max_output_tokens": 512,
"prompt_template_sha": "9f2c1b7",
"retrieval_index_snapshot": "2026-05-02",
"frozen_clock": "2026-05-02T00:00:00Z",
"judge_model_snapshot": "provider-judge-2026-01-20",
"judge_rubric_version": "v4"
}go deeper
Know that setting temperature 0 removes sampling randomness but not all variation, and that eval configuration should name an exact dated model version rather than a moving alias.
Explain why greedy decoding still produces different text — batch-dependent floating-point arithmetic flipping a near-tied argmax, then autoregressive divergence — and list the pins a run should record, judge model included.
Demonstrate the diagnosis: an unexplained score move sends you to the recorded pins first, not to prompt bisection. Show how you measure the residual noise band and feed it into the gate's margin and repeat policy.
Own the framing that an eval is an instrument with an error bar, and set the organisation's policy on model-version upgrades: pinned snapshots, upgrades as reviewed pull requests with a measured delta, never a silent alias move.
## What temperature 0 actually buys you Temperature scales the logits before sampling. At 0 the sampler degenerates to argmax: always take the highest-probability token. That removes *sampling* randomness, which is the largest and most obvious source of run-to-run variation, and it is the right default for an eval run. But it does not make the pipeline a pure function, and treating it as if it did is the most common mistake in eval engineering. ## Sources of variation that survive greedy decoding **Serving-stack nondeterminism.** Inference runs on accelerators in batches assembled from whatever requests arrive together. Floating-point addition is not associative, so reduction order changes the last bits of a logit; kernel selection and tensor parallel splits vary with batch shape; mixture-of-experts architectures route tokens to experts in a way that can depend on the composition of the batch. None of this matters until two candidate tokens are nearly tied — then a bit-level difference flips the argmax, and because generation is autoregressive, one flipped token sends the rest of the completion down a different path. This is why identical requests at temperature 0 occasionally return materially different text. **Seeds are best-effort.** Where a provider exposes a seed, it constrains the sampler, not the arithmetic underneath it. Vendors typically document it as best-effort and pair it with a fingerprint of the backend configuration precisely because they cannot promise reproducibility across serving changes. A seed is worth setting; it is not a determinism guarantee. **Model version drift — the expensive one.** If your configuration names a floating alias rather than a dated snapshot, the provider decides when you change models. The classic incident: a suite sits green for weeks, then overnight drops six points with no commit in the diff. Hours go into bisecting prompt changes that are not there before someone checks the response metadata and finds the alias now resolves to a new snapshot. Pinning a dated snapshot converts that from a mystery into a scheduled, reviewable upgrade: you bump the pin in a pull request, the eval runs, and you see the delta before it reaches production. **Everything around the model.** Retrieval indexes get re-embedded and re-ranked; tool endpoints return live data; any case whose expected answer depends on the current date rots; a judge model has all of the above problems plus its own version drift. A prompt containing a timestamp or a randomly ordered document list changes the input on every run. ## The pin set A disciplined suite records, and asserts on, a fixed set of pins for every run: - exact **dated model snapshot** (never a floating alias), and the backend fingerprint if the provider returns one; - **decoding parameters**: temperature, top-p, max output tokens, stop sequences, and any reasoning-effort setting, since effort level changes both output and cost; - **prompt template version or content hash**, so a whitespace edit is visible in the run record; - **fixtures**: retrieval index snapshot, recorded tool responses, frozen conversation histories, a frozen clock for anything date-dependent; - **judge model snapshot and rubric version**, treated with the same rigour as the system under test. Writing these into the run's output turns an unexplained score move into a diff: you can see which pin changed. A run whose pins differ from the baseline's pins is not comparable, and a good harness says so loudly rather than reporting a delta as if it meant something. ## Living with the residue After pinning, some variance remains — irreducibly, because it lives in someone else's serving infrastructure. The engineering response is not to chase bitwise reproducibility but to *measure* the residue and design around it. Run the unchanged suite several times and record the spread of the aggregate score; that spread is your noise band. Anything smaller than the band is not evidence of a regression, and the gate's decision rule has to reflect that — a margin sized above the band, and repeated runs with a majority rule rather than a single pass or fail. A useful mental frame: the eval is a measurement instrument with a known error bar, not a unit test. Pinning shrinks the error bar; it never removes it. Teams that understand this ship confident gates. Teams that assume temperature 0 means determinism spend their time re-running builds and eventually turn the gate off.
- Your suite drops six points overnight with no commits to the prompt. What do you check first?The pins recorded by the run, starting with the model snapshot. If the configuration named a floating alias, the provider may have repointed it to a new model — the response metadata or backend fingerprint usually shows it. Next check the retrieval index and any fixture that refreshes on a schedule, then the judge model. Bisecting prompt history first wastes hours when the change was not in your repository.
- Should the judge model be pinned as strictly as the model under test?Yes, and it is frequently forgotten. A judge on a floating alias can shift its own grading behaviour and move every score in the suite while the system under test is unchanged, which reads exactly like a product regression. Pin the judge snapshot, version the rubric, run it at temperature 0, and treat any judge upgrade as its own change that requires re-baselining.
- If some nondeterminism is irreducible, why bother pinning at all?Pinning shrinks the error bar and, more importantly, makes score moves attributable. With everything pinned, a change in the score has exactly one candidate cause — the diff under review. Without pins, every red build starts an investigation into whether anything even changed. The residual noise is then handled by measuring it and setting the gate's margin and repeat rule above it.
saying these in an interview costs you the question
- Temperature 0 makes the model fully deterministic
- A seed parameter guarantees identical output
- Pointing evals at a floating model alias
- Forgetting to pin the judge model and rubric
- Re-running until green instead of measuring the noise band