skip to content

Why can an LLM still vary at temperature 0 with a fixed seed?

level: seniorimportance: should knowfreq 44%

answer

  1. the seed pins draws, not numbers
  2. floating point addition is not associative
  3. your batch depends on other people's traffic
  4. aliases move under you silently
  5. assert properties, not exact strings

basics

~20 s

A seed only pins the sampler's random draws; it does not pin the numbers the model computes. Floating-point reductions are not associative, so results shift with how requests are batched on the server, and version, hardware and routing changes shift them further.

solid answer

~50 s

Temperature 0 removes randomness from the *choice*, and a seed removes randomness from the draw — but neither touches how the logits are produced. GPU kernels compute sums in whatever order the parallel decomposition dictates, and floating-point addition is not associative, so a different reduction order gives a slightly different logit. That order depends on the batch the server assembled, which depends on other people's concurrent traffic. When two candidate tokens are close, a difference in the last bits flips the argmax and the sequence diverges from there. Add moving model aliases, mixed GPU generations behind one endpoint, and expert-routing effects in mixture-of-experts models, and identical bytes stop being a reasonable expectation. Building a dialogue-regression suite on exact string equality therefore fails intermittently. Pin the model version, keep temperature 0 and a seed for what they do buy, and assert on properties — schema validity, required facts, a judge rubric — with repeated runs rather than one.

code

python · 6 lines
python
a = 1e16
b = -1e16
c = 1.0

print((a + b) + c)
print(a + (b + c))

go deeper

for a junior

Know that temperature 0 and a seed make output much more consistent but do not guarantee the identical response from a hosted service, and that tests should not compare exact strings.

for a middle

Explain the split: the seed controls sampling draws while the logits themselves can shift, because floating-point sums are order-dependent and the order follows how the server batched requests.

for a senior

Diagnose it in production — batch composition, unpinned version aliases, mixed hardware — and design the eval suite around it with property assertions and repeated runs instead of exact match.

for a principal

Own the tradeoff: batch-invariant serving buys reproducibility with throughput, so decide where in the estate reproducibility is worth paying for, and hold the line that reproducibility is a debuggability property, never evidence of correctness.

## Two different kinds of randomness There are two independent sources of variation in a generation, and people conflate them. The first is *sampling randomness*: given a probability distribution, which token gets drawn. Temperature 0 eliminates it by always taking the argmax; a seed eliminates it at higher temperatures by fixing the pseudo-random stream. Both are fully under your control and both work. The second is *numerical variation*: what the probability distribution is in the first place. Neither knob touches it, and it is where hosted-endpoint nondeterminism actually comes from. ## Why the numbers move Floating-point addition is not associative: (a + b) + c and a + (b + c) can differ in the final bits. A GPU computes a matrix multiplication or a softmax by splitting the work across many threads and combining partial results, and the *order* of that combination depends on how the kernel decomposed the problem — which depends on tensor shapes, and therefore on how many requests the server batched together and how long each one is. Batch composition is a property of live traffic. Your request runs alongside whatever else arrived in the same window, so the same prompt at 9am and at 3am can be processed with different shapes and produce logits differing in their last bits. Usually that changes nothing visible. But when the top two candidates are nearly tied, a last-bit difference flips which one is the argmax, and because generation is autoregressive, one flipped token sends the rest of the response down a different path. That is why the divergence is bursty rather than gradual: identical for many calls, then wholly different. This is the *batch-invariance* problem, and it is addressable: kernels can be written so their reduction order does not depend on batch composition, and open-weight serving stacks have shipped batch-invariant modes at some throughput cost. It is a deliberate tradeoff — you buy reproducibility with performance — and it is only available where you control the serving stack. ## The other sources Batching is the subtle one. Several coarser sources matter just as much in practice. **Moving model pointers.** If you call an alias rather than a pinned version, the weights under you can change without any change on your side. This is the single most common cause of a suite that was stable for weeks failing overnight. **Heterogeneous fleets.** A hosted endpoint may run several GPU generations with different kernels and numerics. Which one serves you is not something you choose. **Mixture-of-experts routing.** In MoE architectures, which experts process a token is decided by a router. Where routing interacts with batch-level capacity constraints, the same token can be handled differently depending on what else is in the batch. Modern load-balancing designs reduce this, but it is another way batch composition reaches the output. **Ties and prompt handling.** Exact ties in the argmax are broken by implementation detail, and any server-side normalisation of your prompt changes the input before the model sees it. ## What to do instead Do not build a regression suite on exact string equality against a hosted endpoint. It will pass locally, pass in CI for a week, and then fail on a day nobody changed anything, which trains the team to ignore it. Do take the cheap wins. Pin an explicit model version rather than an alias. Set temperature 0 and pass a seed where the endpoint supports it: they remove real variance even though they do not remove all of it. Keep prompts byte-stable, since a whitespace change is a genuine input change. Then assert on the right thing. Validate structure — does it parse, does it satisfy the schema, are the required fields present. Validate content properties — does the answer contain the entity, is the classification label correct, is the tool call the expected one with the expected arguments. Where output is free prose, score it with a rubric rather than comparing it. And run each case several times rather than once, because a single sample of a stochastic system tells you very little; treat a case as passing only if it passes consistently, and track that consistency as the metric. Finally, be clear internally that determinism and correctness are different goals. A perfectly reproducible pipeline that reproduces the same wrong answer has bought you debuggability, not quality. Debuggability is worth real money — it is what makes an incident reproducible — but it is not a substitute for evaluation.

  • If you control the serving stack yourself, can you get bit-exact reproducibility?
    Much closer, yes. Pin the weights, the engine version and the hardware; fix temperature and seed; and enable a batch-invariant kernel mode if the stack offers one, so reduction order no longer depends on batch composition. The cost is throughput, because batch-invariant kernels give up some scheduling freedom. It is a reasonable trade for a debugging or forensic environment, rarely for production serving.
  • How should an eval suite be structured given that exact-match assertions are unreliable?
    Assert on properties rather than strings: schema validity, presence of required facts, the expected label or tool call, absence of banned content. Run each case several times and require consistent passing rather than a single green run, so a flaky case is visible as flakiness rather than as random failure. Reserve exact match for genuinely closed outputs like a single classification token.
  • Your suite was stable for a month and started failing overnight with no code change. Where do you look first?
    At the model pointer. Calling an alias rather than a pinned version means the weights can change under you, and that is the most common cause by a wide margin. Check whether the endpoint or default version moved, then whether prompts changed upstream — a template edit or an injected timestamp counts as a real input change — before investigating numerics.

saying these in an interview costs you the question

  • Believes temperature 0 guarantees identical output
  • Thinks a seed pins the model's computation, not just sampling
  • Blames the model for flakiness caused by a moving version alias
  • Builds regression tests on exact string equality
  • Says reproducible output means correct output

context