skip to content

Can you reproduce a sampled self-consistency run in CI by fixing the seed?

level: middleimportance: should knowfreq 34%

answer

  1. seeds are best-effort, not a contract
  2. your call shares a batch with strangers
  3. floating-point reductions are order-sensitive
  4. assert on aggregates, not exact chains

basics

~20 s

No. A seed pins the sampling draw only if everything else is identical, and on hosted inference it is not: batch composition, kernel and hardware changes, and silent model updates all shift the output. Assert on aggregate metrics with tolerance, not on exact chains.

solid answer

~50 s

Seeds on hosted LLM APIs are best-effort, not a contract. Two things break them. First, the server batches your request with other users' traffic, and floating-point reduction order inside the kernels depends on batch shape — so the same tokens can produce marginally different logits, which occasionally flips a sampled token and then diverges the whole chain. Second, the serving stack moves underneath you: kernel versions, hardware, routing and model snapshots change without your code changing. I have seen a "same seed, same output" regression test pass locally for weeks and then fail nightly for exactly this reason. The fix is not to chase determinism; it is to stop asserting on it. Run the eval over fixed inputs, aggregate across repeated runs, and assert that the metric sits inside a tolerance band. Reserve exact-output assertions for recorded fixtures where the model is not in the loop at all.

go deeper

for a junior

Know that repeated calls to a hosted model can differ even with the same input and seed, and that tests should therefore not assert on an exact generated string.

for a middle

Explain why: your request is batched with other traffic, kernel reduction order depends on batch shape, floating-point addition is not associative, and a flipped token diverges the whole chain.

for a senior

Show how you would run evals under this reality — repeated runs, tolerance bands sized from observed variance, pinned model snapshots, and exact assertions confined to the deterministic prompt-building and parsing code.

for a principal

Decide whether reproducibility is worth buying at all: batch-invariant serving on owned infrastructure trades throughput for bit-exactness, and for most teams the right call is to design the eval and rollout process to tolerate variance instead.

## What a seed actually buys A seed initializes the pseudo-random generator that draws tokens from the model's output distribution. Given identical logits at every step, identical sampling parameters and identical PRNG state, the same seed reproduces the same chain. That chain of "identicals" is the catch: on a hosted API you control only the last two, and self-consistency deliberately runs at non-zero temperature, so any perturbation in the logits has a live path to a different token and therefore a completely different reasoning route. ## Why hosted inference is not bitwise reproducible **Batching non-determinism.** Serving stacks batch concurrent requests together for throughput. The batch's shape and composition depend on who else is calling at that instant — a thing you cannot control or even observe. Many GPU kernels are not batch-invariant: the reduction order in matrix multiplications and attention changes with batch shape, and floating-point addition is not associative, so the same inputs can yield logits that differ in the last bits. At temperature 0 that rarely changes the argmax but occasionally does, near a tie; at the non-zero temperatures self-consistency requires, it changes draws more readily. One flipped token early in a chain and the rest of the trace is unrecognizable. **Moving infrastructure.** Kernel and driver upgrades, different GPU models in a heterogeneous fleet, changes in tensor- or expert-parallel routing, speculative decoding acceptance patterns, and quantization or compilation changes all alter numerics. None of these are announced as behaviour changes because, from the provider's perspective, they are not. **Model snapshots.** An alias that points at "the latest" version can move under you. Pinning a dated snapshot removes this specific source but not the others, and snapshots are eventually retired. The consequence is a familiar failure: a CI test that recorded one exact chain, passed for weeks, and then failed intermittently — not because the prompt or the code regressed, but because a batch happened to be composed differently. Teams typically waste a day hunting a code change that does not exist. ## When determinism is achievable It is achievable in principle, and only on infrastructure you own. Serving with batch-invariant kernels — implementations whose reduction order does not depend on batch shape — plus a pinned model, pinned kernels, fixed hardware and controlled batching can give bit-identical results, at a measurable throughput cost. That work exists and is a legitimate answer to "could you", but it is not a thing you obtain by passing a seed to a public API, and it is rarely worth buying for a test suite. ## What to do instead **Assert on aggregates with tolerance.** Run the eval set, compute the metric you actually care about — vote accuracy, agreement rate, cost per item — and assert it is within a band. The band should be derived from observed run-to-run variance on a stable build, not guessed. **Repeat and report a distribution.** A single run of a stochastic system is a sample, not a measurement. Repeat the eval, report a mean and interval, and gate on the interval. This also protects you from the opposite error: celebrating a two-point improvement that is pure noise. **Size the eval set for the noise.** If the metric moves by three points between identical runs on 50 items, no threshold on 50 items means anything. Either enlarge the set or widen the band honestly. **Separate the deterministic parts and test them exactly.** Prompt construction, answer normalization, the parsing that extracts a final answer from a chain, retry and deadline logic — all of these are ordinary code and should have ordinary exact-assertion unit tests against fixed recorded chains. Keep the model out of those tests entirely; that is where reproducibility genuinely lives. **Pin what you can and record what you cannot.** Pin a dated model snapshot, pin sampling parameters, and record the model version, parameters and observed variance alongside every eval result, so a future failure can be attributed rather than guessed at. ## The special case of temperature 0 People reach for temperature 0 to "make the test deterministic". It does not: it removes the sampling stochasticity but leaves the numerical non-determinism intact, so near-tie tokens still flip. And in a self-consistency context it is self-defeating — the technique needs varied paths, so a test that runs at temperature 0 is no longer testing the system you ship. Test the sampled configuration and accept the variance in your assertions.

  • What would it actually take to make sampled inference reproducible?
    Control of the serving stack: batch-invariant kernels whose reduction order does not depend on batch shape, plus pinned model weights, pinned kernel and driver versions, fixed hardware and controlled batch composition. That is achievable when self-hosting, at some throughput cost, and effectively unavailable behind a public API where a seed parameter is documented as best-effort.
  • Does setting temperature to 0 at least make the run stable enough to assert on?
    No on both counts. It removes sampling randomness but not numerical non-determinism, so tokens near a tie can still flip and diverge the chain. And it destroys the path diversity self-consistency depends on, so the test would exercise a configuration you do not ship. Test the sampled configuration and widen the assertion instead.
  • How should the CI suite be structured around a stochastic sampling step?
    Split it. Prompt building, answer normalization, parsing and retry logic are deterministic code — unit-test them exactly against recorded chains, with no model call. Everything that involves the model becomes a metric evaluated over a fixed input set, repeated enough times to estimate variance, and gated on a tolerance band derived from that variance rather than on any single exact output.

saying these in an interview costs you the question

  • Believing a seed parameter guarantees identical hosted-API output
  • Blaming a prompt regression for what is batching non-determinism
  • Setting temperature to 0 to make self-consistency tests reproducible
  • Asserting on one exact reasoning chain in CI
  • Reading run-to-run metric noise as a real quality change

context