In a finding against a hosted chat endpoint you write 'set temperature to 0' as the reproduction instruction. Why is that not the same as recording a seed, and what should the finding say about determinism instead?
answer
- seed pins the draw; temperature shapes the distribution
- request parameter, not a determinism contract
- batching, routing, served build
- report attempts and successes, not 'deterministic'
- effective settings on the wire, not config file
basics
~20 sTemperature 0 records what you sent, not a guarantee you get the same tokens back. A hosted endpoint gives no seed you control and contracts no bit-identical output. State the settings you used, how many attempts you ran, and how many succeeded, so a reader re-runs the attempt the same way you did.
solid answer
~50 sA seed pins a pseudo-random draw; temperature is a setting that shapes the distribution the draw comes from. Setting temperature to 0 pushes decoding toward the highest-probability token and removes most of the sampling variance, but it is a request parameter, not a determinism contract: on a hosted endpoint you do not control batching, routing or the served build, and providers generally do not promise bit-identical outputs. Where an endpoint accepts a seed at all, treat it as best-effort and record it as one more input rather than as proof. So the finding should not say "deterministic at temperature 0". It should say: these were the decoding settings sent, this many attempts were made under them, this many produced the behaviour, and here is the verbatim successful transcript. That framing survives being re-run by someone else. It also makes the record honest about the thing you actually measured — attempts and outcomes — rather than claiming a repeatability the endpoint never offered.
go deeper
Knows temperature controls randomness and that low temperature makes outputs more repeatable, but may believe zero means deterministic.
Separates seed from temperature, knows temperature is a request parameter with no determinism guarantee on a hosted endpoint, and reports attempts and successes instead.
Enumerates the non-sampling sources of drift — served build, routing, batching, retrieval and tool results — and writes the reproduction section as a protocol with timestamps and request ids.
Makes attempts-and-successes plus effective on-the-wire settings a required field in every finding the org files, so downstream rating never has to guess the denominator.
### Two mechanisms that get conflated Generation picks one token at a time. The model produces a score for every token in its vocabulary; those scores are turned into a probability distribution; the distribution is truncated (top-p, top-k); and then one token is **drawn** from what remains. - **Temperature** is a request parameter that reshapes the distribution *before* the draw. High temperature flattens it, low temperature sharpens it. At zero the implementation ordinarily stops sampling and takes the highest-scoring token — greedy decoding. - **A seed** pins the pseudo-random draw itself. Same distribution plus same seed gives the same token. They are not substitutes. Temperature narrows the space the draw happens in; a seed fixes which draw you get. Writing "set temperature to 0" in a reproduction section records the first and claims the second. ### Why temperature 0 still is not determinism on a hosted endpoint Greedy decoding removes the *sampling* randomness. It removes nothing else, and on a hosted endpoint the rest is substantial: - **Numerics.** Your request is batched with other traffic. Batch composition changes the order of floating-point reductions inside the model's arithmetic, which perturbs scores in the last decimal places. Where the top two tokens are near-tied, that perturbation flips the argmax — and one flipped token early in a generation sends the rest of the output somewhere else entirely. - **Hardware and routing.** Different replicas, different accelerator generations, different kernel versions produce slightly different numerics for the same weights. - **The served build.** The provider can roll a new build behind a stable endpoint name without telling you. - **Everything outside the model.** Retrieval results, tool call results, and a moderation layer that may or may not fire — none of which any decoding parameter touches. Some providers accept a seed parameter and return a backend fingerprint so you can at least detect when the serving stack changed. Treat both as best-effort: the seed is documented as an attempt at determinism, not a contract, and the fingerprint tells you *that* something moved, not what. ### What the honest reproduction section says, and what it costs Not "deterministic at temperature 0", but a protocol: the decoding settings actually sent on the wire, the number of attempts run under them, the number that produced the behaviour, the verbatim successful transcript, at least one verbatim failure, and the timestamps and request ids bracketing the window. That costs calls. A single success proves the behaviour is reachable but supports no rate at all. Twenty attempts against a hosted chat endpoint is a few cents to a couple of dollars for short single-turn prompts, and minutes of wall clock; for a multi-turn attack at eight turns each, twenty attempts is a hundred and sixty calls and the wall clock is dominated by the serial turns, so budget tens of minutes rather than seconds. Spread them over time rather than firing them in one burst — twenty calls in thirty seconds are likely to land on one replica and one batch shape, which measures that route, not the endpoint. ### Where the number misleads Three specific misreadings, in the order they cause damage: 1. **"5/5 at temperature 0, therefore deterministic."** Five consecutive calls in one minute are close to one sample of one route. The claim survives your desk and dies on the first reader who gets a refusal, and it takes the report's credibility with it. 2. **A temperature-0 *negative* read as absence.** Greedy decoding samples the mode of the distribution. A behaviour that appears in three per cent of samples at the application's real decoding settings is essentially invisible at temperature zero — and three per cent of production traffic is a lot of traffic. "Does not reproduce at temperature 0" is not evidence the defect is gone; it is evidence about the mode only. 3. **Measuring at your settings, not the product's.** If the application ships at its own temperature and top-p, a rate measured at zero describes a configuration no user will ever hit. Report the rate at the production settings and, if you like, the temperature-0 result as a separate line. ### What you would check Confirm what the harness actually transmitted rather than what your config file says. Check whether the response carries a build or backend fingerprint, and whether it changed across your attempts — if it did, your attempts do not share a denominator. Re-run a slice at the application's real decoding settings. And when you write the rate, write the denominator beside it every time; a bare percentage with no attempt count behind it is the number that later cannot be defended.
- The endpoint does accept a seed parameter. Do you now claim reproducibility?No. Record it as another input. It constrains sampling only, and leaves the served build, routing and any retrieval or tool results untouched.
- Why include a verbatim failed attempt alongside the successful one?It shows the reader what non-reproduction looks like, so a refusal on their first try reads as expected variance rather than as a broken report.
Setting temperature to 0 is like switching to a heavily weighted die rather than not rolling at all. It almost always lands the same way, which is not a promise about any particular roll.
saying these in an interview costs you the question
- Writes 'deterministic at temperature 0' in the finding.
- Treats a seed parameter on a hosted endpoint as a reproducibility guarantee.
- Reports only the successful attempt and omits how many attempts were made.
- Records config-file settings rather than what the harness actually sent.
- Blames the reader's environment when a re-run does not reproduce, with no attempt count to point at.