What does temperature do to an LLM's next-token distribution during sampling?
answer
- a scalar applied before the softmax
- division changes gaps, not order
- below one focuses, above one wanders
- zero means greedy argmax
- rescales the tail, never removes it
basics
~20 sTemperature divides the model's raw scores (logits) before they are turned into probabilities. Below 1 it sharpens the distribution toward the top-scoring tokens; above 1 it flattens it, so rarer tokens get picked. At 0 it collapses to always taking the highest-scoring token.
solid answer
~50 sAt each step the model emits one score, a logit, per vocabulary token, and a softmax turns those scores into probabilities. Temperature T divides every logit by T first. T below 1 exaggerates the gaps, so probability mass piles onto the leading candidates and the output becomes conservative and repeatable; T above 1 compresses the gaps, so mid- and low-ranked tokens get realistic odds and the text gets more surprising. T = 0 is normally implemented as plain argmax — greedy decoding — because the formula itself would divide by zero. A game studio generating NPC barks and side-quest text feels this directly: around 1.1 with a nucleus cutoff you get genuine variety, while 0.2 gives you twelve near-identical fetch quests. Crucially, temperature rescales but never reorders: the most likely token is the same token at every temperature.
code
python · 16 linesimport math
def softmax_with_temperature(logits, t):
if t == 0:
top = max(range(len(logits)), key=lambda i: logits[i])
return [1.0 if i == top else 0.0 for i in range(len(logits))]
scaled = [x / t for x in logits]
m = max(scaled)
exps = [math.exp(x - m) for x in scaled]
total = sum(exps)
return [e / total for e in exps]
logits = [3.0, 2.0, 1.0, 0.0]
print([round(p, 3) for p in softmax_with_temperature(logits, 0.5)])
print([round(p, 3) for p in softmax_with_temperature(logits, 1.0)])
print([round(p, 3) for p in softmax_with_temperature(logits, 2.0)])go deeper
Be able to say that temperature rescales the model's scores before they become probabilities, that low values give focused repeatable text and high values give varied text, and that 0 means always take the top token.
Explain the mechanics: logits divided by T, then softmax, so only the gaps change. Point out that ranking is preserved and that no token is ever removed — removal is a truncation control's job, not temperature's.
Show judgement about which tasks get which value and why, and be ready to say that low temperature buys consistency rather than correctness. Mention pairing high temperature with a truncation cutoff so the flattened tail cannot inject junk.
Own the position that temperature is a per-task setting owned by the feature, not a global default, and that on endpoints which lock sampling the variability strategy has to move into prompt design, validation and model choice instead.
## Where temperature sits in the pipeline A language model is autoregressive: to produce one token it runs a forward pass and emits a vector of raw scores, one per vocabulary entry. Those raw scores are called logits, and they are unbounded real numbers, not probabilities. A softmax converts them into a probability distribution over the vocabulary. Temperature is applied between those two steps: every logit is divided by a scalar T, and the softmax is taken over the rescaled values. That single division is the whole mechanism. Everything temperature does follows from what dividing by T does to the *gaps* between scores, because softmax cares only about differences. ## The three regimes **T = 1** leaves the logits untouched. You sample from the distribution the model actually learned. **T < 1** magnifies every gap. If the top token led the runner-up by 2 logits, at T = 0.5 it leads by 4, and after softmax it takes a far larger share of the mass. The tail is squeezed toward zero. Output becomes focused, conventional and much more consistent across repeated calls. This is what you want for extraction, classification, routing decisions and anything where you would rather be boring than wrong. **T > 1** shrinks every gap. The distribution flattens toward uniform, and tokens the model considered mediocre start winning often enough to matter. Output becomes varied and, past roughly 1.3-1.5 on most models, incoherent — you are asking the model to act against its own judgement. **T = 0** is a special case. The formula divides by zero, so serving stacks special-case it into greedy decoding: take the argmax at every step. This is the most repeatable setting available, though — as with any hosted service — it is not a guarantee of identical bytes across calls. ## What temperature does not do Three misconceptions are worth naming explicitly. First, temperature does not remove any token from consideration. Even at 0.1 every vocabulary entry keeps a nonzero probability; it is just vanishingly small. Removing candidates outright is the job of truncation controls, which cut the tail before sampling. Temperature and truncation are complementary, and stacks apply them as a chain. Second, temperature does not reorder the candidates. Division by a positive constant is monotonic, so the rank order of tokens is identical at 0.1 and at 2.0. Raising temperature does not make the model "prefer" different tokens; it makes the sampler pick further down the same list more often. Third, temperature is not a truthfulness dial. Lowering it makes hallucinations more *consistent*, not less likely — you get the same confident wrong answer every time instead of a different one each call. If the model's leading candidate is wrong, greedy decoding will take it with certainty. ## Choosing a value The honest rule is that the right value is a property of the task, not of the model. Deterministic-feeling work (structured extraction, tool argument filling, grading, routing) sits at or near 0. Ordinary assistant prose sits somewhere in 0.5-0.8. Creative generation where variety across calls is the product — dialogue lines, flavour text, brainstorm lists — sits near or above 1.0, usually paired with a truncation cutoff so the flattened tail cannot inject genuine nonsense. One practical note for creative work: if repeated calls with the same prompt feel same-y, that is often not a temperature problem at all but a prompt problem, because the prompt itself pins the distribution hard. Varying the prompt (different constraints, different seed facts) changes the distribution being sampled from, which is a stronger lever than turning the knob further up. ## Where it applies in 2026 Temperature remains a first-class control in open-weight serving stacks that you run yourself. On hosted frontier reasoning endpoints the picture has changed: several vendors now lock sampling parameters and expose their own coarse control instead, so code that assumes it can always pass a temperature will break against those endpoints. Treat availability as a property of the target you are calling, not as a universal.
- If temperature never changes the ranking of tokens, why does high temperature produce visibly different text?Because generation is sequential. A single off-rank token early in the sequence conditions everything after it, so the model continues coherently from a different starting point. The ranking is preserved at each individual step, but the *path* through the sequence diverges, and divergence compounds over hundreds of tokens.
- Would you ever set temperature above 1 in production, and how would you guard it?Yes, for generation where variety is the product — flavour text, alternative phrasings, idea lists. Guard it by pairing it with a truncation cutoff so the flattened tail cannot admit genuine junk, and by validating output before it ships: length checks, banned-content checks, or a schema. High temperature without truncation is the setting that produces word salad.
- Does lowering temperature reduce hallucination?No. It makes the model's existing leading candidate win more often, so a wrong answer becomes a consistently wrong answer rather than a rarer one. Hallucination is addressed by grounding, retrieval, verification and better prompts. Low temperature buys reproducibility and stylistic conservatism, not factuality.
Think of the logits as heights of hills the sampler rolls a ball down. Low temperature makes the tallest hill tower over everything so the ball almost always lands there; high temperature levels the landscape so smaller hills win their share.
saying these in an interview costs you the question
- Says temperature 0 makes the model factually correct
- Claims temperature filters out low-probability tokens
- Thinks higher temperature changes which token ranks first
- Treats temperature as a creativity slider with no distribution meaning
- Assumes every endpoint accepts a temperature parameter