skip to content

Inference and Sampling

What happens per generated token: autoregressive decoding, the temperature, top-k and top-p knobs, greedy versus beam search, and the KV cache and batching that set your latency and throughput. Expect to be asked which knob to turn for a given symptom.

on this pageshow

questions

15

What does temperature do to an LLM's next-token distribution during sampling?

level: juniorimportance: must knowfreq 82%

answer

  1. a scalar applied before the softmax
  2. division changes gaps, not order
  3. below one focuses, above one wanders
  4. zero means greedy argmax
  5. rescales the tail, never removes it

basics

~20 s

Temperature divides the model's raw scores (logits) before they are turned into probabilities. Below 1 it sharpens the distribution toward the top-scoring tokens; above 1 it flattens it, so rarer tokens get picked. At 0 it collapses to always taking the highest-scoring token.

solid answer

~50 s

At each step the model emits one score, a logit, per vocabulary token, and a softmax turns those scores into probabilities. Temperature T divides every logit by T first. T below 1 exaggerates the gaps, so probability mass piles onto the leading candidates and the output becomes conservative and repeatable; T above 1 compresses the gaps, so mid- and low-ranked tokens get realistic odds and the text gets more surprising. T = 0 is normally implemented as plain argmax — greedy decoding — because the formula itself would divide by zero. A game studio generating NPC barks and side-quest text feels this directly: around 1.1 with a nucleus cutoff you get genuine variety, while 0.2 gives you twelve near-identical fetch quests. Crucially, temperature rescales but never reorders: the most likely token is the same token at every temperature.

code

python · 16 lines
python
import math

def softmax_with_temperature(logits, t):
    if t == 0:
        top = max(range(len(logits)), key=lambda i: logits[i])
        return [1.0 if i == top else 0.0 for i in range(len(logits))]
    scaled = [x / t for x in logits]
    m = max(scaled)
    exps = [math.exp(x - m) for x in scaled]
    total = sum(exps)
    return [e / total for e in exps]

logits = [3.0, 2.0, 1.0, 0.0]
print([round(p, 3) for p in softmax_with_temperature(logits, 0.5)])
print([round(p, 3) for p in softmax_with_temperature(logits, 1.0)])
print([round(p, 3) for p in softmax_with_temperature(logits, 2.0)])

go deeper

for a junior

Be able to say that temperature rescales the model's scores before they become probabilities, that low values give focused repeatable text and high values give varied text, and that 0 means always take the top token.

for a middle

Explain the mechanics: logits divided by T, then softmax, so only the gaps change. Point out that ranking is preserved and that no token is ever removed — removal is a truncation control's job, not temperature's.

for a senior

Show judgement about which tasks get which value and why, and be ready to say that low temperature buys consistency rather than correctness. Mention pairing high temperature with a truncation cutoff so the flattened tail cannot inject junk.

for a principal

Own the position that temperature is a per-task setting owned by the feature, not a global default, and that on endpoints which lock sampling the variability strategy has to move into prompt design, validation and model choice instead.

## Where temperature sits in the pipeline A language model is autoregressive: to produce one token it runs a forward pass and emits a vector of raw scores, one per vocabulary entry. Those raw scores are called logits, and they are unbounded real numbers, not probabilities. A softmax converts them into a probability distribution over the vocabulary. Temperature is applied between those two steps: every logit is divided by a scalar T, and the softmax is taken over the rescaled values. That single division is the whole mechanism. Everything temperature does follows from what dividing by T does to the *gaps* between scores, because softmax cares only about differences. ## The three regimes **T = 1** leaves the logits untouched. You sample from the distribution the model actually learned. **T < 1** magnifies every gap. If the top token led the runner-up by 2 logits, at T = 0.5 it leads by 4, and after softmax it takes a far larger share of the mass. The tail is squeezed toward zero. Output becomes focused, conventional and much more consistent across repeated calls. This is what you want for extraction, classification, routing decisions and anything where you would rather be boring than wrong. **T > 1** shrinks every gap. The distribution flattens toward uniform, and tokens the model considered mediocre start winning often enough to matter. Output becomes varied and, past roughly 1.3-1.5 on most models, incoherent — you are asking the model to act against its own judgement. **T = 0** is a special case. The formula divides by zero, so serving stacks special-case it into greedy decoding: take the argmax at every step. This is the most repeatable setting available, though — as with any hosted service — it is not a guarantee of identical bytes across calls. ## What temperature does not do Three misconceptions are worth naming explicitly. First, temperature does not remove any token from consideration. Even at 0.1 every vocabulary entry keeps a nonzero probability; it is just vanishingly small. Removing candidates outright is the job of truncation controls, which cut the tail before sampling. Temperature and truncation are complementary, and stacks apply them as a chain. Second, temperature does not reorder the candidates. Division by a positive constant is monotonic, so the rank order of tokens is identical at 0.1 and at 2.0. Raising temperature does not make the model "prefer" different tokens; it makes the sampler pick further down the same list more often. Third, temperature is not a truthfulness dial. Lowering it makes hallucinations more *consistent*, not less likely — you get the same confident wrong answer every time instead of a different one each call. If the model's leading candidate is wrong, greedy decoding will take it with certainty. ## Choosing a value The honest rule is that the right value is a property of the task, not of the model. Deterministic-feeling work (structured extraction, tool argument filling, grading, routing) sits at or near 0. Ordinary assistant prose sits somewhere in 0.5-0.8. Creative generation where variety across calls is the product — dialogue lines, flavour text, brainstorm lists — sits near or above 1.0, usually paired with a truncation cutoff so the flattened tail cannot inject genuine nonsense. One practical note for creative work: if repeated calls with the same prompt feel same-y, that is often not a temperature problem at all but a prompt problem, because the prompt itself pins the distribution hard. Varying the prompt (different constraints, different seed facts) changes the distribution being sampled from, which is a stronger lever than turning the knob further up. ## Where it applies in 2026 Temperature remains a first-class control in open-weight serving stacks that you run yourself. On hosted frontier reasoning endpoints the picture has changed: several vendors now lock sampling parameters and expose their own coarse control instead, so code that assumes it can always pass a temperature will break against those endpoints. Treat availability as a property of the target you are calling, not as a universal.

  • If temperature never changes the ranking of tokens, why does high temperature produce visibly different text?
    Because generation is sequential. A single off-rank token early in the sequence conditions everything after it, so the model continues coherently from a different starting point. The ranking is preserved at each individual step, but the *path* through the sequence diverges, and divergence compounds over hundreds of tokens.
  • Would you ever set temperature above 1 in production, and how would you guard it?
    Yes, for generation where variety is the product — flavour text, alternative phrasings, idea lists. Guard it by pairing it with a truncation cutoff so the flattened tail cannot admit genuine junk, and by validating output before it ships: length checks, banned-content checks, or a schema. High temperature without truncation is the setting that produces word salad.
  • Does lowering temperature reduce hallucination?
    No. It makes the model's existing leading candidate win more often, so a wrong answer becomes a consistently wrong answer rather than a rarer one. Hallucination is addressed by grounding, retrieval, verification and better prompts. Low temperature buys reproducibility and stylistic conservatism, not factuality.

Think of the logits as heights of hills the sampler rolls a ball down. Low temperature makes the tallest hill tower over everything so the ball almost always lands there; high temperature levels the landscape so smaller hills win their share.

saying these in an interview costs you the question

  • Says temperature 0 makes the model factually correct
  • Claims temperature filters out low-probability tokens
  • Thinks higher temperature changes which token ranks first
  • Treats temperature as a creativity slider with no distribution meaning
  • Assumes every endpoint accepts a temperature parameter

context

open as a page

In LLM serving, what do time-to-first-token and inter-token latency each measure?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Time-to-first-token is the wait from sending a request until the first output token arrives, and it grows with prompt length. Inter-token latency is the gap between successive tokens after that, and it sets how fast the answer streams.

open as a page

How do top-k, top-p and min-p differ as token truncation strategies?

level: middleimportance: must knowfreq 62%

basics

~20 s

Top-k keeps a fixed number of highest-probability tokens. Top-p (nucleus) keeps the smallest set whose probabilities sum past a threshold, so the set size shrinks when the model is confident. Min-p keeps every token above a fraction of the top token's probability.

open as a page

Why can an LLM cache keys and values across decoding steps but not queries?

level: middleimportance: must knowfreq 62%

basics

~20 s

Under a causal mask, a past token's key and value vectors never change as new tokens arrive, so every later step reuses them. That token's query was consumed once, to produce its own output, and is never read again.

open as a page

Why is LLM prefill compute-bound while token-by-token decode is bandwidth-bound?

level: middleimportance: must knowfreq 70%

basics

~20 s

Prefill pushes every prompt token through the model in one pass, so each weight read is amortized over many tokens and the math units saturate. Decode produces one token per pass, re-reading all weights and the cache for that single token, so memory bandwidth is the ceiling.

open as a page

Why is single-stream LLM decoding limited by memory bandwidth rather than FLOPs?

level: middleimportance: must knowfreq 60%

basics

~20 s

Generating one token for one request reads every parameter that token's forward pass needs out of GPU memory but does only a couple of arithmetic operations per parameter. The accelerator therefore sits waiting on memory, so tokens per second track memory bandwidth, not peak FLOPs.

open as a page

Why does the KV cache, not model weights, cap how many long sessions fit on a GPU?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Model weights are loaded once and shared by every request, so they are a fixed cost. The KV cache is per session and grows with every token, so on long contexts it quickly dwarfs the weights and consumes whatever memory is left.

open as a page

How do repetition, frequency and presence penalties differ in LLM sampling?

level: middleimportance: should knowfreq 48%

basics

~20 s

A presence penalty subtracts a flat amount from any token that has already appeared, once, regardless of count. A frequency penalty subtracts an amount that grows with how many times the token appeared. A repetition penalty scales the token's score multiplicatively instead of subtracting.

open as a page

Why can an LLM still vary at temperature 0 with a fixed seed?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A seed only pins the sampler's random draws; it does not pin the numbers the model computes. Floating-point reductions are not associative, so results shift with how requests are batched on the server, and version, hardware and routing changes shift them further.

open as a page

How do grouped KV heads, latent compression or recurrent layers change cache bytes per token?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Each design changes the per-token constant. Sharing one key/value head across a group of query heads divides the bytes by the group size; storing a single compressed latent per token replaces the key/value pair entirely; recurrent layers keep a fixed-size state, so their cost does not grow with length at all.

open as a page

Raising server batch size from 8 to 64 triples throughput but slows each user — how do you choose?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Fix the user-facing latency target first, then sweep concurrency and pick the largest batch size whose p95 first-token and per-token latencies still meet it. Optimise for throughput that satisfies the target, not raw tokens per second, and re-check the tail rather than the mean.

open as a page

How should a product adapt when its endpoint rejects temperature and top_p?

level: principalimportance: should knowfreq 32%

basics

~20 s

Control of output variability moves out of the sampler and into the system around it: prompt specificity, constrained output, validators, and generating several candidates and selecting. Keep an open-weight lane only where tuned sampling is a genuine requirement, and never silently drop rejected parameters.

open as a page

How do you serve one LLM for interactive chat and an overnight 900K-listing scoring job?

level: principalimportance: should knowfreq 40%

basics

~20 s

Treat them as two deployments of the same model tuned to opposite ends of the latency-throughput frontier: the chat path optimised for tail first-token latency at modest concurrency, the scoring job optimised for tokens per accelerator-hour at maximum concurrency, with the bulk work scheduled into off-peak capacity.

open as a page

When is quantizing the KV cache to FP8 the right way to buy serving capacity?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Cache quantization is right when memory, not quality headroom, is the binding constraint and your own evals show the loss is tolerable at your longest contexts. Halving bytes per entry roughly doubles resident sessions and also speeds bandwidth-bound decode.

open as a page