skip to content

How do top-k, top-p and min-p differ as token truncation strategies?

level: middleimportance: must knowfreq 62%

answer

  1. three ways to draw the same line
  2. fixed count versus running total
  3. one measures mass, one measures ratio
  4. confidence should change the set size
  5. the fat tail after a near-certain token

basics

~20 s

Top-k keeps a fixed number of highest-probability tokens. Top-p (nucleus) keeps the smallest set whose probabilities sum past a threshold, so the set size shrinks when the model is confident. Min-p keeps every token above a fraction of the top token's probability.

solid answer

~50 s

All three cut the tail before sampling, but they measure the cut differently. **Top-k** uses a fixed count: keep the 40 best candidates whatever the shape of the distribution — which is too narrow when the model is genuinely uncertain and too wide when one token is nearly certain. **Top-p** uses cumulative mass: sort descending and keep tokens until the running sum passes p, so the set adapts to confidence. **Min-p** uses a relative floor: with min_p = 0.1 and a top token at 0.9, only tokens at or above 0.09 survive. The difference shows on a low-entropy continuation. With top_p = 0.99 and a near-certain leading token, the nucleus must still absorb the leftover mass, so dozens of junk candidates ride along; min-p prunes them because they are far below the leader. That is why min-p is the preferred pairing with high temperature in open-weight stacks.

code

python · 19 lines
python
probs = {"of": 0.90, "in": 0.02, "for": 0.015, "on": 0.01}
fringe = {f"t{i}": 0.0011 for i in range(50)}
probs.update(fringe)

def top_p(d, p):
    kept, total = [], 0.0
    for tok, q in sorted(d.items(), key=lambda kv: -kv[1]):
        kept.append(tok)
        total += q
        if total >= p:
            break
    return kept

def min_p(d, m):
    floor = max(d.values()) * m
    return [tok for tok, q in d.items() if q >= floor]

print(len(top_p(probs, 0.99)))
print(len(min_p(probs, 0.1)))

go deeper

for a junior

Know that these controls cut off unlikely tokens before sampling, and that top-k keeps a fixed count while top-p keeps a share of the probability mass. Say plainly that top-p adapts to how confident the model is.

for a middle

Explain each rule precisely — fixed count, cumulative sum past a threshold, fraction of the top token's probability — and give the concrete case where they diverge: a near-certain leading token, where a high nucleus still admits a long fringe.

for a senior

Show that you have tuned these on real traffic: which cutoff you pair with which temperature, why you disable the ones you are not using, and why the same numbers do not transfer between serving engines that apply the filters in different orders.

for a principal

Own the position that sampling settings are per-workload and only tunable where you control serving; on hosted endpoints that lock these parameters, output-shape control has to be bought with prompts, validation and model choice instead.

## Why truncate at all A model's output distribution has a very long tail: tens of thousands of tokens each holding a tiny slice of probability. Summed, that tail is not tiny. Sample straight from the raw distribution often enough and you will eventually draw something the model considered nearly absurd — and because generation is autoregressive, one absurd token derails everything after it. Truncation removes the tail before the draw, so randomness stays inside the region the model actually endorses. Temperature and truncation solve different halves of the problem: temperature controls *how much* the sampler prefers the leaders, truncation controls *which candidates exist at all*. ## Top-k: a fixed count Sort tokens by probability, keep the first k, renormalize, sample. Simple and cheap, and it was the first widely used cutoff. Its weakness is that k is blind to the shape of the distribution. Consider two steps in the same generation. At one, the model is finishing the word "Constantino-" and one continuation carries essentially all the mass; keeping 40 candidates means 39 of them are noise that the renormalization now hands real probability to. At the next, the model is choosing the opening adjective of a sentence and a hundred options are genuinely plausible; k = 40 arbitrarily amputates sixty of them. One fixed number cannot serve both. ## Top-p (nucleus): cumulative mass Sort descending, accumulate probabilities, and keep tokens until the running total passes p. The candidate set is the *nucleus* — the smallest prefix of the sorted list holding at least p of the mass. This adapts: when the model is confident the nucleus is one or two tokens, when it is unsure the nucleus grows to dozens. That adaptivity is why nucleus sampling replaced top-k as the default and why it remains the most commonly configured cutoff. Its weakness is that it counts mass, not quality. If p is set high — 0.95 or 0.99, which is common — then after a leading token at 0.9 the nucleus still has to absorb another 0.05 to 0.09 of mass, and it collects that from whatever comes next: a long fringe of tokens each worth a fraction of a percent. Those tokens are, relative to the leader, nine hundred times less likely, yet they are now legal draws. Raise temperature on top and their odds grow further. ## Min-p: a relative floor Min-p sets the threshold as a fraction of the top token's probability. With min_p = 0.1 and a leader at 0.9, the floor is 0.09 and only tokens meeting it survive; with a flat distribution whose leader is 0.05, the floor drops to 0.005 and a wide set survives. So min-p is strict exactly where nucleus sampling is loose — on peaked, low-entropy steps — and permissive exactly where the model has genuine options. This is why min-p became the recommended companion to temperatures above 1 in open-weight serving: it lets you flatten the distribution for variety without opening the door to the fringe, because the fringe is judged against the leader rather than against a mass budget. ## Combining them, and ordering These filters are not exclusive; serving stacks apply them as a chain, and the order in which a given stack applies temperature and each cutoff is part of its semantics. Two practical consequences. First, stacking aggressive settings compounds: top_k = 10 with top_p = 0.5 and min_p = 0.2 can leave a single legal candidate, which is greedy decoding with extra steps and a puzzling one to debug. Second, because implementations differ in ordering, the same nominal numbers can behave differently across stacks — settings are not portable by default and should be re-tuned when you change engine. ## How to pick Disable what you are not deliberately using; many defaults leave k or p at a value nobody chose. For deterministic-feeling work, temperature near 0 makes cutoffs almost irrelevant. For ordinary prose, a nucleus around 0.9-0.95 at moderate temperature is the safe baseline. For high-variety generation, high temperature plus min-p is the modern pairing. And treat any of these as tunable only where you control serving: hosted frontier reasoning endpoints increasingly reject sampling parameters outright, so a cutoff strategy is something you own on open-weight infrastructure.

  • If min-p handles peaked distributions better, why is top-p still the default almost everywhere?
    Inertia and ecosystem coverage. Nucleus sampling arrived years earlier, every API and framework exposes it, and prompts, presets and tutorials are written against it. Min-p is newer, is not offered by every hosted endpoint, and its benefit is most visible at temperatures above 1 — a regime many production systems never enter. For temperature at or below 1, the two often behave similarly.
  • What goes wrong if you set top_k, top_p and min_p aggressively at the same time?
    They intersect. Each filter removes candidates the others kept, so the surviving set can collapse to one or two tokens and the output becomes effectively greedy despite a high temperature. The symptom — repetitive text that ignores your temperature setting — is confusing to debug. Pick one primary cutoff and leave the others disabled unless you have measured a reason.
  • How does the choice of cutoff interact with temperature?
    Temperature reshapes the distribution and the cutoff decides what survives, so the order the stack applies them changes the result. Practically: high temperature widens what a mass-based cutoff admits, because the flattened tail carries more mass, while a relative floor like min-p tracks the leader and stays comparatively stable. That is the core reason high-temperature setups favour min-p.

Top-k invites a fixed number of guests; top-p keeps inviting until the room holds a set share of the total importance; min-p only invites people at least a tenth as important as the guest of honour.

saying these in an interview costs you the question

  • Says top-p keeps a fixed number of tokens
  • Thinks top-p and temperature are the same control
  • Claims min-p uses an absolute probability floor
  • Assumes settings transfer unchanged across serving stacks
  • Sets top_k, top_p and min_p aggressively all at once

context