How do repetition, frequency and presence penalties differ in LLM sampling?
answer
- once versus counted versus scaled
- the loop-breaker escalates with count
- lower temperature makes loops worse
- blind to meaning, punishes required tokens
- code and JSON are the casualties
basics
~20 sA presence penalty subtracts a flat amount from any token that has already appeared, once, regardless of count. A frequency penalty subtracts an amount that grows with how many times the token appeared. A repetition penalty scales the token's score multiplicatively instead of subtracting.
solid answer
~50 sAll three discourage the model from reusing tokens, but they measure reuse differently. **Presence** is a step function: a token that has appeared at all gets a constant subtraction from its score, so it nudges the model toward new vocabulary without escalating. **Frequency** scales with the occurrence count, so a token seen five times is penalised five times as hard — this is the one that breaks degenerate loops. **Repetition penalty**, the formulation inherited from open-weight tooling, is multiplicative: a seen token's score is divided by the factor when positive and multiplied by it when negative, which affects tokens differently depending on their sign. A quest generator emitting "the ancient ruins, the ancient ruins, the ancient ruins" is fixed with a modest frequency penalty, not by lowering temperature — sharpening the distribution makes a loop more likely, not less. The cost of all three is that they are blind to meaning: they punish a required JSON key or a character's name exactly as hard as filler.
code
python · 14 linesdef apply_penalties(logits, counts, presence=0.0, frequency=0.0, repetition=1.0):
out = dict(logits)
for tok, n in counts.items():
if n <= 0 or tok not in out:
continue
if repetition != 1.0:
out[tok] = out[tok] / repetition if out[tok] > 0 else out[tok] * repetition
out[tok] -= presence
out[tok] -= frequency * n
return out
logits = {"ruins": 4.0, "caverns": 1.5, "spire": 1.2}
counts = {"ruins": 3}
print(apply_penalties(logits, counts, presence=0.3, frequency=0.5))go deeper
Know that these settings discourage the model from repeating itself, and that the count-scaled one is what you reach for when output starts looping. Do not confuse them with temperature.
State the three rules precisely: constant subtraction once seen, subtraction proportional to count, and multiplicative scaling of the score. Explain why lowering temperature makes loops worse rather than better.
Show the production judgement: penalties are off by default, applied at the smallest value that clears an observed symptom, and never enabled on code or schema-bound paths where repeated tokens are required.
Own the framing that repetition is usually a prompt or task-decomposition problem before it is a sampling problem, and that per-route sampling policy — rather than one global default — is what keeps a shared inference layer from breaking one team's output to help another's.
## The failure they exist for Autoregressive models fall into loops. Once a phrase appears, it becomes part of the context conditioning the next step, which raises its own probability, which makes it appear again. The feedback is self-reinforcing and it is worse at low temperature, because a sharpened distribution keeps taking the same leading token. "The ancient ruins, the ancient ruins, the ancient ruins" is the canonical shape. The instinct to fix this by lowering temperature is exactly backwards. Loops are a symptom of the model over-committing to its leading candidate; sharpening makes it commit harder. The penalty family attacks the cause directly by modifying scores of tokens that already occurred. ## Presence penalty A presence penalty applies a constant subtraction to the logit of any token that has appeared in the text so far. It is binary in its trigger: appeared once or appeared fifty times, the penalty is identical. That makes it a *topic-broadening* control rather than a loop-breaking one. It pushes the model toward vocabulary it has not used yet, which is useful when you want a list of genuinely distinct items rather than five rephrasings of the first. It will not, on its own, stop a hard loop, because a constant subtraction is a constant the loop can outrun. ## Frequency penalty A frequency penalty subtracts an amount proportional to how often the token has already occurred. The second occurrence is discouraged a little, the tenth a lot. This escalation is exactly what a loop needs: each repetition raises the cost of the next one until the loop cannot sustain itself. It is the right knob for degenerate repetition, and typically the wrong knob for content variety, where presence is gentler. In APIs that expose both, they are independent and additive. ## Repetition penalty The repetition penalty from open-weight tooling works multiplicatively rather than additively. For a token that has already appeared, a positive score is divided by the penalty factor and a negative score is multiplied by it, so in both directions the score moves down. Values live near 1.0 — 1.0 means off, and small increments like 1.05-1.15 are the usual working range, with anything much higher visibly degrading fluency. Two things follow from the multiplicative form. First, the effect size depends on the magnitude of the logit, so it is not uniform across tokens. Second, it does not escalate with count in the way a frequency penalty does; a token seen once and a token seen ten times are scaled by the same factor. ## Where they hurt Every member of this family operates on tokens, with no notion of meaning or structure. That is the whole problem. In code generation, identifiers and keywords are *supposed* to repeat. A penalty strong enough to break prose loops will make the model rename a variable halfway through a function or reach for a synonym of a language keyword, producing code that does not compile. In structured output, repeated field names, brackets and quotes are mandatory. Penalties directly fight the format, and the failures are ugly: mismatched braces, invented key names, truncated arrays. In long-form narrative, a character's name legitimately recurs on every page. A frequency penalty accumulated over thousands of tokens will eventually push the model to stop using it, and the prose degrades into pronoun soup. The practical rule: default all of these to off, apply the smallest value that resolves an observed symptom, scope them to the generation type that has the symptom, and never leave them enabled on code or schema-constrained paths. ## Neighbouring controls Two related controls are worth distinguishing. **No-repeat n-gram blocking** is a hard constraint rather than a soft penalty: it forbids any n-gram from appearing twice, which reliably kills loops but also forbids legitimate repeated phrases, so it is rarely appropriate for open-ended text. **Stop sequences** are not penalties at all — they terminate generation when a given string is produced, which is how you keep a chat model from writing the other speaker's turn. And as with the rest of the sampling surface, availability is not universal: penalties remain live in open-weight serving stacks, while several hosted frontier reasoning endpoints have narrowed or removed the sampling parameters they accept.
- A model looping on a phrase is fixed by lowering temperature, according to a teammate. What do you say?That it usually makes things worse. Loops come from the model over-committing to its leading candidate, and lowering temperature sharpens that commitment. The targeted fixes are a frequency penalty, which escalates with each repeat, or a wider candidate set. Lowering temperature can appear to help only when the loop was driven by a flattened tail feeding junk back into the context.
- Why would you disable penalties entirely on a code-generation endpoint?Because penalties are token-level and structure-blind. Code requires repeated identifiers, keywords and punctuation; penalising them pushes the model to rename variables mid-function or substitute synonyms for keywords, producing output that does not compile. The same applies to JSON and any schema-bound output, where repeated field names and brackets are mandatory.
- How would you stop a dialogue model from also writing the player's next line?With a stop sequence on the string that begins the other speaker's turn — for example a newline followed by the player's label. Generation halts as soon as that text is produced and the stop text itself is normally not returned. This is a termination control, not a penalty: it does not change any probabilities, it just ends the response at a boundary you define.
saying these in an interview costs you the question
- Fixes a repetition loop by lowering temperature
- Treats presence and frequency penalties as the same control
- Leaves penalties on for JSON or code generation
- Thinks a repetition penalty always subtracts from the score
- Assumes penalties understand phrases rather than tokens