skip to content

Why must Llama 3 Instruct generation stop on <|eot_id|>, not just EOS?

level: middleimportance: should knowfreq 48%

answer

  1. two different end tokens exist
  2. document end versus turn end
  3. the model writes the next turn itself
  4. pass a list of terminator ids
  5. <|eot_id|> versus <|end_of_text|>

basics

~20 s

Llama 3 Instruct closes each turn with <|eot_id|>; the base end-of-sequence token <|end_of_text|> almost never appears in chat output. Stop only on the base EOS and the model runs past its answer and writes the next user turn itself.

solid answer

~40 s

Llama 3's chat format has two distinct terminators. `<|end_of_text|>` is the pre-training end-of-sequence token; `<|eot_id|>` is the end-of-turn token the instruct tuning taught the model to emit when it finishes speaking. In a chat, the model emits `<|eot_id|>` — so that is the token your stopping criteria must include. If your decode loop only halts on the tokenizer's `eos_token_id` and that is still `<|end_of_text|>`, generation keeps going: the model happily produces `<|start_header_id|>user<|end_header_id|>` and hallucinates the rest of the dialogue until max_tokens. In transformers this is handled by passing a list of terminator IDs to `generate`, and official Llama 3 Instruct repos ship a `generation_config.json` whose `eos_token_id` is a list covering both. The symptom to recognise in an interview is a correct answer followed by an invented conversation.

code

python · 10 lines
python
terminators = [
    tokenizer.eos_token_id,
    tokenizer.convert_tokens_to_ids("<|eot_id|>"),
]

outputs = model.generate(
    **inputs,
    max_new_tokens=512,
    eos_token_id=terminators,
)

go deeper

for a junior

Know that Llama 3 chat replies end with <|eot_id|> and that your generation call must be told to stop there, otherwise the output keeps going past the answer.

for a middle

Explain the difference between the pre-training EOS <|end_of_text|> and the instruct turn terminator <|eot_id|>, and show how to pass both as terminator IDs to a generate call.

for a senior

Diagnose it in production: recognise fabricated follow-up turns as a stopping bug, check whether generation_config carried a multi-ID EOS, and note the wasted token budget and latency it causes on every request.

for a principal

Treat stop-token configuration as part of the model contract you validate when adopting a checkpoint or a community conversion, with a golden-output test that fails when a re-upload flattens the terminator list.

## Two terminators, one of which you rarely see Llama 3's vocabulary contains both `<|end_of_text|>` and `<|eot_id|>`. They mean different things: - `<|end_of_text|>` ends a **document**. It is the classic EOS from pre-training and is what a base (non-instruct) checkpoint emits. - `<|eot_id|>` ends a **turn**. Instruction tuning taught the chat model to emit it after each assistant reply, because the conversation continues afterwards with another role header. An instruct model in a chat context therefore almost always finishes with `<|eot_id|>`. Treating only the pre-training EOS as "stop" means the sequence never terminates on its own. ## The failure it produces Generation does not crash. The model has just closed its turn, and the most likely continuation given the training data is the next role header. So you get: ``` 4.<|eot_id|><|start_header_id|>user<|end_header_id|> And what about 5+5?<|eot_id|><|start_header_id|>assistant<|end_header_id|> 10. ``` Users see a real answer glued to a fabricated dialogue, or — if your client strips special tokens on decode — a wall of text where the model appears to interview itself. It also burns every remaining token of your `max_new_tokens` budget on every request, which shows up as latency and cost regressions long before anyone reads the output carefully. ## Getting it right In transformers, pass every acceptable terminator to `generate`: ``` terminators = [ tokenizer.eos_token_id, tokenizer.convert_tokens_to_ids("<|eot_id|>"), ] out = model.generate(**inputs, eos_token_id=terminators, max_new_tokens=512) ``` The official instruct repos already set a multi-ID `eos_token_id` in `generation_config.json`, so a plain `generate()` on freshly downloaded weights usually behaves. You hit the bug when: the config was not loaded (you built the model manually, or passed your own `GenerationConfig`); you are using a community re-upload or a quantized conversion whose config was flattened to a single ID; or you wrote your own sampling loop and compared against `tokenizer.eos_token_id` alone. Llama 3.1 adds a third terminator, `<|eom_id|>`, emitted when the assistant pauses for a tool result rather than yielding the turn. A tool-calling loop must treat that as a stop too — and must resume the conversation rather than returning to the user. ## Debugging it Decode with special tokens visible (`skip_special_tokens=False`) and look at the tail of the raw output. If you see `<|eot_id|>` sitting *inside* the returned text, the token was generated and simply not honoured as a stop — that pinpoints the stopping criteria rather than the prompt format. If you see no terminator at all and the text just runs to the length cap, suspect the prompt layout instead: a model that never saw a proper header sequence may never produce a clean turn boundary. ## The neighbouring mistake Do not "fix" this with a string-level stop sequence on `"user"` or `"\n\n"`. It half-works, cuts legitimate answers containing those strings, and hides the real cause. Stop on the token the format defines.

  • How would you tell whether the model failed to emit a stop token or your loop failed to honour it?
    Decode with `skip_special_tokens=False` and read the raw tail. If `<|eot_id|>` appears inside the returned text, the model stopped correctly and your stopping criteria ignored it. If no terminator appears anywhere and output simply hits the length cap, suspect the prompt layout — a malformed header sequence means the model never learned where this turn ends.
  • Llama 3.1 introduces a third terminator. What is it and how does it change the loop?
    `<|eom_id|>` — end of message. The assistant emits it when it has produced a tool call and expects execution output rather than handing the turn back to the user. Your loop must stop on it, run the tool, append the result as an `ipython` message, and continue generation, whereas `<|eot_id|>` means return control to the user.
  • Why is a string-level stop sequence a poor substitute here?
    Because it matches on decoded text, not on the boundary the format defines. Stopping on a literal like "user" or a double newline truncates legitimate answers that contain those strings, still lets a few wrong tokens be sampled and billed before the match, and papers over the real defect. Stop on the terminator token ID instead.

saying these in an interview costs you the question

  • Says <|eot_id|> and <|end_of_text|> are the same token
  • Blames rambling output on temperature rather than stop tokens
  • Fixes runaway generation with a string stop sequence on "user"
  • Assumes tokenizer.eos_token_id always covers the turn terminator
  • Thinks the model errors rather than continuing past its answer

context