Why can a Llama prompt end up with two BOS tokens, and how do you avoid it?
answer
- the template already emits it
- then you encode the string again
- defaults add specials on encode
- exactly one, at position zero
- add_special_tokens=False
basics
~10 sThe rendered chat template already contains the model's beginning-of-sequence token. Tokenizing that string again with default settings prepends a second one. Pass add_special_tokens=False when encoding pre-rendered text, or let the template tokenize directly.
solid answer
~50 sA Llama chat template renders BOS into its output — `<s>` for Llama 2, `<|begin_of_text|>` for Llama 3. The classic two-step pipeline is `apply_chat_template(..., tokenize=False)` to get a string, then `tokenizer(prompt, return_tensors="pt")` to encode it. That second call adds special tokens by default, so BOS appears twice at position 0 and 1. The fix is either to skip the round-trip (`apply_chat_template(..., tokenize=True)`) or to encode with `add_special_tokens=False`. The same duplication happens when you concatenate a BOS yourself in front of a rendered prompt, or when an OpenAI-compatible server applies the template to a string you already rendered. What makes it interview-worthy is that nothing errors. Recent transformers versions log a warning, but the observable effect is a small, hard-to-attribute quality regression — which is why you assert on the leading token IDs in a test rather than eyeballing output.
code
python · 11 lines# Wrong: BOS twice
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
bad = tok(text, return_tensors="pt")
# Right, option A: let the template tokenize
good_a = tok.apply_chat_template(
msgs, add_generation_prompt=True, return_tensors="pt"
)
# Right, option B: keep the string, skip the extra specials
good_b = tok(text, add_special_tokens=False, return_tensors="pt")go deeper
Know that the chat template already includes the beginning-of-sequence token, so you should not add one yourself or re-tokenize the rendered string with default settings.
Explain that encoding defaults to add_special_tokens=True and name both fixes: tokenize inside apply_chat_template, or encode the rendered string with add_special_tokens=False.
Describe how you would find this in a live system — inspecting leading token IDs, logging the rendered prompt head at startup, treating the transformers duplicate-BOS warning as a build failure — and why the symptom is quiet quality loss.
Set the invariant across teams: exactly one BOS, produced by exactly one layer of the stack. Decide whether rendering happens client-side or in the inference server, and enforce it with contract tests so the two never both apply the template.
## Where the duplicate comes from Hugging Face tokenizers have a `add_special_tokens` argument that defaults to `True` on the encoding call. For Llama tokenizers that means "prepend BOS". Chat templates, meanwhile, render BOS themselves because BOS is part of the trained layout. Put the two together in the obvious order and you get it twice: ``` text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True) ids = tok(text, return_tensors="pt") # BOS from the template + BOS from here ``` Three other routes to the same bug: 1. **Manual concatenation** — prepending `"<s>"` or `"<|begin_of_text|>"` to a rendered prompt "to be safe". 2. **Double template application** — rendering client-side, then posting the string to a chat endpoint whose server renders again. That duplicates the whole envelope, BOS included. 3. **Fine-tuning data prep** — building training examples with the template and then letting a collator or a trainer add specials again. The model then *learns* the doubled prefix, and inference without it becomes the mismatch. ## Why it matters, and how much BOS at position 0 is a strong positional anchor: attention sinks concentrate on it, and every layer has seen exactly one of them in every training sequence. Two shifts the whole sequence by one position and inserts a token the model never saw in that spot. The effect is not catastrophic — the model still answers — but it is a real distribution mismatch, and it is the kind of thing that turns a benchmark score down a few points with no other explanation. Because it degrades rather than fails, it can live in a codebase for months. ## Detecting it The cheap, deterministic check is to inspect the first two token IDs of what you are about to send: ``` ids = tok(text, add_special_tokens=False)["input_ids"] assert ids[:2].count(tok.bos_token_id) == 1 ``` Or decode the head of the prompt with special tokens visible and look at it once, at startup, in a log line. In a serving stack, log the first 40 characters of the rendered prompt on boot — that single line catches double-BOS, a missing generation prompt, and a wrong template all at once. Modern transformers releases detect the pattern and emit a warning along the lines of a duplicated BOS being found in the input; treat that warning as an error in CI rather than noise. ## The correct shapes - **One step:** `inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")` — the template tokenizes, nothing is added twice. - **Two steps, explicit:** render with `tokenize=False`, then encode with `add_special_tokens=False`. Use this when the string crosses a process boundary or you want it in a log. - **Server-side:** send the `messages` array to a chat endpoint and let the server render; never send a pre-rendered string to an endpoint that also templates. ## The mirror-image bug The opposite mistake is stripping BOS entirely — for example encoding with `add_special_tokens=False` a string you built by hand without the template's BOS. Zero BOS is as off-distribution as two. The invariant to state in an interview is simply: **exactly one BOS, at position zero, and it comes from exactly one place in your pipeline.**
- How would you catch this in CI rather than in production?Assert on token IDs, not on text: render a fixed set of messages, encode exactly as the serving path does, and check that the BOS ID appears exactly once and at index 0. Pair it with a golden-string test on the rendered prompt so a checkpoint swap or a template edit fails the build instead of quietly shifting quality.
- Is one extra BOS actually harmful, or is this pedantry?It is a genuine distribution mismatch. Every training sequence had exactly one BOS at position 0, and it acts as an attention anchor; a second one shifts positions and inserts an unseen pattern. The damage is small and non-deterministic rather than catastrophic, which is what makes it dangerous — it shows up as an unexplained few-point quality drop, not as a bug report.
- Where does the same duplication bite during fine-tuning?When you build training examples with the chat template and then let the tokenizer or collator add special tokens again. The model is trained on the doubled prefix, so inference with a correctly rendered single-BOS prompt becomes the mismatched case. Prepare training text with add_special_tokens=False and verify the first tokens of a decoded batch before launching a long run.
saying these in an interview costs you the question
- Prepends BOS manually to be safe on top of the template
- Thinks an extra BOS is harmless because output looks fine
- Assumes the tokenizer never adds special tokens by default
- Confuses add_special_tokens on encode with skip_special_tokens on decode
- Sends a pre-rendered prompt to an endpoint that templates again