skip to content

Why can a fine-tuned Gemma 3 degrade at inference if the chat template is applied wrong?

level: seniorimportance: should knowfreq 38%

answer

  1. Trained format is the served format
  2. The assistant role is not called assistant
  3. One BOS, exactly one
  4. Something has to stop the turn
  5. Loss belongs on completions only

basics

~20 s

Gemma is trained on an exact turn format — start_of_turn markers with user and model roles, closed by end_of_turn, and a single leading BOS token. Serving prompts that differ from the training format, or that duplicate BOS, puts the model off-distribution and quality drops silently.

solid answer

~50 s

Gemma's conversation format is literal: each turn opens with `<start_of_turn>` plus a role and closes with `<end_of_turn>`, and the only two roles are **user** and **model** — not "assistant". There is no separate system role; a system instruction is folded into the first user turn rather than getting its own marker. Three failure modes follow. **Format drift** — you fine-tune with the tokenizer's template but serve with a hand-built string, so training and inference disagree and the model produces rambling or role-confused output. **Double BOS** — `apply_chat_template` already emits the leading `<bos>`, so tokenising its output again with special tokens enabled prepends a second one, which measurably hurts quality. **Wrong stop token** — generation must stop at `<end_of_turn>`; miss it and the model happily writes the user's next turn for them. None of these throw an error, which is why they survive to production.

code

python · 11 lines
python
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("google/gemma-3-4b-it")
messages = [{"role": "user", "content": "Summarise this in one line."}]

# Correct: let the template produce ids, so BOS is added exactly once
ids = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=True)

# Wrong: the rendered string already starts with <bos>, and this adds another
text = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
double_bos = tok(text)["input_ids"]

go deeper

for a junior

Know that instruction-tuned models expect one exact prompt format and that you should render it with the tokenizer's chat template rather than assembling the string yourself.

for a middle

Explain Gemma's turn grammar — start_of_turn with user or model, closed by end_of_turn, a single leading BOS — and name what breaks when serving deviates from the training format.

for a senior

Show you diagnose silent quality loss by diffing token IDs between the evaluation and production paths, and that you configure stop conditions and loss masking deliberately rather than by default.

for a principal

Own the systemic fix: prompt rendering should live in one shared component used by training, evaluation and serving, so a template change cannot desynchronise them across teams or languages.

## Why a template exists at all An instruction-tuned model has been trained on text laid out in one specific way: special tokens that mark where a turn begins, who is speaking, and where the turn ends. The model's behaviour — following instructions, stopping when it is done, staying in character — is learned *conditional on that layout*. Feed it a different layout and you are running the model off-distribution. Nothing errors; the weights happily process any tokens. The output just gets worse in ways that are hard to attribute. ## Gemma's specific format The shape is: - A single **`<bos>`** token at the very start of the sequence. - Each turn wrapped as **`<start_of_turn>`role**, newline, content, then **`<end_of_turn>`**. - Exactly two role names: **`user`** and **`model`**. - To elicit a reply, the prompt ends with an opened, empty model turn — the "generation prompt". Two details trip people who arrive from other families. First, the assistant role is spelled **model**, not `assistant`; code that hardcodes the other name silently produces a role Gemma never saw in training. Second, Gemma has **no dedicated system role** in its turn grammar — a system instruction is merged into the first user turn by the template rather than emitted as its own marked turn. ## Failure mode 1: training and serving disagree The canonical bug. Fine-tuning is done with the tokenizer's own chat template, which produces exactly the right token sequence. Then the serving path — a different service, maybe a different language — builds the prompt with string concatenation someone wrote from memory. A missing newline, `assistant` instead of `model`, a forgotten `<end_of_turn>`: each is a small distributional shift, and the fine-tune's gains are the first thing to disappear because they were the most narrowly learned behaviour. The fix is structural, not diligent: **render prompts through the tokenizer's template on both paths**, or better, let the serving stack apply the model's own template and never hand it a pre-formatted string. ## Failure mode 2: double BOS The template output already begins with `<bos>`. If you then pass that string to the tokenizer with special-token addition left enabled, the tokenizer prepends another one. The prompt now starts with two BOS tokens, a sequence the model has never seen. This one is especially nasty because it is invisible in logs — the decoded text looks right — and it is a genuine quality regression, not a cosmetic one. The fix: when tokenising template output, disable automatic special tokens; or ask the template to return token IDs directly rather than a string. ## Failure mode 3: the stop token A turn ends at `<end_of_turn>`. If your generation configuration stops only at the classic end-of-sequence token or at nothing at all, the model finishes its answer and then keeps going — often inventing the user's next message and answering that too. Users see a reply with a hallucinated conversation stapled to the end. Set the stop condition to include `<end_of_turn>`. ## Failure mode 4: training on the wrong tokens During supervised fine-tuning you want the loss computed on the **model turns only**. If you compute loss over the whole rendered conversation, the model spends capacity learning to generate user messages and the structural markers, which is not the behaviour you are paying for. Mask the prompt portion so the gradient sees only the completions. ## Diagnosing it in production Symptoms that point at template drift rather than at the fine-tune itself: - The model answers correctly but does not stop. - It emits role markers as visible text in the reply. - It addresses itself, or continues the user's message. - Base-model behaviour reappears — generic, unfollowed instructions — despite loading the tuned checkpoint. - Evaluation harness scores are fine but production quality is not; the harness uses the template and production does not. The cheapest diagnostic is to **log the exact token IDs** of one production prompt and diff them against the token IDs your evaluation path produces for the same conversation. Every one of the failures above shows up in that diff immediately, and none of them show up in a decoded-string comparison. ## The rule to state in an interview "Whatever format the model was trained on is the format you must serve, byte for byte, and the only reliable way to guarantee that is to render both paths through the tokenizer's own template rather than reconstructing it by hand." That sentence, plus the double-BOS example, is what separates someone who has shipped a fine-tune from someone who has read about one. ## Version caution Template details are tied to a model generation and to the packaged tokenizer configuration. Read the template shipped with the checkpoint you are actually loading rather than reusing one from an earlier generation.

  • How would you prove template drift is the cause rather than the fine-tune itself?
    Capture one real production request, render the same conversation through the tokenizer's chat template in your evaluation path, and diff the two token ID sequences. Template drift shows up as differing IDs at the turn boundaries — a doubled leading token, a missing end-of-turn, a different role string. If the IDs match exactly and quality still differs, the problem is genuinely the model or the decoding parameters, and you can move on.
  • Gemma has no dedicated system role. Where does a system instruction go?
    The template merges it into the first user turn, so the model sees the instruction as the opening of the user's message rather than as a separately marked channel. Practically this means system content is not privileged the way a distinct role would make it, and long system instructions consume the same context as user text. If you fine-tune, keep the same placement in training data that your serving template will use.
  • Why mask the loss to model turns during supervised fine-tuning?
    Because you want the model to learn how to answer, not how to write user messages or emit turn markers. Computing loss over the full rendered conversation spends gradient on reproducing prompts, which dilutes the signal you care about and can make the model volunteer the next user turn at inference. Masking the prompt region is the default in mature training harnesses for exactly this reason.

saying these in an interview costs you the question

  • Using the assistant role name in Gemma prompts
  • Hand-building prompt strings instead of using the template
  • Assuming a duplicated BOS token is harmless
  • Leaving end_of_turn out of the stop conditions
  • Training loss computed over the entire conversation

context