skip to content

In SFT, why must training data use the model's own chat template?

level: middleimportance: should knowfreq 55%

answer

  1. surface form the model was aligned on
  2. train and serve must render alike
  3. the token that says 'I'm done'
  4. no error, just a quiet gap
  5. render once, from the model's own code

basics

~20 s

Serving renders every request with the model's specific role markers and end-of-turn tokens. Training on a hand-rolled format teaches a different surface form, so the model meets an unfamiliar prompt shape at inference and quality drops quietly, with no error anywhere.

solid answer

~50 s

An instruction-tuned chat model was aligned against one exact rendering of a conversation: particular role markers, particular special tokens, a particular end-of-turn token. That rendering ships with the model as its chat template, and the serving stack applies it to every request. If your SFT data is built by concatenating hand-written strings like `Human:` and `Assistant:`, you have trained on a format that never appears in production, so at inference the model is conditioned on a prompt shape it has weak priors for. The result is not a crash — it is a quiet quality gap that is easy to misattribute to the dataset. The second, sharper failure is the end-of-turn token: if it is not appended to your training targets, the model never learns to stop and rambles past its answer. The fix is to render every example with the same template call the serving path uses, and to check one decoded example by eye before launching.

go deeper

for a junior

Know that a chat model expects one specific text layout of roles and turn markers, that this layout ships with the model, and that you should generate it with the provided template rather than writing the strings yourself.

for a middle

Explain that the template inserts role markers and the end-of-turn token, that serving applies it on every request, and that a training/serving mismatch degrades quality without erroring. Name the missing-stop-token symptom.

for a senior

Demonstrate the operational discipline: one shared rendering function for training, eval and serving; raw structured messages as the dataset of record; a decoded golden example inspected before every run; assertions on marker and stop-token counts.

for a principal

Own it as a platform concern. Argue that rendering belongs in one versioned library shared by data generation, training and inference, with a fixture test per supported base model, so a model-family upgrade cannot silently invalidate every dataset a team owns.

## What a chat template is A base language model sees only a flat token stream. Everything that makes a conversation a conversation — who is speaking, where a turn ends, where the system instruction sits — is encoded as literal tokens in that stream. A **chat template** is the deterministic function that turns a structured list of messages (`system`, `user`, `assistant`) into that flat string, inserting the model family's specific markers and special tokens and appending the end-of-turn token after each turn. The template is model-specific and it ships with the model. Two models from different families use different markers; two versions of the same family sometimes differ too. That is precisely why you must apply *the model's own* template rather than one you like. ## Why a mismatch hurts An instruction-tuned model has been trained, at scale, that the tokens following its assistant marker are helpful assistant text. That is an enormously strong prior, and it is attached to that exact surface form. When you fine-tune on `Human: ... Assistant: ...` plain text and then serve behind a chat API that renders the model's real markers, three things go wrong at once. **You fight the existing prior instead of building on it.** Your examples land in a region of format space the model has weak associations for, so more of the gradient budget goes into re-learning "this shape means assistant turn" and less into learning your actual task. **Train and serve disagree.** The behaviour you measured on your own rendered eval prompts is not the behaviour production gets, because production renders differently. This is the classic train/serve skew, and it makes offline results untrustworthy in a way that is very hard to notice — nothing errors, tokenisation succeeds, output is fluent. **System-prompt handling degrades.** Models are trained to give the system turn a particular authority. Fold your system instruction into the user turn as plain prose and you lose that, which shows up first as weaker instruction adherence and looser refusal behaviour. ## The end-of-turn token This is the single highest-yield detail. The template appends an end-of-turn token after the assistant turn, and that token must be part of the *supervised* target, not just present in the input. If it is missing, or if it is present but sits outside the trained span, the model has never been shown how a turn terminates. At inference it produces a good answer and then keeps going — inventing a follow-up question, restating the answer, drifting into unrelated text — until it hits the maximum token limit. Teams routinely misread this as a decoding-parameter problem and spend a day on sampling settings before finding the data bug. A related trap is stop-token drift: if a model family changes its end-of-turn token between versions and your pipeline hard-codes the old string, you silently train the wrong stop behaviour. ## Getting it right in practice The hygiene is unglamorous and it is most of the job. 1. **Render with the tokenizer's own template function, never with string concatenation.** The template is code that ships with the model; call it. 2. **Render exactly once.** The most common double-template bug is applying the template to text that already contains markers, producing nested system blocks. Because the result is still valid text, nothing complains. 3. **Use the same rendering code path for training data, evaluation prompts, and any offline scoring.** If the serving stack adds a generation prompt suffix, your eval must add it too. 4. **Decode and read one full example before launching.** Print the rendered string with special tokens visible, and confirm: one system block, correct role markers, an end-of-turn token after the assistant turn, and no stray leading whitespace or duplicated markers. 5. **Keep the raw structured messages as the dataset of record** and render at load time. Storing pre-rendered strings means the day you switch base models, your whole dataset is silently formatted for the wrong family. ## Where this sits relative to other choices Template correctness is independent of everything people usually argue about. It does not depend on how many parameters you update, on how much data you have, or on the training schedule. It is pure data hygiene — and in real SFT projects, formatting defects (templates, turn boundaries, stop tokens, mask boundaries) account for a larger share of failed runs than any modelling decision. Raw exports make this worse: a contact-centre transcript dump with inconsistent turn boundaries has to be restructured into clean role-tagged messages *before* any template can be applied correctly. An interviewer asking this wants to hear that you treat rendering as a shared, tested code path rather than a scripting detail, and that you can name the concrete symptom — a model that will not stop — that follows from getting the stop token wrong.

  • What is the concrete symptom of omitting the end-of-turn token from the training targets?
    The model produces a correct answer and then does not stop — it invents a follow-up question, restates itself, or drifts into unrelated text until the token cap. Teams often misdiagnose this as a sampling or stop-sequence configuration issue. The fix is in the data: append the model's end-of-turn token and make sure it falls inside the supervised span.
  • Why store raw message lists rather than pre-rendered strings in your dataset?
    Rendering is model-specific, so a pre-rendered corpus is silently bound to one model family. Keeping structured `role`/`content` messages and applying the template at load time lets you retarget a different base model without regenerating data, keeps training and evaluation on the same rendering code path, and makes double-templating easy to detect.
  • How would you catch a double-templating bug?
    Decode one rendered example with special tokens visible and read it. Double templating shows up as nested or duplicated system blocks and repeated role markers. It never raises an error because the output is still valid text, so an automated check should assert the expected count of role markers and end-of-turn tokens per example rather than relying on the run failing.

saying these in an interview costs you the question

  • Assumes any 'Human:'/'Assistant:' format works as well as the real template
  • Thinks a format mismatch produces an error rather than silent degradation
  • Forgets the end-of-turn token in the training targets
  • Applies the template twice to already-rendered text
  • Renders training data one way and evaluation prompts another

context