skip to content

What does Qwen's ChatML chat template wrap around each message, and why does it matter?

level: middleimportance: must knowfreq 55%

answer

  1. turn markers, not prose
  2. role name on the opening line
  3. the open assistant header at the end
  4. stop token is the turn terminator
  5. lives in tokenizer_config.json

basics

~20 s

Qwen uses ChatML: every message is rendered as <|im_start|>role, a newline, the content, then <|im_end|>. Generation starts from a trailing <|im_start|>assistant header, and <|im_end|> is the stop token. Wrong formatting means the model rambles or never stops.

solid answer

~40 s

Qwen models are trained on ChatML, so a conversation is flattened into blocks of `<|im_start|>{role}\n{content}<|im_end|>` for system, user and assistant turns, and the prompt ends with an open `<|im_start|>assistant\n` header — the *generation prompt* — that the model completes. `<|im_end|>` is the end-of-turn token the sampler stops on. The template itself is a Jinja string shipped in `tokenizer_config.json`, and `tokenizer.apply_chat_template(messages, add_generation_prompt=True)` renders it; GGUF builds carry the same template in file metadata. Any server's `/v1/chat/completions` applies it for you, which is exactly why hand-rolling prompts against `/v1/completions` is the classic self-hosting bug: without the special tokens the model has no turn structure, so it continues past the answer, invents both sides of the dialogue, or emits literal role markers. Qwen's template also renders `tools` and `role: "tool"` messages into its own tag syntax.

go deeper

for a junior

Know that Qwen expects ChatML with explicit turn markers and that the chat endpoint builds them for you — so send a messages array rather than a hand-written prompt string.

for a middle

Be able to write out the block structure from memory, name where the template lives in the repo, and explain what add_generation_prompt does and why the turn terminator is the stop token.

for a senior

Diagnose from symptoms: output that never stops or leaks role markers points at prompt rendering, not sampling. Show that you render the template to inspect what the model actually receives.

for a principal

Own the rule that prompts are rendered from the served checkpoint's own template across serving, evaluation and fine-tuning, so a model swap does not silently change what every pipeline sends.

## The template is part of the model A chat model is not trained on raw dialogue text; it is trained on a specific serialisation of dialogue. For Qwen that serialisation is ChatML, the same scheme popularised for GPT-style chat models. Getting it wrong is not a cosmetic issue — it puts the model off-distribution, and the symptoms look like the model being bad rather than the prompt being malformed. ## What ChatML looks like Each message becomes a delimited block: <|im_start|>system You are a helpful assistant.<|im_end|> <|im_start|>user What is the capital of Peru?<|im_end|> <|im_start|>assistant Three things to notice. **`<|im_start|>` and `<|im_end|>` are single tokens**, not literal character sequences the model reads letter by letter. They exist in the vocabulary as special tokens, which is what lets the model learn hard turn boundaries. Typing them as plain text into a raw completion endpoint may or may not tokenise identically, depending on how the server handles special tokens. **The role name sits on the opening line**, followed by a newline, then the content. Roles are `system`, `user`, `assistant` and `tool`. **The prompt ends with an open assistant header.** That trailing `<|im_start|>assistant\n` is what `add_generation_prompt=True` appends. Without it the model sees a completed conversation and is as likely to start a new user turn as to answer. ## Where the template lives The Jinja source is the `chat_template` field in `tokenizer_config.json` inside the model repo. In Python: prompt = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True ) That is the ground truth — if you ever wonder what a server is actually sending the model, render it yourself and print it. GGUF conversions embed the same template in the file's key-value metadata, so llama.cpp and runtimes built on it apply it without a separate file. Serving frameworks let you override it with a `--chat-template` argument, which you should only do when the shipped one is missing or lacks a feature you need. ## Why it matters operationally When you call `/v1/chat/completions`, the server applies the template for you: it takes your `messages` array, renders ChatML, appends the generation prompt, and strips the trailing `<|im_end|>` from the output before returning `message.content`. That is the entire value of the chat endpoint. When you call `/v1/completions` you get none of that. You send a string, the model continues it. If that string is plain prose, the failure modes are recognisable: - **No stopping.** The model answers and then keeps writing, often inventing the next user question, because nothing produced the `<|im_end|>` boundary the sampler stops on. - **Role bleed.** Literal `<|im_start|>assistant` or `user` text appears in the output. - **Ignored system prompt.** Instructions placed as prose at the top carry far less weight than a properly delimited system turn. - **Degraded instruction following in general**, because the input does not resemble the fine-tuning distribution. ## Tools and extended turns Qwen's template does more than three roles. When you pass a `tools` list, the template renders the tool schemas into the system turn, and it defines how an assistant's function call is serialised and how a `role: "tool"` result is rendered back into the conversation using its own tag syntax. Reasoning-capable Qwen releases additionally use a thinking segment inside the assistant turn, and the template exposes a toggle for whether that segment is opened. All of this is why replacing the shipped template with a generic one silently removes capabilities: the model still emits the syntax it was trained on, but your prompt no longer sets it up and your server no longer parses it. ## Practical rules Use the chat endpoint and let the server template for you. If you must build prompts yourself — for batch scoring, evaluation harnesses, or fine-tuning data — render them with `apply_chat_template` from the exact checkpoint you serve, never from a hand-copied string, because templates differ between model generations and between base and instruct variants. And when output looks unhinged on a self-hosted model that behaves fine on a hosted endpoint, check the rendered prompt before you touch sampling parameters.

  • What is the practical symptom of forgetting the generation prompt when you build the string yourself?
    The model sees what looks like a finished conversation, so instead of answering it commonly starts a fresh turn — emitting a new user message, restating the question, or producing a role marker as literal text. Setting `add_generation_prompt=True` appends the open assistant header, which is the cue that the next tokens are the assistant's reply.
  • Why can a base Qwen checkpoint behave badly with the same ChatML prompt that works on the instruct variant?
    Base checkpoints are pretrained on text and never instruction-tuned on ChatML, so the special turn tokens carry no learned meaning and there is no trained habit of stopping at the end of a turn. Some base repos do not ship a chat template at all. Base weights are for continued pretraining or your own fine-tune, not for chat serving.
  • You need a custom system-prompt policy on every request. Do you edit the chat template?
    No — pass it as a `system` message in the `messages` array, which the shipped template already renders correctly. Editing the template forks it from the checkpoint and risks losing tool and reasoning rendering. Template overrides are for a missing template or a documented feature gap, and then you start from the model's own template rather than a generic one.

saying these in an interview costs you the question

  • Calling the ChatML markers ordinary text the model reads literally
  • Building prompts as plain prose and blaming the model for rambling
  • Reusing a Llama or Mistral prompt format for Qwen
  • Assuming any chat template works as long as roles are labelled
  • Dropping the trailing assistant header and expecting an answer

context