skip to content

In vLLM, what breaks at /v1/chat/completions when a model ships no chat template?

level: middleimportance: must knowfreq 62%

answer

  1. Messages become one flat token sequence
  2. Jinja, shipped with the tokenizer
  3. Base checkpoints have no roles
  4. Server errors rather than guessing
  5. --chat-template overrides tokenizer_config.json

basics

~20 s

Chat requests fail outright. vLLM renders the messages array into a single prompt using the Jinja chat template stored in the model's tokenizer_config.json; a base checkpoint has none, so you must supply one with --chat-template or use /v1/completions instead.

solid answer

~50 s

The model does not see a `messages` array — it sees one flat token sequence. The chat template is a Jinja template shipped in the model's `tokenizer_config.json` that turns roles and contents into that sequence, inserting the exact special tokens the model was instruction-tuned on and appending the generation prompt that cues the assistant turn. vLLM applies it server-side on every `/v1/chat/completions` call. Base or completion-only checkpoints ship no template, so vLLM cannot render the request and returns an error; `/v1/completions`, which takes a raw `prompt`, still works. The fix is `--chat-template` pointing at a Jinja file (or the template inline). The dangerous case is not the error but the near-miss: a template that renders but uses the wrong role markers produces no error at all, just quality that quietly degrades and tool calls that stop parsing. Post the same `messages` to `/tokenize` to see exactly what the server built.

code

bash · 3 lines
bash
vllm serve mistralai/Mistral-7B-v0.1 \
  --chat-template /srv/templates/mistral-instruct.jinja \
  --served-model-name base-with-template

go deeper

for a junior

Know that a chat template turns the messages list into one prompt string using the model's own special tokens, and that it normally arrives with the model rather than being something you write.

for a middle

Explain where the template lives (the tokenizer config), what the generation prompt does, why a base checkpoint fails the chat route, and that --chat-template is the override. Mention that /v1/completions bypasses the whole issue.

for a senior

Demonstrate the diagnosis: a wrong-but-valid template returns 200 and silently degrades quality, so verify the rendered prompt with /tokenize and treat any overridden .jinja file as versioned production config.

for a principal

Frame the template as part of the prompt contract across teams: pin it with the checkpoint, review changes to it like code, and decide whether overrides live with the model artifact or the deployment so a model swap cannot silently change behaviour.

## Why a template exists at all A language model consumes a single sequence of token ids. The `messages` array in a chat request is an API convenience; something must flatten it into that sequence, and it must do so exactly the way the model was fine-tuned, because instruction tuning teaches the model specific delimiter tokens for "here begins a system instruction", "here begins the user turn", "now you speak". Those markers differ per model family and are not interchangeable. The convention the ecosystem settled on is the **chat template**: a Jinja template stored under the `chat_template` key in the model's `tokenizer_config.json` (some repositories ship it as a separate template file). It iterates the messages, emits the model's own role markers, and — because the request wants a continuation, not a completed transcript — appends the *generation prompt*, the opening marker of an assistant turn, so decoding starts in the right place. vLLM applies this template inside the server on every `/v1/chat/completions` request. Nothing about it is vLLM-specific magic: it is the tokenizer's template, run by the server, which is precisely why serving a model works out of the box when the repository is well-formed. ## The failure when there is no template Base checkpoints — the pretrained, non-instruction-tuned models — have no chat template, because they have no notion of roles. Neither do some completion-only or older releases. Post to `/v1/chat/completions` against such a model and vLLM has nothing to render with, so the request fails with an error rather than guessing. That refusal is deliberate: silently concatenating message contents would produce a prompt shaped like nothing the model has ever seen, and the resulting garbage would be much harder to diagnose than a startup-time error. Two ways out. If you genuinely want raw continuation behaviour, use `/v1/completions`, which takes a `prompt` string and needs no template. If you want the chat surface, pass `--chat-template` at launch with a path to a Jinja file (or the template text itself); vLLM then uses that instead of looking in the tokenizer config. The same flag is how you override a template that ships broken, which does happen with fresh releases. ## The quieter, worse failure A missing template is an error you cannot ignore. A *wrong* template is not. If you attach a template built for a different model family, every request succeeds — the server renders, the model generates, you get a 200. What you lose is fidelity: the model sees delimiters it was not tuned on, so it drifts off-format, ignores the system message, keeps talking past the stop point, or refuses to emit tool calls in the shape the parser expects. Symptoms look like "the model is worse than the benchmark said", and teams often chase sampling parameters for days before checking the prompt that actually reached the GPU. The direct check is `POST /tokenize`. It accepts the same `messages` payload and returns the token ids the server would feed the engine, so you can confirm the rendered prompt contains the delimiters you expect. Do that once when onboarding a new checkpoint and you will never spend that week. ## Two knobs worth knowing `--chat-template-content-format` controls how `content` is handed to the template. Modern requests may send `content` as a list of typed parts (text, image) rather than a plain string; templates differ in which form they expect. The setting accepts `auto`, `string` and `openai`, letting you force flattening to a plain string or preserve the structured list when the template iterates parts. Getting this wrong on a multimodal model shows up as images being dropped from the prompt. Requests can also carry `chat_template_kwargs`, a dictionary passed into the template's rendering context. Templates for models that support an optional thinking block, or an optional tool-preamble, read flags from there, so this is the per-request hook for behaviour the template author made configurable. It is a passthrough: what the keys mean is entirely a property of the model's template, not of vLLM. ## Practical rules Prefer instruction-tuned checkpoints whose repository ships a template — that is the path with no configuration. When you must override, keep the `.jinja` file in version control next to your deployment manifest, because it is now part of your prompt contract; a change to it is a behaviour change with no code diff anywhere else. And whenever tool calling misbehaves, check the template first: a template that never renders the `tools` argument means the model was never told the tools exist, no matter what the request contained.

  • How would you confirm what prompt the server actually built from a messages array?
    POST the same `messages` payload to the server's `/tokenize` endpoint. It runs the chat template and returns the token ids the engine would receive, so you can verify the role delimiters and the generation prompt are present. It is a cheap, GPU-free check and the fastest way to catch a wrong or overridden template.
  • If a wrong template still returns 200 responses, what symptoms tell you it is wrong?
    Output that ignores the system message, rambles past where it should stop, opens with an echoed role marker in the visible text, or emits tool calls as prose so `tool_calls` stays empty. Quality looks broadly worse than the model's published benchmarks. Compare the rendered prompt against the model card's documented format to confirm.
  • What is chat_template_kwargs in a request body for?
    It is a dictionary passed into the template's rendering context, so per-request flags that the template author exposed — for example an optional thinking block or a preamble toggle — can be set by the caller. The meaning of every key belongs to that model's template; vLLM only forwards it.

saying these in an interview costs you the question

  • Thinks vLLM invents a generic template for any model
  • Believes the model receives the messages array directly
  • Assumes templates are interchangeable across model families
  • Blames sampling settings when the template is wrong
  • Forgets /v1/completions needs no template at all

context