skip to content

How does Llama 3's chat template differ from Llama 2's [INST] and <<SYS>> format?

level: middleimportance: must knowfreq 68%

answer

  1. one format per model generation
  2. plain-text markers versus real vocab tokens
  3. system prompt nested versus its own role
  4. role headers closed by an end-of-turn token
  5. <<SYS>> inside [INST] vs <|start_header_id|>

basics

~20 s

Llama 2 wraps each user turn in [INST] ... [/INST] plain-text markers and nests the system prompt inside the first one with <<SYS>>. Llama 3 drops both, using real special tokens: a role header per turn, closed by <|eot_id|>.

solid answer

~40 s

Llama 2 Chat renders as `<s>[INST] <<SYS>>\n{system}\n<</SYS>>\n\n{user} [/INST] {assistant} </s>`, and each further exchange starts a fresh `<s>...</s>` pair. `[INST]`, `[/INST]`, `<<SYS>>` and `<</SYS>>` are ordinary text that tokenizes like any other string; only `<s>`/`</s>` are special tokens. Llama 3 Instruct replaces that with a header grammar: one `<|begin_of_text|>` at the very start, then per turn `<|start_header_id|>{role}<|end_header_id|>` followed by two newlines, the content, then `<|eot_id|>`. All four markers are genuine entries in the 128k vocabulary. The system prompt becomes a normal `system` turn instead of being nested inside the first user message, and BOS appears once for the whole conversation rather than once per exchange. To sample a reply you append the assistant header and stop generating at `<|eot_id|>`.

code

python · 8 lines
python
LLAMA2_CHAT = (
    "<s>[INST] <<SYS>>\n"
    "You are a helpful assistant.\n"
    "<</SYS>>\n\n"
    "What is 2+2? [/INST] 4 </s>"
    "<s>[INST] And 3+3? [/INST]"
)
print(LLAMA2_CHAT)

go deeper

for a junior

Know that open-weight models need the prompt formatted in a specific way, that Llama 2 and Llama 3 use different markers, and that you should get the format from the tokenizer rather than typing it yourself.

for a middle

Be ready to write both layouts from memory: [INST]/<<SYS>> with per-exchange <s>...</s> for Llama 2, and <|begin_of_text|> plus role headers closed by <|eot_id|> for Llama 3, including where the system prompt sits in each.

for a senior

Explain the failure mode: a wrong template never errors, it just degrades quality, so you need a rendered-prompt log or golden-string test. Mention the injection angle of plain-text markers and the single-BOS rule.

for a principal

Own the policy that prompt rendering is a versioned artifact tied to the checkpoint, not code duplicated across services. Argue for one rendering path, contract tests on the rendered bytes, and a model-swap checklist that treats the template as part of the model interface.

## Why the layout is your problem at all When you call a hosted API you hand over a list of `{role, content}` objects and the provider assembles the prompt. With open weights there is no such layer: the model is a next-token predictor over one flat token stream, and instruction tuning taught it exactly one way that a conversation looks. Feed it a different layout and nothing errors — you simply land off-distribution and the answers get worse in ways that look like "the model is dumb" rather than "my prompt is malformed". That silent failure mode is why this is asked in interviews. ## The Llama 2 Chat layout ``` <s>[INST] <<SYS>> You are a helpful assistant. <</SYS>> What is 2+2? [/INST] 4 </s><s>[INST] And 3+3? [/INST] ``` Properties worth knowing: - `[INST]` / `[/INST]` and `<<SYS>>` / `<</SYS>>` are **plain ASCII text**, not special token IDs. They occupy several ordinary tokens each. - `<s>` (BOS) and `</s>` (EOS) *are* special tokens in the Llama 2 SentencePiece vocabulary. Each user/assistant exchange is wrapped in its own `<s> ... </s>` pair, so a five-turn chat contains five BOS tokens. - The system prompt has no slot of its own. It is spliced **inside the first `[INST]` block**, between `<<SYS>>` and `<</SYS>>`, with a blank line before the user text. - Because the markers are plain text, untrusted user input containing the literal string `[/INST]` is indistinguishable from a real turn boundary — a genuine prompt-injection surface. ## The Llama 3 Instruct layout ``` <|begin_of_text|><|start_header_id|>system<|end_header_id|> You are a helpful assistant.<|eot_id|><|start_header_id|>user<|end_header_id|> What is 2+2?<|eot_id|><|start_header_id|>assistant<|end_header_id|> ``` Properties: - `<|begin_of_text|>`, `<|start_header_id|>`, `<|end_header_id|>`, `<|eot_id|>` and `<|end_of_text|>` are **reserved special tokens** in the 128k vocabulary — one token each, not spelled-out text. - A turn is `header + \n\n + content + <|eot_id|>`. The two newlines after `<|end_header_id|>` are part of the format; dropping them is a real (if subtle) deviation. - `<|begin_of_text|>` appears **once**, at the very start of the whole conversation — not per exchange as in Llama 2. - `system` is just another role header, alongside `user` and `assistant`. It is no longer nested inside a user turn, so it cannot be accidentally terminated by user content and does not have to be re-emitted per exchange. - To generate, you append `<|start_header_id|>assistant<|end_header_id|>\n\n` and let the model write until it emits `<|eot_id|>`. ## What actually changed, conceptually Role became a first-class, addressable slot rather than an in-band text convention. That buys three things: the system prompt stops competing with the first user message for the same block; conversation structure survives user text that happens to contain the marker strings; and turn boundaries cost one token instead of several. It does not make you automatically safe. Hugging Face tokenizers parse special-token *strings* found in raw input back into their token IDs by default, so if you concatenate untrusted text yourself, a user who types `<|eot_id|>` can still forge a boundary. Render through the template and tokenize untrusted content with `split_special_tokens=True` (or strip the markers) if you are stitching strings by hand. ## Version notes Every Llama 3 generation — 3, 3.1, 3.2, 3.3 — shares the header grammar. Llama 3.1 added `<|eom_id|>` and an `ipython` role for tool calling; Llama 3.2's vision variants add an `<|image|>` token. The Llama 2 `[INST]` layout applies only to the Llama 2 chat models: applying it to a Llama 3 checkpoint (or vice versa) is exactly the quiet-degradation bug described above. Code Llama's instruct variants inherit the Llama 2 `[INST]` style. ## The practical rule Never hand-write either layout in production code. Call the tokenizer's chat template so the format travels with the checkpoint, and log the rendered string once at startup so you can eyeball the boundaries. If you must hardcode — for a C++ or Rust serving path with no Python tokenizer — copy the reference layout token for token, including the blank lines, and diff it against the template output before shipping.

  • Why is it risky that Llama 2's [INST] markers are plain text rather than special tokens?
    Because user content containing the literal string `[/INST]` or `<<SYS>>` is byte-identical to a real boundary, so untrusted input can forge a turn break or open a fake system block. Llama 3's markers are reserved vocabulary entries, which removes the ambiguity — provided you render through the template and do not let raw user text be tokenized with special-token parsing enabled.
  • In Llama 3's format, how many times does the BOS token appear in a ten-turn conversation?
    Once. `<|begin_of_text|>` opens the whole token stream and turns are separated by `<|eot_id|>` plus the next role header. Llama 2 is the opposite: each `[INST] ... [/INST] answer </s>` exchange is wrapped in its own `<s> ... </s>`, so BOS repeats per exchange. Emitting a BOS per turn on Llama 3 is a formatting bug.
  • What symptom would you expect if you served a Llama 3 checkpoint using the Llama 2 [INST] layout?
    No error at all — the model sees `[INST]` as ordinary text. You would see degraded instruction following: ignored or weakly-honoured system prompts, chattier and less aligned answers, occasional echoing of the markers, and unreliable stopping because no `<|eot_id|>` boundary was ever established. It typically reads as "this model is worse than the benchmarks claim".

saying these in an interview costs you the question

  • Thinks [INST] is a special token ID rather than plain text
  • Says Llama 3 still uses <<SYS>> for the system prompt
  • Claims a wrong template raises an error instead of degrading quietly
  • Repeats <|begin_of_text|> at the start of every turn
  • Assumes one universal chat format works across all open models

context