skip to content

In Llama 3.1's chat format, what do <|eom_id|> and the ipython role mean?

level: seniorimportance: nice to knowfreq 26%

answer

  1. a turn can pause mid-way
  2. two terminators, not one
  3. who speaks after the assistant
  4. the role that carries tool output
  5. end-of-message versus end-of-turn

basics

~20 s

<|eom_id|> ends an assistant message that is waiting on a tool result instead of handing the turn back to the user. The tool's output is fed back as a message with the ipython role, and the model then continues.

solid answer

~50 s

Llama 3.1 extended the header format for tool use. Two additions matter. `<|eom_id|>` — end of message — is emitted instead of `<|eot_id|>` when the assistant has produced a tool call and expects execution output before it can finish. `<|eot_id|>` still means "turn complete, user speaks next". A tool loop must stop on both but react differently: on `<|eom_id|>` you execute and continue the same turn, on `<|eot_id|>` you return the answer. The execution result comes back as a message whose role is `ipython`, alongside `system`, `user` and `assistant`. For Meta's built-in tools the model prefixes the call with `<|python_tag|>`; for custom tools declared in the system prompt it typically emits a JSON object and closes with `<|eot_id|>`. Whether you see end-of-message or end-of-turn therefore depends on how the tools were declared, which is why the loop must handle both.

code

python · 8 lines
python
TOOL_TURN = (
    "<|start_header_id|>assistant<|end_header_id|>\n\n"
    '<|python_tag|>get_weather(city="Berlin")<|eom_id|>'
    "<|start_header_id|>ipython<|end_header_id|>\n\n"
    '{"temperature": 12, "unit": "C"}<|eot_id|>'
    "<|start_header_id|>assistant<|end_header_id|>\n\n"
)
print(TOOL_TURN)

go deeper

for a junior

Recognise that Llama 3.1 added tool-calling markers to the chat format, and that a tool's output is fed back as another message rather than through a separate API field.

for a middle

Explain the split between end-of-message and end-of-turn, and name the ipython role as the carrier of tool results in the token stream.

for a senior

Describe the loop you would actually write: stop on both terminators, distinguish a built-in call prefixed with the python tag from a JSON custom-function call, execute, splice the result back as an ipython turn, and continue.

for a principal

Weigh owning this raw protocol against adopting a serving layer that parses tool calls for you — the tradeoff is portability and debuggability versus the maintenance cost of tracking format changes across model generations.

## The extension, in one paragraph Llama 3.0 shipped three roles — `system`, `user`, `assistant` — and one turn terminator, `<|eot_id|>`. Llama 3.1 added the pieces needed to express a conversation that pauses mid-turn to run code: a fourth role, `ipython`, and a second terminator, `<|eom_id|>`. Everything else about the format is unchanged: still one `<|begin_of_text|>`, still `<|start_header_id|>{role}<|end_header_id|>` followed by a blank line. ## End of message versus end of turn Think of a turn as "the assistant's whole contribution before control returns to the user", and a message as one contiguous chunk within it. - `<|eot_id|>` — the assistant is done; the next thing in the stream is a user (or system) header. - `<|eom_id|>` — the assistant has said all it can *for now* and needs an external result; the next thing in the stream is an `ipython` header carrying that result, after which the assistant speaks again. This is a token-level encoding of the agent loop. Without it, a client could not tell "finished" from "blocked on a tool" except by parsing the content. ## The ipython role Tool output is injected as a message with role `ipython`: ``` <|start_header_id|>ipython<|end_header_id|> {"temperature": 12, "unit": "C"}<|eot_id|><|start_header_id|>assistant<|end_header_id|> ``` The name is historical — it reflects the code-execution framing of the built-in tools — but it is used for any tool result, not only Python. If you are used to other vendors' shapes, note that this is where Llama diverges: there is no `tool` role and no per-call correlation ID in the base format, so the association between a call and its result is positional. That has a practical consequence for parallel tool calls: you cannot rely on IDs to match results back, so keep the ordering strict. ## Built-in tools and the python tag Meta's reference format distinguishes built-in tools (enabled by putting `Environment: ipython` in the system prompt, which unlocks helpers such as a code interpreter and search) from custom functions you declare yourself in the system prompt as JSON schemas. - Built-in path: the assistant emits `<|python_tag|>` followed by the call, and closes with `<|eom_id|>`. - Custom-function path: the assistant emits a JSON object naming the function and its arguments, and typically closes with `<|eot_id|>`. So a loop that only stops on `<|eom_id|>` will hang on custom tools, and a loop that only stops on `<|eot_id|>` will treat a built-in call as a finished answer and show the raw call to the user. Handle both terminators, then inspect the content to decide whether it is a call or a reply. ## Why this shows up in interviews It is the point where "open weights mean you own the protocol" becomes concrete. With a hosted API the vendor parses tool calls for you and hands back a structured object. With Llama you are looking at raw tokens and deciding, in your own code, what counts as a call, when to execute, and how to splice the result back into the token stream. Serving stacks vary in how much of this they do for you — some parse tool calls into a structured field, some hand you the raw text — so a candidate who knows the underlying token grammar can debug across all of them. ## Version caveat All of this is Llama 3.1 and later. A Llama 3.0 checkpoint has no `ipython` role and no `<|eom_id|>` in its trained format; asking it to use them is just off-distribution text. And the reserved special tokens exist in the vocabulary of 3.x models generally, so their presence in the tokenizer is not evidence that a given checkpoint was tuned to use them — check the model card, and check the chat template that ships with the weights.

  • What breaks if your loop only stops on <|eom_id|> and ignores <|eot_id|>?
    Custom function calls typically close with <|eot_id|>, not <|eom_id|>, so the loop never sees its stop token and generation runs to the length cap — often producing an invented tool result and a fabricated follow-up turn. Handle both terminators and decide by inspecting the content whether the assistant produced a call or a final answer.
  • How does Llama's tool-result shape differ from a vendor API's tool_call_id convention?
    Llama's base format has no correlation identifier: the result is injected as an ipython-role message and the association with its call is positional. Vendors that mint a call ID let you return results out of order and match them explicitly. With Llama you must preserve ordering yourself, which matters as soon as you allow more than one call before resuming.
  • You see <|eom_id|> and <|python_tag|> in a checkpoint's tokenizer. Does that mean it supports tool calling?
    No. Those IDs are reserved across the Llama 3 vocabulary, so their presence proves nothing about training. What matters is whether the checkpoint was tuned for tool use — read the model card and the chat template shipped with the weights. A Llama 3.0 model has the tokens in vocabulary but no trained behaviour behind them.

saying these in an interview costs you the question

  • Says <|eom_id|> and <|eot_id|> are interchangeable
  • Expects a tool role because other vendors use one
  • Assumes every tool call ends with <|eom_id|>
  • Believes the tokens existing in the vocabulary implies trained tool use
  • Applies the 3.1 tool format to a Llama 3.0 checkpoint

context