skip to content

Which hidden tokens does an LLM chat request add beyond your message text?

level: middleimportance: should knowfreq 42%

answer

  1. the request is one rendered sequence
  2. structure is billed like content
  3. overhead repeats per message
  4. tool schemas ride along every call
  5. apply the template, then tokenize

basics

~20 s

Chat models are fed one rendered sequence, not a list of strings. A template wraps every message in role and turn-delimiter tokens and prepends the system prompt and tool schemas, all billed as input and invisible to a plain character count.

solid answer

~50 s

An instruction-tuned chat model expects a specific serialization it was trained on: each message is surrounded by special tokens marking who is speaking and where the turn begins and ends, and the whole conversation is preceded by the system prompt and any tool definitions. That framing is a **fixed overhead per message**, typically a handful of tokens, plus a one-off prefix. Two practical effects follow. First, a naive count over the concatenated message strings always undercounts, and the error grows with the number of turns — a hundred short exchanges pay the overhead a hundred times, so chatty turn structures cost more than the same words in one message. Second, tool and function schemas are prompt text re-sent every call, often the largest hidden line item. To count correctly, render the request through the same template the model uses, then tokenize; the usage figures the response reports are the ground truth to reconcile against.

go deeper

for a junior

Know that the system prompt and the conversation structure are billed too, so the tokens you pay for are always more than the text you typed.

for a middle

Explain that the request is rendered through a chat template adding role and turn-delimiter tokens per message, that tool schemas are re-sent every call, and that this is why local counts undershoot reported usage.

for a senior

Demonstrate that you budget and trim against the templated count, watch the estimate-versus-usage delta as a regression signal, and treat verbose tool descriptions as a recurring per-request cost worth engineering down.

for a principal

Own the structural choices: how many tools stay loaded, how instruction is consolidated, and how history is summarized — decisions whose token consequences compound across every turn of every session at fleet scale.

## The model does not see a list of messages An API that takes an array of role-tagged messages is a convenience over something flatter. Under it, the conversation is serialized into one token sequence using a **chat template** — a fixed format the model was instruction-tuned on. The template inserts special tokens the tokenizer reserves for structure rather than content: a marker that a turn begins, an indication of the role, and a marker that the turn ends. Those special tokens exist precisely so the model can tell your instructions from the user's text and from its own prior replies, which is also why prompt-injection defence begins with never letting untrusted text forge them. ## The overhead is per message, and it compounds Roughly, the rendered request is: a prefix (system prompt, tool definitions, any provider preamble), then for each message a small constant of template tokens plus that message's content, and finally an opening marker inviting the assistant to speak. The constant is small — a handful of tokens — but it is charged once per message. One 2,000-character message pays it once. Two hundred ten-character messages pay it two hundred times, and the structure can then outweigh the content. This is the direct explanation for the most common support question in LLM apps: *why does the provider report more input tokens than I counted?* You counted content; you were billed content plus structure. ## Tool schemas are the biggest hidden item Tool and function definitions are not metadata. Their names, human-readable descriptions, parameter names, enums and nested JSON Schema are rendered into the prompt as text and re-sent on **every** request in the conversation, because the model is stateless across calls. A dozen well-documented tools can occupy thousands of input tokens before the user has typed a character, paid on every turn. Trimming verbose descriptions or deferring rarely used tools is one of the highest-leverage cost reductions available, and it is invisible to anyone who budgets by message length alone. ## Why your local count can disagree with the provider's Three causes, in order of frequency. You tokenized raw strings instead of the templated sequence, so you missed the per-message and prefix overhead. You used a tokenizer from a different model family, so the vocabulary differs. Or the provider adds something you did not author — a system preamble, a tool-use scaffold, an injected date. The fix is the same in all three cases: render through the correct template for the deployed model, tokenize with the matching tokenizer, and then reconcile against the usage numbers in the response. Treat the reported usage as ground truth and your estimator as an approximation whose error you monitor. ## Consequences for design Budgeting. If a trimming policy assumes the window holds the sum of the message lengths, it will overflow on long, chatty threads exactly when history is most valuable. Budget against the templated count. Turn granularity. Splitting an instruction into many small messages costs more than the same words in one, with no benefit to the model — the template overhead is real tokens. This argues for consolidating system-level instruction into one system message rather than sprinkling it across turns. History management. Because history is re-sent whole on each call, per-turn overhead accumulates quadratically over a long conversation: turn *n* pays for all *n* prior turns and their markers. Summarizing or dropping old turns is a token decision, not only a quality one. Structure integrity. Special tokens are reserved. Text that merely looks like a role marker is normally encoded as ordinary content by the tokenizer, which is the point — but any pipeline that concatenates untrusted text into the raw templated string, rather than passing it as message content, risks letting that text impersonate a role boundary. ## What to say in an interview State plainly that the request is a rendered sequence, not a list; that role and turn-delimiter tokens plus the system prompt and tool schemas are billed input; that overhead is per message so many short turns cost more; and that the way to count is to apply the template, then tokenize, then reconcile with the reported usage. That answer covers the mechanism, the cost consequence and the operational check.

  • Why does a conversation of many short turns cost more than the same words in one message?
    Because the template overhead is charged per message, not per conversation. Each turn adds its own begin, role and end markers, so a hundred ten-character messages pay that constant a hundred times while one long message pays it once. On short exchanges the structural tokens can approach or exceed the content tokens, which is why consolidating instructions into a single system message is cheaper for identical text.
  • Your local estimate is consistently a few percent below the provider's reported input tokens. Where do you look first?
    First, check that you tokenized the templated request rather than the concatenated strings, since per-message markers and the prefix are the usual gap. Second, confirm the tokenizer matches the deployed model family. Third, account for anything the provider injects that you did not author, such as a preamble or tool scaffold. Then keep logging estimate against reported usage so the residual error is monitored rather than rediscovered.
  • How does per-turn overhead behave as a conversation grows long?
    It accumulates. The model is stateless across calls, so each request re-sends the whole history, meaning turn n pays for all previous turns plus their markers — total tokens across a session grow roughly quadratically in turn count. That is why dropping or summarizing old turns is a cost decision as much as a quality one, and why per-turn overhead that looks negligible early becomes material late.

saying these in an interview costs you the question

  • Believes only the user's text is billed
  • Counts message strings and expects it to match reported usage
  • Thinks tool schemas are metadata rather than prompt text
  • Assumes the API stores history so it is not re-sent
  • Says splitting one instruction across turns costs the same

context