skip to content

How do you count the tokens an LLM prompt will use before sending it?

level: juniorimportance: must knowfreq 68%

answer

  1. price, latency, and fit
  2. estimate before you dispatch
  3. tokenizer belongs to the model
  4. count the whole rendered request
  5. four characters per token is prose-only

basics

~20 s

Run the target model's own tokenizer over the fully assembled request, or call a provider token-counting endpoint when no offline tokenizer exists. Character heuristics such as roughly four characters per token are planning guardrails, not exact counts.

solid answer

~50 s

Token counts decide three things before a call: **price**, because input and output are billed per token; **latency**, which tracks how much text is prefilled and how much is generated; and **fit**, because the input plus the output you reserve must sit inside the context window. Exact counting means running the tokenizer that belongs to that specific model over the fully rendered request — system prompt, every message with its chat-template markers, and any tool schemas — not just the user's text. Some providers ship an offline tokenizer for their open models; others expose a counting endpoint you call first. Heuristics like ~4 characters per token, or ~0.75 words per token, hold only for English prose and belong in capacity planning, never in a hard fit decision. Reconcile estimates against the token usage each response reports, and re-measure whenever you switch model families, since vocabularies differ.

go deeper

for a junior

Know that billing and context limits are measured in tokens, not characters, and that the way to get a real number is to run the model's own tokenizer rather than dividing the character count by four.

for a middle

Be ready to explain what the ~4-characters-per-token rule actually approximates, where it breaks, and why the count must be taken over the fully rendered request including the system prompt and tool schemas.

for a senior

Show that you instrument it: pre-flight estimate logged against reported usage, headroom reserved for output, worst-case content exercised in CI, and a model swap treated as a re-measurement rather than a config change.

for a principal

Own the economics — how token accounting feeds unit-cost models, which budget you enforce where (per request, per session, per tenant), and how you keep a cost model honest as prompts, toolsets and model choices change underneath it.

## Why anyone counts tokens A token count is not trivia; it is the unit three separate budgets are denominated in. Providers bill per token, with input and output priced differently. Latency scales with tokens: the prefill phase processes the input, then each output token is generated in sequence, so a prompt twice as long costs roughly twice the prefill work. And every model has a hard context limit expressed in tokens — exceed it and the request is rejected or history is silently dropped. Estimating tokens before dispatch is therefore ordinary engineering work, the same way you would size a payload before a network call. ## A token is not a character, a word, or a byte A token is an entry in the vocabulary that a particular tokenizer learned. Common English words usually map to a single token, leading whitespace normally attaches to the following word, and rarer words break into pieces. The crucial consequence is that a token count is a property of a *(text, tokenizer)* pair, not of the text alone. The same paragraph can differ by tens of percent between two model families, and even between generations of one family, because vocabularies are retrained. ## Exact counting: two routes The first route is an offline tokenizer: the provider or the open-weight release publishes the vocabulary and the merge rules, and you tokenize locally at zero marginal cost. The second is a counting endpoint: you send the assembled request to the provider and get back the number of input tokens it would consume. Offline is cheaper and works on the hot path; an endpoint costs a round trip but is authoritative for closed models whose tokenizer is not published. Either way, count the *rendered* request. Chat models are not fed a list of strings; they are fed one sequence produced by applying a chat template that wraps each message in role and turn-delimiter tokens and prepends the system prompt and tool definitions. Tokenizing the raw concatenated message strings undercounts by a fixed amount per message — a small error on one long message, a large one on a hundred short ones. ## Heuristics and their honest range The ~4-characters-per-token rule (equivalently ~0.75 words per token) is a reasonable prior for English prose and nothing else. Observed ratios in practice: ordinary English prose lands near 3.5–5 characters per token; source code with punctuation and indentation lands near 2.5–3.5; minified code, base64 blobs, UUIDs and hex hashes fall toward 1–2 because random character sequences share no learned merges; and non-Latin scripts such as Japanese or Thai often approach one token per character. A per-conversation cost model built on the flat 4x rule will be wrong by a multiple on some traffic. When you genuinely cannot tokenize on the hot path, do not use the generic constant. Sample real traffic, measure a characters-per-token ratio *per locale and per content type*, take a pessimistic percentile rather than the mean, add a fixed per-message overhead, and validate the estimator against the usage numbers the API reports. ## What must be inside the count Everything the model reads is input: the system prompt, every retained turn of history, retrieved documents pasted into context, tool and function schemas (descriptions and JSON Schema are prompt text and are re-sent on every call), and the template markers. Tool schemas are the most commonly forgotten line item — a large toolset can occupy thousands of tokens on every request before the user has typed anything. ## Output shares the budget Input and output normally compete for the same window, so the number of output tokens you allow must be subtracted from what the input may use; many APIs reject a request whose input plus requested maximum output exceeds the limit. On models that expose a reasoning-depth dial, the hidden reasoning tokens are generated and billed as output too, so a high effort setting silently enlarges the output side of the budget. ## Operational practice Log the reported input and output usage for every call alongside your pre-flight estimate. The delta is a cheap regression signal: it jumps the moment someone adds a tool, expands the system prompt, or switches model. Re-run your estimator over a worst-case corpus in CI, and treat a model swap as a re-measurement event rather than a drop-in. ## Common mistakes Assuming one token equals one word; reusing another family's tokenizer because it was handy; counting only the user message; forgetting the output reservation; and treating the character heuristic as a guarantee rather than a guardrail.

  • Why can the identical string count differently on two models from the same provider?
    Because each model family is trained with its own tokenizer. Vocabulary size, merge rules and the special tokens in the chat template all differ, so the same characters segment differently. Counts are a property of the text plus the tokenizer, never of the text alone, which is why a model swap requires re-measuring cost and fit rather than assuming the old numbers carry over.
  • You cannot afford a tokenizer call on the hot path. How do you keep a safe margin?
    Calibrate offline: sample real traffic, measure characters per token separately per locale and content type, and pick a pessimistic percentile rather than the mean. Add a fixed per-message overhead for the chat template, subtract the reserved output tokens, and leave explicit headroom. Then log the estimate next to the usage the API reports and alert when the ratio drifts, which is what catches a new tool schema or a model change.
  • What does a large tool or function schema do to your input budget?
    It is prompt text, re-sent on every request in the conversation. Names, descriptions, enums and nested JSON Schema all tokenize, so a broad toolset can consume thousands of input tokens before the user says anything — paid on every turn, and subtracted from the window available for history and retrieved context. It is one of the most commonly missed line items in a token budget.

saying these in an interview costs you the question

  • Says one token is roughly one word, universally
  • Applies the same characters-per-token ratio to every language
  • Counts only the user message, ignoring system prompt and tool schemas
  • Assumes one provider's tokenizer works for every model
  • Thinks output tokens do not consume the context window

context