skip to content

How do you count an OpenAI request's tokens with tiktoken before sending?

level: middleimportance: should knowfreq 52%

answer

  1. Local estimate, not the authoritative number
  2. Model name picks the encoding
  3. Messages carry framing beyond their text
  4. Tool schemas and images are invisible to it
  5. Output length cannot be known in advance

basics

~20 s

Get the encoding for your model with tiktoken.encoding_for_model, then take len(encoding.encode(text)) for each message. Add the per-message framing overhead, because raw text encoding alone undercounts a chat request and ignores tool schemas and images entirely.

solid answer

~40 s

`tiktoken` is OpenAI's tokenizer library, and the safe entry point is `tiktoken.encoding_for_model("gpt-4o")`, which picks the right encoding rather than making you remember that the gpt-4o generation uses `o200k_base` while gpt-4 and gpt-3.5-turbo used `cl100k_base`. Counting a single string is `len(encoding.encode(text))`. Counting a *chat request* is more than that: the API wraps each message in role framing worth roughly three extra tokens per message plus a few for priming the reply, so summing bare content undercounts. Tool and function JSON schemas are serialised into the prompt and billed too, and image inputs are charged by a tile-based formula that `tiktoken` cannot compute at all. So treat a local count as a good pre-flight estimate for truncation and budget checks, and reconcile against `usage.prompt_tokens` from the real response for anything financial.

code

python · 28 lines
python
import tiktoken


def encoding_for(model: str) -> tiktoken.Encoding:
    try:
        return tiktoken.encoding_for_model(model)
    except KeyError:
        # Model newer than this tiktoken release; fall back deliberately.
        return tiktoken.get_encoding("o200k_base")


def count_chat_tokens(messages: list[dict], model: str = "gpt-4o") -> int:
    enc = encoding_for(model)
    tokens = 0
    for message in messages:
        tokens += 3  # per-message role and separator framing
        for key, value in message.items():
            tokens += len(enc.encode(value))
            if key == "name":
                tokens += 1
    tokens += 3  # priming for the assistant reply
    return tokens


print(count_chat_tokens([
    {"role": "system", "content": "You are terse."},
    {"role": "user", "content": "Summarise the CAP theorem."},
]))

go deeper

for a junior

Know that tiktoken turns text into tokens, that encoding_for_model picks the right one for your model, and that len(encoding.encode(text)) is the basic count.

for a middle

Explain why a chat request costs more than the sum of its message texts — per-message framing, tool schemas, image tiles — and which encoding belongs to which model generation.

for a senior

Show you use local counts for context guards, history trimming and limiter admission, then reconcile against usage.prompt_tokens and alert on drift between estimate and actual.

for a principal

Treat token accounting as measurable infrastructure: a single counting path shared across services, versioned with the tokenizer, feeding both capacity planning and per-feature cost attribution.

## What tiktoken is for `tiktoken` is the byte-pair-encoding tokenizer OpenAI ships, and it exists so you can answer two questions *before* spending money or hitting a limit: will this prompt fit in the context window, and roughly what will it cost. Both are pre-flight checks. Neither is a substitute for the authoritative number, which is the `usage` object the API returns. ## Getting the right encoding Two entry points matter: - `tiktoken.encoding_for_model(model_name)` — maps a model name to its encoding. Prefer this; it keeps the mapping knowledge in the library rather than in your head. - `tiktoken.get_encoding(name)` — loads an encoding directly, for example `get_encoding("o200k_base")`. The encodings you will meet are `o200k_base`, used by the gpt-4o generation and later models, and `cl100k_base`, used by gpt-4 and gpt-3.5-turbo. They tokenize differently, so counting gpt-4o traffic with `cl100k_base` gives a systematically wrong number — usually an overcount on non-English text, where the newer encoding is more efficient. There is one operational trap worth knowing: `encoding_for_model` raises `KeyError` for a model name your installed version of `tiktoken` has never heard of. Models ship faster than tokenizer releases, so a brand-new model id can break a counting path that has worked for months. Handle it: catch the `KeyError` and fall back to the newest encoding you know, and treat upgrading `tiktoken` as part of adopting a new model. ## Counting a chat request, not a string `len(encoding.encode(text))` counts a string. A Chat Completions request is not a string — it is a structured conversation the API renders into tokens with framing around each message. OpenAI's own counting example adds a fixed overhead per message (about three tokens) for the role and separators, an extra token when a message carries a `name`, and about three more to prime the assistant's reply. Summing raw content therefore undercounts by roughly three to four tokens per message, which is negligible for one long document and material for a hundred short turns. Two further gaps are larger than the framing: - **Tool and function definitions.** When you pass a `tools` array, the JSON Schemas are serialised into the model's input and billed as prompt tokens. A rich tool catalogue can add thousands of tokens to *every* request, and naive counters miss it entirely because the schemas never appear in `messages`. - **Images and other non-text inputs.** Image inputs are priced through a tile-based formula driven by resolution and detail setting, not by BPE encoding. `tiktoken` has no way to compute this. If your feature is multimodal, take the image token cost from the response's `usage` and model it separately. ## Output tokens cannot be counted in advance `tiktoken` measures only what you send. The completion length is unknown until the model generates it, so a pre-send budget is always "known input + a bounded guess at output". That bound is exactly what your completion cap expresses, which is why a realistic cap is what makes a spend forecast tractable: input measured, output bounded. On reasoning models, remember the hidden reasoning tokens count as output as well, and can dwarf the visible answer. ## Where local counting genuinely pays 1. **Context-window guards.** Reject or truncate before sending, so users get a clear message instead of a 400 about exceeding the model's context length. 2. **History trimming.** Drop or summarise oldest turns until the transcript fits a token budget you chose. 3. **Chunking for embeddings and RAG.** Chunk sizes are naturally expressed in tokens, and encoding then decoding lets you split on real token boundaries rather than guessing with characters. 4. **Rate-limit shaping.** A client-side limiter needs an estimate of a request's token cost to admit or delay it before the call goes out. 5. **Pre-flight cost estimates.** Multiply estimated input tokens by the model's input rate to sanity-check a batch job before launching it. ## Reconcile, always Because of framing, tool schemas, images and model-side details, a local count is an estimate. Log both your estimate and the returned `usage.prompt_tokens`, and watch the ratio. A drift in that ratio is an excellent early signal that something changed — a new tool was added to the catalogue, a template grew, or the model's encoding moved under you. Teams that only track dollars discover these weeks later; teams that track tokens-per-request discover them the same day.

  • Your counting code raises KeyError after you switch to a newly released model. What happened, and what do you do?
    tiktoken.encoding_for_model only knows models present in the installed release, so a model shipped after that release is unmapped and raises KeyError. Upgrade tiktoken as part of adopting the model, and make the code defensive: catch the KeyError and fall back to the newest encoding you know, logging that it happened. Silent fallback with no log is worse — you get numbers that look fine and drift.
  • Why can a local count be thousands of tokens below usage.prompt_tokens on a tool-using request?
    Because the tools array is serialised into the model's input and billed as prompt tokens, but it lives outside the messages array most counters walk. A catalogue of a dozen tools with detailed JSON Schemas and descriptions can add thousands of tokens to every single request. Count the serialised schemas too, or reconcile against the returned usage and treat the gap as a fixed per-request overhead.
  • Is tiktoken any use for estimating what a response will cost?
    Only indirectly. It measures text you already have, and the completion does not exist yet. The practical approach is to measure input exactly, bound output with a realistic completion cap, and calibrate the expected output length from observed usage.completion_tokens on a pilot. On reasoning models add the reasoning tokens, which are billed as output and are frequently several times larger than the visible answer.

saying these in an interview costs you the question

  • Counting characters divided by four instead of encoding
  • Using cl100k_base for gpt-4o generation models
  • Forgetting per-message framing overhead in chat requests
  • Ignoring tool schema tokens that are billed every request
  • Assuming tiktoken can price image inputs

context