How do you attribute token cost in a LlamaIndex app using TokenCountingHandler?
answer
- a callback, not a return value
- counters keep counting until told otherwise
- prompt and completion priced differently
- embeddings counted separately
- global handler cannot attribute per request
basics
~20 sAttach a TokenCountingHandler to a CallbackManager (usually via Settings.callback_manager), give it a tokenizer matching your model, then read prompt_llm_token_count, completion_llm_token_count and total_embedding_token_count. Counters accumulate until you call reset_counts(), so scope or reset them per request.
solid answer
~40 s`TokenCountingHandler` lives in `llama_index.core.callbacks` and is wired in through a `CallbackManager` — most commonly `Settings.callback_manager = CallbackManager([token_counter])`, which makes it observe every LLM and embedding event in the process. Construct it with a `tokenizer` callable matching the model you actually call (for OpenAI models, `tiktoken.encoding_for_model(...).encode`); a mismatched tokenizer gives you numbers that are directionally right and financially wrong. After a query you read `prompt_llm_token_count`, `completion_llm_token_count`, `total_llm_token_count` and `total_embedding_token_count`, and `llm_token_counts` for the per-event breakdown. The counters are cumulative — they never reset themselves — so call `reset_counts()` at a boundary you define. The trap in a server is that a globally installed handler counts every concurrent request into the same accumulator, so you cannot attribute cost per request from it; scope a handler per request, or take the attribution from your tracing backend instead.
code
python · 20 linesimport tiktoken
from llama_index.core import Document, Settings, VectorStoreIndex
from llama_index.core.callbacks import CallbackManager, TokenCountingHandler
token_counter = TokenCountingHandler(
tokenizer=tiktoken.encoding_for_model("gpt-4o").encode
)
Settings.callback_manager = CallbackManager([token_counter])
index = VectorStoreIndex.from_documents(
[Document(text="Refunds are accepted within 30 days of purchase.")]
)
print("index embedding tokens:", token_counter.total_embedding_token_count)
token_counter.reset_counts()
index.as_query_engine().query("What is the refund window?")
print("prompt:", token_counter.prompt_llm_token_count)
print("completion:", token_counter.completion_llm_token_count)
print("query embedding:", token_counter.total_embedding_token_count)
print("llm calls this query:", len(token_counter.llm_token_counts))go deeper
Know that token usage comes from a callback handler you install, not from the query result, and that you read named counters off the handler afterwards.
Explain how the handler attaches through a CallbackManager, why prompt, completion and embedding tokens are counted separately, and why the tokenizer must match the model.
Show the server-side judgement: cumulative counters plus a global handler means no per-request attribution, so scope a handler per request or take token attribution from tracing spans.
Frame it as unit economics — cost per query as a tracked metric next to quality, the levers that move it (top-k, chunk size, synthesizer mode, rerank), and what a quality gain is worth per token.
## Why this exists A RAG query is not one LLM call. It is one or more embedding calls to encode the query, possibly a rerank call, and then one or more generator calls — several, if the response synthesizer refines across chunks. The bill is the sum, and nothing in the query's return value tells you what that sum was. `TokenCountingHandler` is the callback that adds it up. ## Wiring it in It is a callback handler, so it goes into a `CallbackManager`, and the manager goes somewhere the framework consults. The global path is `Settings.callback_manager = CallbackManager([token_counter])`, after which every component constructed from those settings emits events into it. Components also accept a `callback_manager` directly, which is how you scope counting to one index or one query engine rather than the whole process. The constructor's important argument is `tokenizer`: a callable that turns a string into a list of tokens. Pass the encoder for the model you actually call — for OpenAI models that is `tiktoken.encoding_for_model(model_name).encode`. If you leave it to a default and then call a model with a different tokenizer, every count is an estimate with a systematic bias, and the error compounds across thousands of requests. `verbose=True` prints each event's counts as it happens, which is useful exactly once, while you are learning what your pipeline calls. ## What it exposes Four aggregate counters matter: - `prompt_llm_token_count` — input tokens sent to the generator. In RAG this is the big one, because it contains all the retrieved chunk text. - `completion_llm_token_count` — output tokens generated. - `total_llm_token_count` — the sum of the two. - `total_embedding_token_count` — tokens embedded, both at index build and at query time. Separating prompt from completion is not pedantry: providers price them differently, usually with output several times more expensive per token, and the two respond to completely different levers. Prompt tokens are driven by `similarity_top_k` and chunk size; completion tokens by your prompt's verbosity instructions and any max-token cap. `llm_token_counts` holds the per-event records rather than the totals, which is what you need when a single query fired several LLM calls and you want to know which one was expensive — a refine-style synthesizer that walks ten chunks makes ten calls, and the aggregate hides that entirely. ## The cumulative trap Counters accumulate for the lifetime of the handler. Nothing resets them between queries. Read them, then call `reset_counts()` — or read deltas. In a script this is a minor discipline. In a server it is a correctness bug waiting to happen: a handler installed on global `Settings` sees every request from every thread or task, so by the time you read the counter after handling request A, requests B and C have added their tokens to it. There is no request identity in the aggregate. The fixes are the obvious ones. Build a fresh `TokenCountingHandler` and `CallbackManager` per request and attach it to the query engine you construct for that request, so the counts are naturally isolated. Or accept the global handler as a process-level meter — useful for "what did this service spend this hour" — and get per-request attribution from a tracing backend, which records token counts per span with a trace id attached. ## What you do with the numbers Token counting turns qualitative pipeline decisions into arithmetic. Doubling `similarity_top_k` roughly doubles the retrieved text in the prompt; the counter tells you exactly what that costs before you argue about whether the quality gain justified it. A tree-summarize or refine synthesizer over many chunks makes multiple generator calls where a compact one makes fewer; the per-event list shows the multiplier. Reranking adds a call but often lets you cut top-k, and the net can go either way — measure it. The strongest use is pairing cost with quality on the same eval set: run the batch, record the pass rate, record the tokens, and report both. A configuration that improves faithfulness by two points while tripling prompt tokens is a decision, not an improvement, and putting both numbers in front of the decision-maker is the job. ## Limits It counts tokens, not money — you apply the price table yourself, and prices differ per model, so a mixed-model pipeline needs per-model attribution that the aggregate counters do not give you. It also counts what your process sent, which can diverge from what the provider bills if the provider counts differently or if retries occurred inside the client below the callback layer.
- You raise similarity_top_k from 3 to 10 and cost jumps. Which counter moved, and why?`prompt_llm_token_count`, because retrieved chunk text is stuffed into the generator's prompt — more chunks means a proportionally larger input. Completion tokens barely move, since the answer length is set by the question and your instructions, not by how much context you supplied. Query-time embedding tokens do not move either; you still embed one query. It is a pure input-side cost, which is why top-k tuning is best argued with the counter in hand.
- A single query shows several entries in llm_token_counts. What does that indicate?That synthesis made more than one generator call. Refine-style synthesis walks the retrieved chunks sequentially, calling the model once per chunk to update the answer, and tree-summarize calls it repeatedly as it collapses chunks pairwise. The aggregate counters hide this multiplier entirely, so the per-event list is where you discover that a ten-chunk retrieval turned into ten LLM calls rather than one.
- Why can a globally installed handler mislead you in a web service?Because the counters are process-wide and cumulative, with no request identity attached. Under concurrency, tokens from other in-flight requests land in the same accumulator between the moment your request starts and the moment you read the total, so any per-request figure you derive is contaminated. Either construct a handler per request and attach it to that request's query engine, or take per-request attribution from tracing spans and treat the global handler as a process meter.
saying these in an interview costs you the question
- Expecting token counts on the query response object
- Forgetting counters are cumulative across queries
- Using a tokenizer that does not match the called model
- Reading a global counter to bill a single request
- Treating token count and dollar cost as the same number