Why meter LLM spend from the provider's reported token usage rather than a local tokenizer estimate?
answer
- the bill comes from the response, not your guess
- some billed tokens are never shown
- output length cannot be predicted
- cache split is decided at serve time
- reconcile the total against the invoice
basics
~20 sThe response's usage record carries the counts you are actually billed for, including tokens you never see, such as reasoning tokens and cached-prefix tokens. A local tokenizer only guesses the prompt, misses server-side additions, and drifts as models change.
solid answer
~50 sEvery mainstream LLM API returns a usage record with the call's token counts, and that record is what the invoice is derived from — so it is the only figure worth putting in a cost dashboard. A client-side tokenizer can only count text you assembled yourself: it cannot know output length, it cannot see a reasoning model's internal thinking tokens, it cannot tell you how much of the prompt was served from a cached prefix versus written into the cache, and its vocabulary drifts as model families change. Estimates still have a job — admission control before a call, context-window fit checks, a live cost indicator while a response streams — but they should be stored as a separate, clearly-labelled series. The real discipline is emitting one usage event per call, from **every** call path, and reconciling the total against the provider invoice each billing period.
go deeper
Know that the API response carries the token counts you are billed for, and that you should log those rather than counting the prompt yourself. Be able to name one billed thing a local count would miss.
Explain why output, reasoning and cached-prefix tokens cannot be predicted client-side, and describe capturing usage in one client wrapper so every call path is metered the same way.
Show the operational habit: reconcile telemetry against the invoice each period, treat a gap as an unmetered call path, and keep estimated and billed counters in separate series so a guess never reaches a margin report.
Own the accounting contract — which figure the business plans on, who is allowed to change the rate card, and why priced history must be pinned to the rates in force at call time rather than re-priced on every query.
## What "reported usage" is Alongside the generated text, an LLM API returns a usage record for that call: input tokens, output tokens, and — where the model and provider support them — reasoning (thinking) tokens and the split between input tokens read from a cached prefix and tokens written into that cache. This record is produced by the same accounting the provider bills from, so it is the authoritative number. Cost telemetry means capturing it on every call, attaching attribution (tenant, feature, model, prompt version), and pricing it with a rate card. ## Why client-side estimation systematically under-counts - **Output is unknowable in advance.** The largest and most variable part of the bill for a chat or agent call is generated tokens, which do not exist until the model has produced them. - **Hidden generated tokens.** On reasoning models, internal thinking tokens are billed but not shown to the user. They can be several times the visible answer, so a counter built from visible text alone can be off by a large multiple. - **Server-side additions.** Tool definitions, system instructions injected by your framework, structured-output scaffolding and retrieved context assembled downstream of your estimator all land in the billed input. - **Cache split depends on the hit.** Whether a prefix was read cheaply from cache or written at a premium is determined at serve time, and the two are priced differently. No client-side count can predict it. - **Tokenizer drift and non-text inputs.** Vocabularies differ between model families and change between generations, and images, audio and documents are not priced by counting characters at all. ## Where estimates legitimately live Estimates are for decisions you must make *before* the usage record exists: refusing a request that will not fit the context window, enforcing a per-request or per-tenant quota pre-flight, showing a running cost indicator during a stream (the usage record typically arrives at the end), and reasoning about a prompt change's token delta in review. Keep them in their own metric names — `estimated_*` versus `billed_*` — so nobody mixes a guess into a margin model. Also record cancelled and errored calls: an aborted generation may still have produced billed tokens, and a dashboard that only logs successes will disagree with the invoice. ## The pipeline that makes it trustworthy 1. Wrap every model client so the usage record is captured in one place rather than at each call site. 2. Emit one event per call: counters, model id, effort or thinking setting, prompt version, attribution keys, latency, status, and the rate-card version used. 3. Price at ingestion or at query time — pinning the rate-card version matters, because provider prices change and a re-priced history quietly rewrites last quarter's unit economics. 4. Aggregate low-cardinality dimensions into metrics for alerting; keep the raw events for attribution queries. 5. **Reconcile monthly.** Compare summed telemetry against the provider invoice and alert when the gap exceeds a few percent. ## The failure this catches Reconciliation drift almost always means an unmetered call path. The usual culprits are retries performed inside an SDK or framework, embedding calls for retrieval, judge or grader calls in an eval job, background and cron work, sub-agent calls that use a differently-configured client, and a debugging script someone left running. A dashboard that shows sixty percent of the invoice is not a rounding problem; it is a call path you do not know about, and it is exactly the path most likely to leak spend.
- Your dashboard shows 62% of what the provider invoiced. Where do you look first?Assume an unmetered call path rather than a pricing bug. Common sources: retries inside an SDK or framework, embedding calls for retrieval, judge or grader calls in eval jobs, cron and background work, sub-agents constructed with their own client, and ad-hoc scripts using production keys. Fix it by capturing usage in a single client wrapper that every path must go through, then re-reconcile.
- A request is cancelled mid-stream. What should the telemetry record?Record the call with a cancelled status and whatever usage the provider reports, because tokens generated before the abort are generally still billed. Dropping cancelled calls is a classic reason telemetry undershoots the invoice, and a rising cancelled-call cost share is itself a useful signal about users abandoning slow responses.
- Would you ever price a stored usage event later, at query time?Yes, but only if you pin which rate card applies to which period. Pricing at query time keeps history flexible and lets you model a price change, but applying today's rates to last quarter's tokens silently rewrites reported margin. Store the counters raw plus the rate-card version in force at call time.
saying these in an interview costs you the question
- Estimating tokens with a local tokenizer and calling it billing data
- Assuming reasoning tokens are free because they are not displayed
- Logging only successful calls, so the invoice always exceeds the dashboard
- Treating a character-count heuristic as accurate enough for margins
- Never reconciling telemetry totals against the provider invoice