How do you get accurate per-call token usage from a LangChain chat model?
answer
- the provider's count, not yours
- it rides on the response object
- streams hold it back until the end
- chunks add together
- estimates are for fitting, not billing
basics
~20 sRead usage_metadata on the returned AIMessage — it carries input_tokens, output_tokens and total_tokens reported by the provider. When streaming, request usage explicitly (ChatOpenAI's stream_usage=True) and sum the chunks, because usage arrives only on the final chunk.
solid answer
~40 sThe authoritative source is `AIMessage.usage_metadata`, populated from the provider's own accounting: `input_tokens`, `output_tokens`, `total_tokens`, plus detail dicts that break out cached input and reasoning tokens where the provider reports them. `response_metadata` keeps the raw provider payload if you need vendor-specific fields. Three gotchas bite in production. First, streaming: most providers omit usage from chunks unless asked, so set `stream_usage=True` on `ChatOpenAI` and add the `AIMessageChunk` objects together (`full = full + chunk`) — the aggregate carries the usage. Second, `with_structured_output(..., include_raw=False)` throws the `AIMessage` away, so usage disappears; pass `include_raw=True` and read `result["raw"].usage_metadata`. Third, `get_num_tokens_from_messages()` is a local pre-call *estimate* for budgeting and truncation, not a billing figure — it does not know about tool schemas or provider-side prompt additions.
code
python · 12 linesfrom langchain_openai import ChatOpenAI
model = ChatOpenAI(model="gpt-4o-mini", stream_usage=True)
full = None
for chunk in model.stream("Summarise photosynthesis in one sentence."):
print(chunk.content, end="")
full = chunk if full is None else full + chunk
print()
print(full.usage_metadata)
# {'input_tokens': 16, 'output_tokens': 31, 'total_tokens': 47, ...}go deeper
Know that the response object carries usage_metadata with input, output and total token counts, and that these come from the provider rather than being computed locally.
Explain how streaming changes it: chunks carry no usage by default, you opt in and aggregate them, and structured output without include_raw drops the message entirely.
Show a working accounting design — pre-call estimation for context fitting, per-call ledger of usage_metadata including cached and reasoning breakdowns, and reconciliation against the provider's invoice.
Own cost observability as a system property: per-tenant and per-feature attribution, drift detection between recorded and billed usage, and the policy decision of what happens when a budget is exceeded mid-conversation.
## Where the numbers live Every LangChain chat-model response is an `AIMessage`, and the standardized accounting hangs off it as `usage_metadata`. The shape is provider-neutral: - `input_tokens` — prompt tokens the provider billed. - `output_tokens` — completion tokens. - `total_tokens` — their sum. - `input_token_details` — breakdowns the provider reports, such as cache-read tokens. - `output_token_details` — breakdowns such as reasoning tokens on reasoning models. Those details matter more than they look. A prompt-cached call can bill a large input at a fraction of the price, and a reasoning model can spend far more output tokens than the visible answer contains. A dashboard that only tracks `total_tokens` will misprice both. `response_metadata` on the same message holds the provider's raw payload — model name actually served, finish reason, the vendor's own usage block. Use it for vendor-specific fields; use `usage_metadata` for anything that must survive a provider swap. ## Streaming is the classic trap When you stream, you receive `AIMessageChunk` objects, and by default providers do not attach usage to them: the numbers are only known when generation ends. `ChatOpenAI` exposes `stream_usage=True` (settable on the constructor or per call) to request the final usage block in the stream. Even with that flag, individual chunks are the wrong place to look. `AIMessageChunk` supports `+`, and the idiom is to fold the stream into one message: ``` full = None for chunk in model.stream(...): full = chunk if full is None else full + chunk ``` The aggregate carries the merged content and the usage. Teams that only forward chunks to the client and never aggregate end up with zero usage recorded for every streamed request — which is exactly the traffic that dominates a chat product. ## The wrapper that eats your metadata `model.with_structured_output(Schema)` returns the parsed object, not the message. That is convenient and it silently discards `usage_metadata`. In a service that meters cost per tenant, this shows up as a whole class of calls reporting nothing. The fix is `include_raw=True`, which returns `{"raw": AIMessage, "parsed": ..., "parsing_error": ...}` so you can log usage and still get your object. The same principle applies anywhere a chain converts the message into something else early — `StrOutputParser`, a custom lambda pulling `.content`. If you need the metadata, capture it before the message is dropped, or use a callback handler that observes model responses and aggregates `usage_metadata` across a run. ## Estimating before you call Chat models expose `get_num_tokens(text)` and `get_num_tokens_from_messages(messages)`. For OpenAI models these use tiktoken with the model's encoding; other integrations approximate. They are the right tool for *pre-call* decisions: will this history fit the context window, how many history turns should I drop, should I summarize before sending. They are the wrong tool for billing. The count omits tool and structured-output schemas serialized into the request, any provider-side prompt scaffolding, and per-message overhead the vendor may change without notice. Treat the estimate as a budget guard with headroom, and the provider's returned usage as the ledger. ## Building it into a service A usable setup has three layers: 1. **Pre-call guard** — estimate with `get_num_tokens_from_messages()`, truncate or summarize history so you stay under the window with margin. 2. **Post-call ledger** — record `usage_metadata` from every response (including the aggregated streamed one), tagged with tenant, feature, model and the detail breakdowns. 3. **Reconciliation** — periodically compare your recorded totals against the provider's invoice or usage dashboard. Drift points at a code path that drops metadata, usually streaming or structured output. Also record the model name from `response_metadata` rather than the one you asked for: aliases and auto-upgrades mean the served model is not always the requested one, and cost per token differs. ## Interview framing Name `usage_metadata` as the single source of truth, then immediately volunteer the two paths that silently lose it — streaming without aggregation, and structured output without `include_raw` — and separate local estimation from provider accounting. That combination is what distinguishes someone who has actually run a billed LLM service from someone who has read the quickstart.
- Why does token usage often come back empty for streamed calls?Providers know the totals only when generation completes, so they omit usage from intermediate chunks unless you opt in — `stream_usage=True` on `ChatOpenAI`. Even then the numbers ride on the final chunk, so code that inspects chunks individually sees nothing. Fold the stream with `full = full + chunk` and read `usage_metadata` off the aggregate.
- A tenant's bill is far above your recorded totals. Where do you look first?Code paths that discard the AIMessage: streamed responses forwarded without aggregation, and `with_structured_output(include_raw=False)`. Then check cached-input and reasoning breakdowns in the detail dicts, and confirm the served model from `response_metadata` matches the one you priced — an alias resolving to a costlier model explains a lot of drift.
- Is get_num_tokens_from_messages() good enough to enforce a hard spend cap?No. It is a local estimate that cannot see tool or structured-output schemas serialized into the request, provider-side prompt scaffolding, or per-message overhead the vendor may change. Use it to decide what fits the context window with headroom, and enforce spend against the provider-reported `usage_metadata` you record after each call.
- What do the input_token_details and output_token_details breakdowns tell you?They separate tokens that are billed differently — cache-read input tokens on a prompt-cached call, and reasoning tokens produced by a reasoning model but not shown in the answer. Ignoring them misprices both directions: cached prompts look more expensive than they were, and reasoning output looks far cheaper.
saying these in an interview costs you the question
- Counting tokens locally and calling it the billed amount
- Reading usage off individual streamed chunks
- Not noticing structured output discards the message
- Tracking only total_tokens and ignoring cached or reasoning detail
- Assuming the served model equals the requested model