skip to content

How do you get accurate per-call token usage from a LangChain chat model?

level: seniorimportance: should knowfreq 52%

answer

  1. the provider's count, not yours
  2. it rides on the response object
  3. streams hold it back until the end
  4. chunks add together
  5. estimates are for fitting, not billing

basics

~20 s

Read usage_metadata on the returned AIMessage — it carries input_tokens, output_tokens and total_tokens reported by the provider. When streaming, request usage explicitly (ChatOpenAI's stream_usage=True) and sum the chunks, because usage arrives only on the final chunk.

solid answer

~40 s

The authoritative source is `AIMessage.usage_metadata`, populated from the provider's own accounting: `input_tokens`, `output_tokens`, `total_tokens`, plus detail dicts that break out cached input and reasoning tokens where the provider reports them. `response_metadata` keeps the raw provider payload if you need vendor-specific fields. Three gotchas bite in production. First, streaming: most providers omit usage from chunks unless asked, so set `stream_usage=True` on `ChatOpenAI` and add the `AIMessageChunk` objects together (`full = full + chunk`) — the aggregate carries the usage. Second, `with_structured_output(..., include_raw=False)` throws the `AIMessage` away, so usage disappears; pass `include_raw=True` and read `result["raw"].usage_metadata`. Third, `get_num_tokens_from_messages()` is a local pre-call *estimate* for budgeting and truncation, not a billing figure — it does not know about tool schemas or provider-side prompt additions.

code

python · 12 lines
python
from langchain_openai import ChatOpenAI

model = ChatOpenAI(model="gpt-4o-mini", stream_usage=True)

full = None
for chunk in model.stream("Summarise photosynthesis in one sentence."):
    print(chunk.content, end="")
    full = chunk if full is None else full + chunk

print()
print(full.usage_metadata)
# {'input_tokens': 16, 'output_tokens': 31, 'total_tokens': 47, ...}

go deeper

for a junior

Know that the response object carries usage_metadata with input, output and total token counts, and that these come from the provider rather than being computed locally.

for a middle

Explain how streaming changes it: chunks carry no usage by default, you opt in and aggregate them, and structured output without include_raw drops the message entirely.

for a senior

Show a working accounting design — pre-call estimation for context fitting, per-call ledger of usage_metadata including cached and reasoning breakdowns, and reconciliation against the provider's invoice.

for a principal

Own cost observability as a system property: per-tenant and per-feature attribution, drift detection between recorded and billed usage, and the policy decision of what happens when a budget is exceeded mid-conversation.

## Where the numbers live Every LangChain chat-model response is an `AIMessage`, and the standardized accounting hangs off it as `usage_metadata`. The shape is provider-neutral: - `input_tokens` — prompt tokens the provider billed. - `output_tokens` — completion tokens. - `total_tokens` — their sum. - `input_token_details` — breakdowns the provider reports, such as cache-read tokens. - `output_token_details` — breakdowns such as reasoning tokens on reasoning models. Those details matter more than they look. A prompt-cached call can bill a large input at a fraction of the price, and a reasoning model can spend far more output tokens than the visible answer contains. A dashboard that only tracks `total_tokens` will misprice both. `response_metadata` on the same message holds the provider's raw payload — model name actually served, finish reason, the vendor's own usage block. Use it for vendor-specific fields; use `usage_metadata` for anything that must survive a provider swap. ## Streaming is the classic trap When you stream, you receive `AIMessageChunk` objects, and by default providers do not attach usage to them: the numbers are only known when generation ends. `ChatOpenAI` exposes `stream_usage=True` (settable on the constructor or per call) to request the final usage block in the stream. Even with that flag, individual chunks are the wrong place to look. `AIMessageChunk` supports `+`, and the idiom is to fold the stream into one message: ``` full = None for chunk in model.stream(...): full = chunk if full is None else full + chunk ``` The aggregate carries the merged content and the usage. Teams that only forward chunks to the client and never aggregate end up with zero usage recorded for every streamed request — which is exactly the traffic that dominates a chat product. ## The wrapper that eats your metadata `model.with_structured_output(Schema)` returns the parsed object, not the message. That is convenient and it silently discards `usage_metadata`. In a service that meters cost per tenant, this shows up as a whole class of calls reporting nothing. The fix is `include_raw=True`, which returns `{"raw": AIMessage, "parsed": ..., "parsing_error": ...}` so you can log usage and still get your object. The same principle applies anywhere a chain converts the message into something else early — `StrOutputParser`, a custom lambda pulling `.content`. If you need the metadata, capture it before the message is dropped, or use a callback handler that observes model responses and aggregates `usage_metadata` across a run. ## Estimating before you call Chat models expose `get_num_tokens(text)` and `get_num_tokens_from_messages(messages)`. For OpenAI models these use tiktoken with the model's encoding; other integrations approximate. They are the right tool for *pre-call* decisions: will this history fit the context window, how many history turns should I drop, should I summarize before sending. They are the wrong tool for billing. The count omits tool and structured-output schemas serialized into the request, any provider-side prompt scaffolding, and per-message overhead the vendor may change without notice. Treat the estimate as a budget guard with headroom, and the provider's returned usage as the ledger. ## Building it into a service A usable setup has three layers: 1. **Pre-call guard** — estimate with `get_num_tokens_from_messages()`, truncate or summarize history so you stay under the window with margin. 2. **Post-call ledger** — record `usage_metadata` from every response (including the aggregated streamed one), tagged with tenant, feature, model and the detail breakdowns. 3. **Reconciliation** — periodically compare your recorded totals against the provider's invoice or usage dashboard. Drift points at a code path that drops metadata, usually streaming or structured output. Also record the model name from `response_metadata` rather than the one you asked for: aliases and auto-upgrades mean the served model is not always the requested one, and cost per token differs. ## Interview framing Name `usage_metadata` as the single source of truth, then immediately volunteer the two paths that silently lose it — streaming without aggregation, and structured output without `include_raw` — and separate local estimation from provider accounting. That combination is what distinguishes someone who has actually run a billed LLM service from someone who has read the quickstart.

  • Why does token usage often come back empty for streamed calls?
    Providers know the totals only when generation completes, so they omit usage from intermediate chunks unless you opt in — `stream_usage=True` on `ChatOpenAI`. Even then the numbers ride on the final chunk, so code that inspects chunks individually sees nothing. Fold the stream with `full = full + chunk` and read `usage_metadata` off the aggregate.
  • A tenant's bill is far above your recorded totals. Where do you look first?
    Code paths that discard the AIMessage: streamed responses forwarded without aggregation, and `with_structured_output(include_raw=False)`. Then check cached-input and reasoning breakdowns in the detail dicts, and confirm the served model from `response_metadata` matches the one you priced — an alias resolving to a costlier model explains a lot of drift.
  • Is get_num_tokens_from_messages() good enough to enforce a hard spend cap?
    No. It is a local estimate that cannot see tool or structured-output schemas serialized into the request, provider-side prompt scaffolding, or per-message overhead the vendor may change. Use it to decide what fits the context window with headroom, and enforce spend against the provider-reported `usage_metadata` you record after each call.
  • What do the input_token_details and output_token_details breakdowns tell you?
    They separate tokens that are billed differently — cache-read input tokens on a prompt-cached call, and reasoning tokens produced by a reasoning model but not shown in the answer. Ignoring them misprices both directions: cached prompts look more expensive than they were, and reasoning output looks far cheaper.

saying these in an interview costs you the question

  • Counting tokens locally and calling it the billed amount
  • Reading usage off individual streamed chunks
  • Not noticing structured output discards the message
  • Tracking only total_tokens and ignoring cached or reasoning detail
  • Assuming the served model equals the requested model

context