skip to content

How do you make OpenRouter return the actual cost and native token counts inline?

level: middleimportance: must knowfreq 50%

answer

  1. One request field switches it on
  2. Cost arrives with the response
  3. Native counts, not normalised
  4. Streaming hides it at the end
  5. Empty choices, populated usage

basics

~20 s

Send "usage": {"include": true} in the chat completion request body. The response's usage object then carries cost in credits alongside token counts from the serving model's own tokenizer, with cost_details breaking out any upstream charge.

solid answer

~50 s

OpenRouter calls this **usage accounting**: add `"usage": {"include": true}` to the `POST /api/v1/chat/completions` body and the response's `usage` object comes back enriched. You get `cost` — the credits actually deducted for that call — plus `cost_details` (which separates any upstream inference charge under a bring-your-own-key route), and token counts reported by the serving model's **native** tokenizer rather than OpenRouter's cross-model normalised counts. Detail fields such as `prompt_tokens_details.cached_tokens` and `completion_tokens_details.reasoning_tokens` show up when the model supports caching or reasoning tokens. In a streamed response the usage block does not ride along with the text: it arrives in a **final SSE chunk whose `choices` array is empty**, after the last content delta. A client that stops reading once it sees a finish reason, or that skips chunks with no choices, silently loses every cost figure — which is how teams end up with a streaming endpoint they cannot bill.

code

python · 15 lines
python
import os, requests

resp = requests.post(
    "https://openrouter.ai/api/v1/chat/completions",
    headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
    json={
        "model": os.environ["OPENROUTER_MODEL"],
        "messages": [{"role": "user", "content": "Say hi."}],
        "usage": {"include": True},
    },
    timeout=60,
).json()

usage = resp["usage"]
print(usage["prompt_tokens"], usage["completion_tokens"], usage["cost"])

go deeper

for a junior

Know that one request field, usage with include set to true, makes OpenRouter report what the call cost, and that the number comes back on the response rather than being calculated by you.

for a middle

Explain the whole usage object — cost, cost_details, native versus normalised token counts, cached and reasoning token details — and describe exactly where the block lands in a streamed response.

for a senior

Demonstrate the operational side: record cost per request with tenant and feature dimensions, alarm on cost per request rather than only total spend, and diagnose the streaming client bug that silently drops all accounting.

for a principal

Make per-request cost a first-class signal in the platform: mandate accounting in the shared client so no team can ship an uninstrumented call path, and decide what cost telemetry is retained, at what cardinality, and who is accountable when it moves.

## The problem usage accounting solves Estimating cost client-side means multiplying a cached price table by a token count you also estimated. Both halves are wrong in interesting ways: the price depends on which provider served the route, and your token count depends on which tokenizer you assumed. Usage accounting removes the estimate — the platform tells you what it charged. ## Turning it on It is one field in the request body: ``` "usage": { "include": true } ``` That is an OpenRouter extension, and it is a common trip-hazard for people arriving from an OpenAI-shaped client, where the analogous streaming knob is `stream_options: {"include_usage": true}`. On OpenRouter the `usage` object in the request is the switch, and it covers both streaming and non-streaming calls. ## What comes back The response `usage` object gains: - `cost` — credits deducted for this request. One credit is one US dollar, so this is a dollar figure. - `cost_details` — a breakdown; the notable member is `upstream_inference_cost`, which is populated on bring-your-own-key routes to show what the upstream provider will bill you separately. - `prompt_tokens`, `completion_tokens`, `total_tokens` — counts as measured by the serving model. - `prompt_tokens_details.cached_tokens` — how much of the input was served from a prompt cache, on models that support one. - `completion_tokens_details.reasoning_tokens` — hidden reasoning tokens on reasoning models, which you pay for but never see as text. Those last two are where surprise bills live. A reasoning model can bill several times the visible output length, and a dashboard that counts only characters returned to the user will never explain the invoice. ## Normalised versus native token counts OpenRouter serves models with radically different tokenizers, so for cross-model comparability it reports normalised counts by default. Billing, however, happens against the tokens the serving model actually consumed — its **native** counts. With usage accounting enabled you see the native figures, which is why a token count from an accounting-enabled response will not always match one you computed with a generic tokenizer library, and why the accounting figures are the ones to reconcile against spend. If you are building a cost model, always compare like with like: native tokens against the provider's per-token rate. ## Streaming: the chunk after the last chunk With `stream: true`, deltas arrive as Server-Sent Events. The content ends, a chunk carries the finish reason, and then **one more chunk arrives whose `choices` array is empty and whose `usage` object holds the accounting**. Two very common client bugs drop it: 1. Breaking out of the read loop the moment a finish reason appears. 2. Guarding the loop body with something like `if chunk.choices:` and skipping everything else. Both produce a system that streams beautifully and reports zero cost. The correct shape is to keep reading until the stream terminator, and to handle a chunk with no choices by looking for `usage` on it. ## Cost of the feature itself Asking for accounting does not change what you are charged for the inference and does not meaningfully change latency — the numbers are already computed on OpenRouter's side. There is no reason to leave it off in production; the only argument for omitting it is that you have nowhere to put the number, which is itself the problem. ## Using it well Log `cost` per request alongside your own dimensions — tenant, feature, model slug, request id — at the moment the response completes. That single line is what lets you answer "which feature doubled our spend last Tuesday" without a reconciliation project. Aggregate it into a rolling spend metric and alarm on the derivative, not just the level: a 10x jump in cost per request usually means routing moved to a pricier provider or somebody enabled a reasoning mode, and both are cheaper to catch in an hour than in a month. ## Pitfalls to name in an interview The accounting reports what was charged **for that request**, not your remaining balance and not account-wide spend. Cached input is cheaper but still billed. Reasoning tokens count. And on a failed or aborted request, whatever tokens were generated before the abort can still be billed, so a client that cancels aggressively should still record whatever accounting it managed to receive.

  • Your streaming endpoint logs cost as null for every request even though accounting is enabled. What is the bug?
    The read loop is discarding the final chunk. With streaming, OpenRouter sends the usage block in one last SSE chunk whose `choices` array is empty, after the finish reason. Clients that break on the finish reason, or that only process chunks where `choices` is non-empty, throw that chunk away. Read until the stream terminator and treat a chunk with no choices but a `usage` object as the accounting record.
  • Why might the token counts from usage accounting disagree with what a local tokenizer library computed for the same prompt?
    Because accounting reports the serving model's native tokenizer counts — the ones billing is computed from — while a local library tokenizes with whatever vocabulary it ships. OpenRouter also normalises counts for cross-model comparability when accounting is off. Reconcile spend against the native figures; use a local tokenizer only for rough pre-flight budgeting, never for invoicing.
  • A reasoning model's bill is far higher than the visible answer length suggests. Where does usage accounting show why?
    Under `completion_tokens_details.reasoning_tokens`. Reasoning models emit hidden chain-of-thought tokens that are billed as output but never surface as text. If your dashboard counts characters returned to the user it will always under-report. Log the reasoning token count separately so a mode change is visible as a line on a chart rather than a surprise at the end of the month.

saying these in an interview costs you the question

  • Estimating cost from a cached price table instead of the response
  • Using OpenAI's stream_options include_usage flag on OpenRouter
  • Skipping streamed chunks whose choices array is empty
  • Assuming usage.cost is the remaining account balance
  • Ignoring reasoning and cached-token detail fields

context