skip to content

What does OpenRouter's GET /api/v1/generation endpoint tell you about a finished call?

level: middleimportance: should knowfreq 32%

answer

  1. Look up a call by its id
  2. The response id is the key
  3. Tells you who actually served it
  4. Native beside normalised counts
  5. Written asynchronously, so poll

basics

~10 s

Called with the completion's id, it returns that generation's audit record: total_cost, native and normalised token counts, which upstream provider served it, latency and generation time, finish reason and any cache discount.

solid answer

~50 s

Every OpenRouter chat completion comes back with an `id`. Pass it as the query parameter to `GET https://openrouter.ai/api/v1/generation?id=...` and you get the server's own record of that call: `total_cost` in credits, `native_tokens_prompt` and `native_tokens_completion` (the serving model's own tokenizer) beside the normalised `tokens_prompt`/`tokens_completion`, `provider_name` for the upstream that actually served it, `latency` and `generation_time`, `finish_reason`, `streamed`, and `cache_discount` where prompt caching applied. That makes it the reconciliation tool: given a request id from your logs, you can answer what it cost and who served it after the fact. The catch is that the record is written **asynchronously** — querying immediately after a stream closes can come back empty, so poll with a short backoff rather than treating the first miss as a lost generation. For live per-request cost, inline usage accounting is the cheaper path; this endpoint is for audits and for questions the response never answered.

code

python · 13 lines
python
import os, time, requests

def generation_stats(gen_id, attempts=5):
    url = "https://openrouter.ai/api/v1/generation"
    headers = {"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"}
    delay = 0.5
    for _ in range(attempts):
        r = requests.get(url, params={"id": gen_id}, headers=headers, timeout=30)
        if r.status_code == 200 and r.json().get("data"):
            return r.json()["data"]
        time.sleep(delay)
        delay *= 2
    return None

go deeper

for a junior

Know that each OpenRouter completion has an id, and that passing that id to the generation endpoint returns what the call cost and how long it took.

for a middle

Enumerate the record's fields and explain the native versus normalised token distinction, the cache discount, and why the stats row may not exist the instant a stream closes.

for a senior

Design the reconciliation path: log ids at request time, take inline cost in the hot path, and use this endpoint for sampled or incident-window forensics — especially provider_name when routing hides who answered.

for a principal

Decide how much cost and route telemetry is worth retaining and at what sampling rate, and set the policy that makes spend attributable across teams without adding a round trip to every user-facing request.

## What this endpoint is for There are two ways to learn what an OpenRouter call cost: ask for it in the response, or look it up afterwards by id. `GET /api/v1/generation` is the second. It exists because some questions only make sense after the fact — reconciling a month of logged request ids against a bill, investigating one slow call, or finding out which upstream provider served a request when your routing rules allowed several. ## Getting the id The `id` on the chat completion response is the handle. Log it. A team that logs prompts and latencies but not the generation id has thrown away its only key into this endpoint, and reconstructing it later is impossible. Under streaming, the id appears on the chunks as well, so capture it from the first one rather than waiting for the end. ## What the record contains The useful fields group into three families. **Money.** `total_cost` is the credits charged for the generation. `cache_discount` shows the reduction attributable to a prompt-cache hit, which is how you prove a caching change actually paid for itself rather than merely appearing to. **Tokens.** `native_tokens_prompt` and `native_tokens_completion` are counted by the serving model's own tokenizer — the counts billing derives from. `tokens_prompt` and `tokens_completion` are OpenRouter's normalised counts, which exist so you can compare volume across models with incompatible vocabularies. Confusing the two is the standard reconciliation bug: you compare a normalised count against a native price and the arithmetic never closes. **Route and timing.** `provider_name` names the upstream that served the call — the single most valuable field here, because with fallbacks or provider preferences in play the slug you asked for does not tell you who answered. `latency` reflects time to the first token and `generation_time` the generation itself, so you can separate a slow provider from a long answer. `finish_reason` and `streamed` complete the picture of how the call ended. ## The asynchronous gap The stats row is written after the generation completes, and it is not instantaneous. Query the endpoint the microsecond your stream closes and you may get nothing back. The correct client is a short retry loop — a few attempts with a growing delay — and the correct mental model is "eventually consistent audit log", not "synchronous read-your-writes". Systems that treat the first miss as a hard failure end up with gaps in their cost data precisely on the fastest calls. ## When to use it versus inline accounting If you want cost on every request in production, enable inline usage accounting and log the number as the response lands: one round trip, no polling, no gap. Reach for the generation endpoint when: - You are auditing historical ids and never captured cost at the time. - You need `provider_name` to explain a quality or latency regression. - You want to prove a cache discount or compare native against normalised counts. - You are debugging one specific call rather than instrumenting all of them. Using it as your primary billing pipeline means one extra HTTP round trip and a retry loop per request, which at volume is real load and real latency for data you could have had for free in the response. ## Practical shape A sensible design records `id`, model slug, and your own dimensions at request time; records inline cost when the response arrives; and keeps a small background job that fills in `provider_name` and timing from the generation endpoint for a sample of requests, or for all requests in an incident window. That gives you cheap always-on cost data plus deep forensics when something goes wrong, without paying for forensics on every call.

  • You call the generation endpoint right after a stream closes and get nothing back. Is the generation lost?
    No. The stats record is written asynchronously once the generation finishes, so an immediate read can race it. Retry a few times with an exponential delay before concluding anything. Treating the first empty read as a failure creates gaps in cost data that correlate with the fastest calls, which is the worst possible sampling bias for a spend dashboard.
  • Which field explains why two calls with the same model slug had very different latency?
    `provider_name`, read together with `latency` and `generation_time`. The same slug can be served by different upstream providers, and they differ in time to first token, throughput and queueing. Pinning the slowdown to one provider is what turns a vague 'the model got slow' report into an actionable routing change.
  • When would you not use this endpoint for cost tracking?
    For always-on production accounting. Inline usage accounting returns the cost on the completion itself with no extra round trip and no polling, so it scales better and has no consistency gap. Keep the generation endpoint for audits of historical ids, incident forensics, and questions the completion response cannot answer, such as which provider served the call.

saying these in an interview costs you the question

  • Not logging the completion id, so lookups are impossible
  • Treating an immediate empty read as a lost generation
  • Comparing normalised token counts against native per-token prices
  • Polling this endpoint for every production request
  • Assuming the requested slug identifies the serving provider

context