How does OpenRouter bill a call, and where do its per-model prices come from?
answer
- One prepaid balance, many vendors
- Prices ship as API data
- Watch the unit on pricing
- Per single token, decimal string
- GET /api/v1/credits for balance
basics
~10 sOpenRouter runs on prepaid credits: you top up one balance and each call deducts the chosen provider's own per-token price. GET /api/v1/models publishes every model's pricing object, quoted in US dollars per single token.
solid answer
~50 sYou buy credits on OpenRouter once, and every request against any of the hundreds of models draws down that single balance — there is no separate account or invoice per vendor. The price applied is the one the serving provider charges for that model; you can read it from `GET /api/v1/models`, where each entry carries a `pricing` object with `prompt` and `completion` rates plus optional `request`, `image`, `web_search` and `internal_reasoning` components. The values are decimal **strings in USD per single token**, so a `"0.000003"` prompt rate is $3 per million tokens — multiplying by 1,000,000 is the usual first step. The same model slug can be served by several providers at different rates, listed per provider under the model's endpoints. Your remaining balance and lifetime spend come from `GET /api/v1/credits`, which returns `total_credits` and `total_usage`; OpenRouter's own margin is documented as a fee on credit purchases rather than a per-token markup on inference.
code
python · 9 linesimport requests
from decimal import Decimal
models = requests.get("https://openrouter.ai/api/v1/models", timeout=30).json()["data"]
for m in models[:5]:
p = m["pricing"]
per_m_in = Decimal(p["prompt"]) * 1_000_000
per_m_out = Decimal(p["completion"]) * 1_000_000
print(f'{m["id"]}: ${per_m_in}/M in, ${per_m_out}/M out')go deeper
Be able to say that OpenRouter uses one prepaid credit balance across every model, and that model prices come from the models endpoint quoted per single token in dollars.
Explain the pricing object field by field, convert per-token strings to per-million figures without slipping a factor of a million, and know that per-provider endpoints for one slug can differ in price.
Show how you keep a live view of spend and solvency: alarm on the credits endpoint, record the reported cost per request, and never reconstruct billing from a cached price table that drifts when routing changes.
Own the buy-versus-integrate argument. Frame the gateway's credit-purchase fee against the cost of maintaining several vendor contracts, keys and invoices, and decide when a workload is large enough to move onto a direct provider agreement.
## One balance in front of many vendors The commercial point of OpenRouter is that you sign one payment relationship instead of sixteen. You buy **credits** (1 credit = 1 USD of inference), and every call to `POST /api/v1/chat/completions` — whichever vendor's model slug you name — is metered against that one balance. You do not hold an OpenAI account, an Anthropic account and a Mistral account; you hold credits. That is the whole reason a team reaches for a gateway when they want to try five models next week without five procurement conversations. ## Where the prices live Prices are data, not documentation. `GET https://openrouter.ai/api/v1/models` returns the catalogue, and each model object carries a `pricing` object. The fields you will actually use are: - `prompt` — cost per input token - `completion` — cost per output token - `request` — a flat per-call charge, `"0"` for most models - `image` — cost per image input, for multimodal models - `web_search`, `internal_reasoning` — charged only when those features are exercised Cache-related rates appear alongside them for models that support prompt caching, so a cached read can be priced differently from a fresh input token. ## Per token, not per million — the classic misread Vendors advertise "$3 per million input tokens". OpenRouter's API reports **per single token, as a decimal string**: `"0.000003"`. Two mistakes follow from ignoring that. The first is arithmetic — a cost estimator that treats the number as per-million is off by a factor of a million and will happily tell you a chat costs three cents when it costs three hundredths of a cent, or vice versa. The second is type — the values are strings, so naive addition in JavaScript or Python concatenates or raises instead of summing. Parse to a decimal type before doing money maths; floating point across a million-token accumulation is exactly the place where cents go missing. ## Same model, different providers, different prices A slug like `vendor/model-name` can be served by several upstream providers, each with its own price, context limit and quantisation. The per-provider rates are exposed on the model's endpoints listing, and this is why a cost model built from a single headline number drifts: the price you actually paid depends on which provider served the request, which can change between calls. If you need to know what a specific call cost, do not recompute it from a price table — read the cost the platform reports for that request. ## Checking the balance `GET /api/v1/credits` returns `data.total_credits` (everything you have ever purchased) and `data.total_usage` (everything you have ever spent); the remaining balance is the difference. `GET /api/v1/auth/key` reports the calling key's own `label`, `usage` and `limit`. Both are cheap enough to poll from a monitoring job, and a balance alarm is the difference between a graceful top-up and a Friday-night outage where every call fails with an insufficient-credits error. ## Where the margin comes from OpenRouter documents its own revenue as a fee taken when you **buy credits**, not as a markup on each token. Practically, the per-token rate you see for a model is the provider's rate, and your effective cost is that rate plus the purchase fee amortised across everything you spend. This matters when you compare a gateway against calling a vendor directly: the comparison is not "gateway tax per token" but "purchase fee versus the engineering cost of maintaining N vendor integrations, N sets of credentials and N billing relationships". ## What this means in practice Build your cost view on three things: the catalogue for planning ("what would this workload cost on each candidate model"), the per-request cost the API reports for accounting ("what did this call actually cost"), and the credits endpoint for solvency ("will the next call succeed"). Teams that only build the first end up with a beautiful spreadsheet and no idea why the balance drained overnight.
- Your cost dashboard is off by a factor of a million against the real balance drain. What is the first thing you check?The unit on the pricing fields. OpenRouter quotes `pricing.prompt` and `pricing.completion` in USD per single token as decimal strings, while vendor marketing pages quote dollars per million tokens. A dashboard that copies the marketing figure into a per-token formula, or feeds the API string into a per-million formula, produces exactly that error. Parse the strings as decimals and multiply by 1,000,000 only when rendering.
- Why can two calls to the same model slug on the same day cost different amounts per token?Because the slug can be served by more than one upstream provider, and each provider sets its own price for that model — they are listed separately on the model's endpoints. Which provider serves a given call depends on availability and your routing preferences, so the effective rate is a property of the route, not just the slug. Read the reported cost per request rather than recomputing from one headline price.
- Does OpenRouter charge you again for the tokens on top of the provider's price?Not as a per-token markup. OpenRouter documents its margin as a fee applied when you purchase credits; the per-token rates shown for a model are the serving provider's own rates. So the right comparison against calling a vendor directly is the purchase fee plus the operational savings of one key and one bill, not a per-token gateway tax.
Credits work like a prepaid transit card that is accepted on every operator's line: you top up one card, and each ride is deducted at whatever that operator charges.
saying these in an interview costs you the question
- Reads pricing.prompt as dollars per million tokens
- Assumes every model costs the same through the gateway
- Thinks you need a separate account per upstream vendor
- Adds pricing strings without parsing them as decimals
- Believes the catalogue price is authoritative for a given call