skip to content

OpenRouter

OpenRouter puts one API key and one request shape in front of hundreds of models from dozens of providers. Interviews use it to probe how you would keep an app answering when a single vendor degrades or gets expensive.

on this pageshow

questions

17

How does OpenRouter bill a call, and where do its per-model prices come from?

level: juniorimportance: must knowfreq 55%

answer

  1. One prepaid balance, many vendors
  2. Prices ship as API data
  3. Watch the unit on pricing
  4. Per single token, decimal string
  5. GET /api/v1/credits for balance

basics

~10 s

OpenRouter runs on prepaid credits: you top up one balance and each call deducts the chosen provider's own per-token price. GET /api/v1/models publishes every model's pricing object, quoted in US dollars per single token.

solid answer

~50 s

You buy credits on OpenRouter once, and every request against any of the hundreds of models draws down that single balance — there is no separate account or invoice per vendor. The price applied is the one the serving provider charges for that model; you can read it from `GET /api/v1/models`, where each entry carries a `pricing` object with `prompt` and `completion` rates plus optional `request`, `image`, `web_search` and `internal_reasoning` components. The values are decimal **strings in USD per single token**, so a `"0.000003"` prompt rate is $3 per million tokens — multiplying by 1,000,000 is the usual first step. The same model slug can be served by several providers at different rates, listed per provider under the model's endpoints. Your remaining balance and lifetime spend come from `GET /api/v1/credits`, which returns `total_credits` and `total_usage`; OpenRouter's own margin is documented as a fee on credit purchases rather than a per-token markup on inference.

code

python · 9 lines
python
import requests
from decimal import Decimal

models = requests.get("https://openrouter.ai/api/v1/models", timeout=30).json()["data"]
for m in models[:5]:
    p = m["pricing"]
    per_m_in = Decimal(p["prompt"]) * 1_000_000
    per_m_out = Decimal(p["completion"]) * 1_000_000
    print(f'{m["id"]}: ${per_m_in}/M in, ${per_m_out}/M out')

go deeper

for a junior

Be able to say that OpenRouter uses one prepaid credit balance across every model, and that model prices come from the models endpoint quoted per single token in dollars.

for a middle

Explain the pricing object field by field, convert per-token strings to per-million figures without slipping a factor of a million, and know that per-provider endpoints for one slug can differ in price.

for a senior

Show how you keep a live view of spend and solvency: alarm on the credits endpoint, record the reported cost per request, and never reconstruct billing from a cached price table that drifts when routing changes.

for a principal

Own the buy-versus-integrate argument. Frame the gateway's credit-purchase fee against the cost of maintaining several vendor contracts, keys and invoices, and decide when a workload is large enough to move onto a direct provider agreement.

## One balance in front of many vendors The commercial point of OpenRouter is that you sign one payment relationship instead of sixteen. You buy **credits** (1 credit = 1 USD of inference), and every call to `POST /api/v1/chat/completions` — whichever vendor's model slug you name — is metered against that one balance. You do not hold an OpenAI account, an Anthropic account and a Mistral account; you hold credits. That is the whole reason a team reaches for a gateway when they want to try five models next week without five procurement conversations. ## Where the prices live Prices are data, not documentation. `GET https://openrouter.ai/api/v1/models` returns the catalogue, and each model object carries a `pricing` object. The fields you will actually use are: - `prompt` — cost per input token - `completion` — cost per output token - `request` — a flat per-call charge, `"0"` for most models - `image` — cost per image input, for multimodal models - `web_search`, `internal_reasoning` — charged only when those features are exercised Cache-related rates appear alongside them for models that support prompt caching, so a cached read can be priced differently from a fresh input token. ## Per token, not per million — the classic misread Vendors advertise "$3 per million input tokens". OpenRouter's API reports **per single token, as a decimal string**: `"0.000003"`. Two mistakes follow from ignoring that. The first is arithmetic — a cost estimator that treats the number as per-million is off by a factor of a million and will happily tell you a chat costs three cents when it costs three hundredths of a cent, or vice versa. The second is type — the values are strings, so naive addition in JavaScript or Python concatenates or raises instead of summing. Parse to a decimal type before doing money maths; floating point across a million-token accumulation is exactly the place where cents go missing. ## Same model, different providers, different prices A slug like `vendor/model-name` can be served by several upstream providers, each with its own price, context limit and quantisation. The per-provider rates are exposed on the model's endpoints listing, and this is why a cost model built from a single headline number drifts: the price you actually paid depends on which provider served the request, which can change between calls. If you need to know what a specific call cost, do not recompute it from a price table — read the cost the platform reports for that request. ## Checking the balance `GET /api/v1/credits` returns `data.total_credits` (everything you have ever purchased) and `data.total_usage` (everything you have ever spent); the remaining balance is the difference. `GET /api/v1/auth/key` reports the calling key's own `label`, `usage` and `limit`. Both are cheap enough to poll from a monitoring job, and a balance alarm is the difference between a graceful top-up and a Friday-night outage where every call fails with an insufficient-credits error. ## Where the margin comes from OpenRouter documents its own revenue as a fee taken when you **buy credits**, not as a markup on each token. Practically, the per-token rate you see for a model is the provider's rate, and your effective cost is that rate plus the purchase fee amortised across everything you spend. This matters when you compare a gateway against calling a vendor directly: the comparison is not "gateway tax per token" but "purchase fee versus the engineering cost of maintaining N vendor integrations, N sets of credentials and N billing relationships". ## What this means in practice Build your cost view on three things: the catalogue for planning ("what would this workload cost on each candidate model"), the per-request cost the API reports for accounting ("what did this call actually cost"), and the credits endpoint for solvency ("will the next call succeed"). Teams that only build the first end up with a beautiful spreadsheet and no idea why the balance drained overnight.

  • Your cost dashboard is off by a factor of a million against the real balance drain. What is the first thing you check?
    The unit on the pricing fields. OpenRouter quotes `pricing.prompt` and `pricing.completion` in USD per single token as decimal strings, while vendor marketing pages quote dollars per million tokens. A dashboard that copies the marketing figure into a per-token formula, or feeds the API string into a per-million formula, produces exactly that error. Parse the strings as decimals and multiply by 1,000,000 only when rendering.
  • Why can two calls to the same model slug on the same day cost different amounts per token?
    Because the slug can be served by more than one upstream provider, and each provider sets its own price for that model — they are listed separately on the model's endpoints. Which provider serves a given call depends on availability and your routing preferences, so the effective rate is a property of the route, not just the slug. Read the reported cost per request rather than recomputing from one headline price.
  • Does OpenRouter charge you again for the tokens on top of the provider's price?
    Not as a per-token markup. OpenRouter documents its margin as a fee applied when you purchase credits; the per-token rates shown for a model are the serving provider's own rates. So the right comparison against calling a vendor directly is the purchase fee plus the operational savings of one key and one bill, not a per-token gateway tax.

Credits work like a prepaid transit card that is accepted on every operator's line: you top up one card, and each ride is deducted at whatever that operator charges.

saying these in an interview costs you the question

  • Reads pricing.prompt as dollars per million tokens
  • Assumes every model costs the same through the gateway
  • Thinks you need a separate account per upstream vendor
  • Adds pricing strings without parsing them as decimals
  • Believes the catalogue price is authoritative for a given call

context

open as a page

How do you point the OpenAI SDK at OpenRouter, and how are models named there?

level: juniorimportance: must knowfreq 78%

basics

~10 s

Set the OpenAI client's base URL to https://openrouter.ai/api/v1 and send your OpenRouter key as a bearer token. Models are addressed by a vendor-prefixed slug such as anthropic/claude-sonnet-4.5, so one key reaches every vendor.

open as a page

How do you make OpenRouter return the actual cost and native token counts inline?

level: middleimportance: must knowfreq 50%

basics

~20 s

Send "usage": {"include": true} in the chat completion request body. The response's usage object then carries cost in credits alongside token counts from the serving model's own tokenizer, with cost_details breaking out any upstream charge.

open as a page

In OpenRouter, what does the models array in a chat completions request do?

level: middleimportance: must knowfreq 62%

basics

~20 s

OpenRouter treats models as an ordered fallback list: it tries the first entry and, if that model errors or is unavailable, retries the next one inside the same request. You are billed for whichever model actually answered.

open as a page

How do OpenRouter's provider.order and allow_fallbacks fields control routing?

level: middleimportance: must knowfreq 55%

basics

~20 s

provider.order lists upstream providers to try in priority order for the chosen model. allow_fallbacks, true by default, decides what happens when none of them can serve: leave it on and OpenRouter tries other providers, set it false and the request fails instead.

open as a page

On OpenRouter, what happens to request parameters the chosen model doesn't support?

level: middleimportance: must knowfreq 55%

basics

~20 s

By default they are dropped, not rejected: OpenRouter forwards only the parameters the target model honours and returns a normal completion. The danger is silence — your setting simply had no effect, so verify support against the model catalog instead of assuming.

open as a page

In OpenRouter, what do the :nitro and :floor model-slug suffixes do?

level: juniorimportance: should knowfreq 45%

basics

~20 s

They are shorthand for a provider-sorting rule on the model slug. Appending :nitro ranks the providers hosting that model by throughput, so the request goes to the fastest; :floor ranks them by price, so it goes to the cheapest.

open as a page

What does OpenRouter's GET /api/v1/generation endpoint tell you about a finished call?

level: middleimportance: should knowfreq 32%

basics

~10 s

Called with the completion's id, it returns that generation's audit record: total_cost, native and normalised token counts, which upstream provider served it, latency and generation time, finish reason and any cache discount.

open as a page

What do OpenRouter's optional HTTP-Referer and X-Title request headers do?

level: middleimportance: should knowfreq 40%

basics

~20 s

They are optional attribution headers: HTTP-Referer carries your app's URL and X-Title its display name, so requests are credited to your app on OpenRouter's public model rankings and in your own dashboard. They are not authentication and do not affect routing.

open as a page

OpenRouter starts returning 402 on some calls and 429 on others — what differs?

level: seniorimportance: should knowfreq 45%

basics

~20 s

402 means the OpenRouter credit balance cannot cover the call — no amount of retrying helps, only a top-up. 429 means a rate limit was hit, on OpenRouter's free-tier caps or an upstream provider, and it clears with time.

open as a page

With BYOK on OpenRouter, who bills the inference and what still leaves your credits?

level: seniorimportance: should knowfreq 26%

basics

~20 s

Under bring-your-own-key, the upstream provider authenticates with your key and invoices you directly for the tokens. OpenRouter still deducts a percentage fee from your credits per request, so a positive credit balance is still required.

open as a page

What does require_parameters do in an OpenRouter provider block?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Set to true, it restricts routing to upstream providers that support every parameter in your request. It defaults to false, which lets OpenRouter route to a provider that quietly ignores a parameter such as tools or a JSON schema, so you get a valid-looking answer in the wrong shape.

open as a page

How does OpenRouter map an OpenAI-shaped request onto a non-OpenAI vendor's API?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The gateway rewrites the request into the vendor's own shape — hoisting system messages, converting tool calls and results into that vendor's block format, supplying required fields the vendor demands — then rewrites the reply back into OpenAI's choices/message/finish_reason structure.

open as a page

In OpenRouter, how do you decide how far a production workload may fall back?

level: principalimportance: should knowfreq 32%

basics

~20 s

Classify traffic by what a wrong-shaped answer costs. Let conversational paths float widely across providers and models; keep structured and regulated paths on a short, evaluated list with fallbacks disabled, so they fail loudly instead of degrading invisibly.

open as a page

When is OpenRouter's unified API the wrong abstraction for a production app?

level: principalimportance: should knowfreq 30%

basics

~20 s

When the app has settled on one vendor and depends on that vendor's deep features, contracts or compliance terms. A common-denominator gateway lags vendor launches, adds a hop and a dependency, and cannot express what the shared shape has no field for.

open as a page

In OpenRouter, what does the provider.data_collection setting control?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

It is a routing filter on provider data policy. Left at its default of allow, any upstream is eligible; set to deny, OpenRouter routes only to providers whose published policy says they do not collect or retain your prompts for their own use.

open as a page

How would you attribute and cap OpenRouter spend per team on one shared account?

level: principalimportance: nice to knowfreq 22%

basics

~10 s

Issue a separate OpenRouter API key per team through the key-provisioning API, each with its own credit limit, and log the per-request cost from usage accounting against your own tenant and feature dimensions.

open as a page