In OpenRouter, what does the models array in a chat completions request do?
answer
- ordered preference, not an ensemble
- one HTTP response, gateway-side retry
- the route fallback flag pairs with it
- billed for whoever actually answered
- response model field names the winner
basics
~20 sOpenRouter treats models as an ordered fallback list: it tries the first entry and, if that model errors or is unavailable, retries the next one inside the same request. You are billed for whichever model actually answered.
solid answer
~50 sInstead of a single `model` slug you can send `models: ["primary", "secondary", "tertiary"]` (documented alongside `route: "fallback"`). OpenRouter attempts them in order and, when an attempt fails — the model is down, rate-limited, or otherwise cannot serve the request — it moves to the next entry and returns the first successful completion. The client sees one HTTP response, not a retry loop, so a vendor outage does not surface to your users as an error. Two consequences matter in production. First, billing follows the model that actually served, so a fallback to a pricier model quietly changes your unit cost. Second, the response's `model` field reports which slug answered (OpenRouter also surfaces the serving upstream provider), so you must log it — otherwise you cannot tell that half your traffic silently moved to a different model with different behaviour.
code
json · 11 lines{
"models": [
"vendor-a/model-primary",
"vendor-b/model-backup",
"vendor-c/model-last-resort"
],
"route": "fallback",
"messages": [
{ "role": "user", "content": "Summarise this ticket in one sentence." }
]
}go deeper
Know that OpenRouter takes a list of model slugs and tries them in order, so one vendor being down does not turn into an error for your users. Be able to say that only one model's answer comes back.
Explain when the next entry is tried (the attempt cannot serve, not that the answer was bad), that it is one HTTP request from the client's side, and that billing follows the model that actually served.
Show the operational side: log the response model, alert on fallback rate, and treat every entry in the chain as something your prompts and parsers were evaluated against — including its tool-calling and context limits.
Own the policy question of which workloads may silently degrade to a different model at all, and which must fail loudly instead; set the cost and behavioural blast radius the chain is allowed to have.
## What the field is OpenRouter's chat-completions body normally carries a single `model` slug of the form `vendor/model-name`. The `models` field replaces it with an **ordered array of slugs**. The documented pairing is `models: [...]` together with `route: "fallback"`, which states explicitly that the array is a fallback chain rather than an ensemble. OpenRouter starts at index 0 and walks forward until one model produces a completion. This is *model-level* fallback. It is distinct from *provider-level* fallback, which OpenRouter does by default for a single model that several upstream providers host: there it can move between providers serving the identical weights. The `models` array is what you reach for when you want to cross to a **different model**, possibly from a different vendor entirely. ## When the fallback fires The next entry is tried when the current attempt cannot produce a completion — the model or all of its providers are unavailable, the request is rate-limited upstream, or the attempt errors out. It is not a quality judgement: OpenRouter has no idea whether the answer was good, so it never falls through because the first model produced a poor answer. It also does not run models in parallel and pick a winner; exactly one model's output is returned. ## What it does not do It does not rewrite your request for the next model. The same messages, the same `tools`, the same `response_format` are replayed. If the fallback model lacks a capability the primary had — tool calling, structured output, a long enough context window, image input — that attempt can fail too, or it can succeed while ignoring the parameter and returning ordinary prose where your code expected a tool call. This is why a fallback list should be *evaluated*, not just typed: every entry must be a model your prompt and your parser actually work with. Pairing the array with `provider.require_parameters` helps keep capability, but it cannot make a model support a feature it does not have. It is also weak protection once bytes are on the wire. Fallback substitutes an attempt that has not yet produced output; after a streaming response has begun emitting tokens, OpenRouter cannot retroactively swap models, so a mid-stream failure surfaces to your client as an error in the stream and you must decide client-side whether to retry or salvage the partial text. ## Billing and observability You pay the rate of the model that served, not the one you listed first. A cheap primary with an expensive fallback therefore has a cost profile that depends on the primary's uptime — a detail that only shows up on the invoice unless you instrument it. Two habits make that visible: log the `model` value from every response (it names the slug that answered, and OpenRouter additionally reports the serving provider), and alert on the *fallback rate* — the share of requests answered by anything other than index 0. A rising fallback rate is an early warning about your primary vendor that no client-side error metric will show you, precisely because the gateway absorbed the failure. ## Ordering strategy Order the array by decreasing preference, not by decreasing price. Typical shapes: - **Same family, different sizes** — the closest behavioural match, so prompts and parsers usually carry over. - **Different vendors, similar class** — real vendor-outage insurance, but expect tone and formatting differences; keep the list short and evaluated. - **Premium primary, cheap safety net** — acceptable for chat, dangerous for extraction pipelines whose downstream code assumes a schema. For workloads where a wrong-model answer is worse than no answer at all, do not build a long chain. Use one or two vetted entries and let the request fail loudly. ## Common mistakes Assuming the array means "try all and pick the best"; assuming the first model's price is what you pay; assuming a fallback rescues a stream that already started; and shipping a chain whose later entries were never run against the prompt suite. Each one is a production incident waiting for the primary vendor's next bad afternoon.
- If the second model in the list answers, how would your service know?Read the `model` field on the response: it reports the slug that actually served, not the one you asked for first, and OpenRouter also reports the upstream provider that handled it. Log both per request and track the share of responses that did not come from index 0. That fallback rate is your outage signal and your cost-drift signal; without it, a vendor degradation is invisible because the gateway already absorbed the error.
- Does the fallback list protect a streaming request that fails halfway through?No. Fallback replaces an attempt that has not yet produced output. Once a stream has started emitting tokens, the model that answered is committed, and a failure after that point arrives as an error inside the stream. Your client has to decide whether to discard the partial text and re-issue the request or to present what it has. Treat model fallback as protection against failure-to-start, not against failure-mid-generation.
- How is this different from OpenRouter's default behaviour with a single model slug?With one slug, OpenRouter still fails over — but only between the upstream providers that host that same model, load-balancing across them by price and observed reliability. The weights are identical, so behaviour is stable. The `models` array is the only mechanism that crosses to a *different* model, which is why it buys wider outage coverage and costs you behavioural consistency.
saying these in an interview costs you the question
- Thinks the models array runs every model and picks the best answer
- Assumes you are billed for the first model even when a later one answered
- Believes fallback can rescue a stream that has already emitted tokens
- Expects a prompt tuned for the primary to behave identically on the backup
- Confuses model fallback with provider-level failover for one model