skip to content

In OpenRouter, how do you decide how far a production workload may fall back?

level: principalimportance: should knowfreq 32%

answer

  1. classify by cost of a wrong-shaped answer
  2. provider swap is cheap, model swap is not
  3. the risk is substitution, not downtime
  4. bound capability, price and perimeter
  5. fallback rate belongs on a dashboard

basics

~20 s

Classify traffic by what a wrong-shaped answer costs. Let conversational paths float widely across providers and models; keep structured and regulated paths on a short, evaluated list with fallbacks disabled, so they fail loudly instead of degrading invisibly.

solid answer

~50 s

There is no single correct fallback depth — the decision is per workload class. Fallback between the providers hosting one model is nearly free: same weights, so widen it and let the gateway absorb outages. Fallback across *models* changes behaviour, so it is a product decision, not an infrastructure one: every entry in a `models` array needs to pass the same evals, prompts and parsers as the primary. Then bound the blast radius: `require_parameters` so capability survives the reroute, a `max_price` ceiling so an outage cannot quietly multiply unit cost, `only`/`order` with `allow_fallbacks: false` where routing outside an approved set is worse than failing. Finally, make it observable — log the model and serving provider on every response, alert on fallback rate, and treat a rising rate as a vendor incident. The failure mode a gateway introduces is not downtime, it is *silent substitution*, and the policy exists to decide where that is acceptable.

go deeper

for a junior

Know that a gateway can quietly answer with a different provider or model, and that your logs should say which one actually served the request.

for a middle

Explain the difference in risk between swapping providers for one model and swapping to a different model, and name the fields that constrain each.

for a senior

Set policy per call site rather than per service, bound capability, price and perimeter explicitly, and stand up fallback-rate and cost-per-served-model monitoring.

for a principal

Own the classification itself — which traffic may silently degrade and which must fail — quantify the availability given up by each fence, and put the fallback list on a review cadence tied to evals and incidents.

## Reframing the question Asked as "how much fallback should we enable", this has no answer. Asked as "what is the cost of being answered by something other than what we asked for", it has a different answer per traffic class, and that is the framing a principal is expected to bring. A gateway converts an availability problem into a **substitution** problem. Without it, a vendor outage is an error your monitoring already understands. With it, the outage becomes a different model or a different provider answering, at a different price, in a slightly different style, possibly without a capability you depended on — and returning HTTP 200 throughout. Every policy decision below is about where that substitution is a gift and where it is a defect. ## Two layers, two different risk profiles **Provider-level fallback** (several upstreams hosting one model) is cheap. The weights are the same; differences are price, throughput, quantization, usable context and supported parameters. For almost all traffic the right posture is permissive: let OpenRouter balance and fail over, and constrain only where you have a specific reason. **Model-level fallback** (a `models` array) is expensive in a way that does not show on a dashboard. A different model has different tool-calling fidelity, different adherence to output schemas, different refusal behaviour, different prompt sensitivity. Adding an entry to that array is shipping a second model to production; it deserves the same evaluation gate as the first. A list nobody evaluated is worse than no list, because it converts a visible outage into an invisible quality regression. ## Classify the traffic A workable taxonomy: - **Conversational / summarisation.** Any coherent answer is acceptable. Widest policy: provider fallback on, a vetted alternative model behind the primary, no hard fences. Availability wins. - **Structured / extraction / agentic.** Downstream code parses the output. `require_parameters: true` is mandatory; the model fallback list must be short and evaluated against the same schema; validate the response shape and fail the request rather than pass malformed data on. - **Regulated / confidential.** Where the request goes is itself a controlled property. Explicit `only`/`order`, `data_collection: "deny"` where relevant, and `allow_fallbacks: false` so exhaustion is an error you handle, not a reroute you never notice. - **Batch / offline.** Latency-insensitive and retryable in your own scheduler. Price-first routing is appropriate; wide fallback is fine because you can re-run. The useful discipline is that these are properties of the *call site*, not of the service. One codebase can and should send different routing policies from different endpoints. ## Bounding the blast radius For each class, three bounds are worth setting explicitly: 1. **Capability** — `require_parameters` so a reroute cannot drop tools or structured output. 2. **Cost** — a `max_price` ceiling on prompt and completion rates, so a failover to a premium upstream cannot multiply unit economics during exactly the hour you are least able to watch it. 3. **Perimeter** — `only`/`order` plus `allow_fallbacks: false` where the set of acceptable counterparties is finite. Each bound trades availability for control. Say the number out loud: if the intersection of your filters is two providers, your effective availability is theirs, and the mitigation is a *broader approved set*, not a looser policy. ## Observability is the deliverable The control that matters most is not a request field, it is the log line. For every response record the model that served, the provider that served, the latency and the cost. From those, three signals: - **Fallback rate** — the share of requests not served by the primary. A rising rate is a vendor incident happening quietly; it should page before your users notice quality drift. - **Cost per request by served model** — catches the expensive-backup problem the moment it starts, not at month end. - **Shape-violation rate by provider** — a schema failure that correlates with one upstream is a capability gap you can fix with `ignore` or `require_parameters`. Without these, a fallback policy is an assertion. With them, it is a control you can review after an incident and adjust. ## When to prefer failing The hardest part of the policy is granting permission to fail. For workloads where a wrong answer propagates — money movement, compliance classification, anything writing to a system of record — an error is cheaper than a plausible answer from an unevaluated model. Encoding that means `allow_fallbacks: false`, a short model list or none at all, and an upstream caller that can degrade gracefully. Teams under-use this because failing feels like losing; the framing that lands is that you are choosing *which* failure you get. ## Review cadence Provider mixes, prices and model line-ups all move. A fallback list is a decision with an expiry date: re-run the evals when you add an entry, when a listed model's generation is superseded, and after any incident where fallback fired. Otherwise the list quietly becomes a set of models nobody has tested against the current prompts.

  • What single metric best tells you a fallback policy is doing something you did not intend?
    Fallback rate — the share of responses not served by the primary model and provider. It is derivable from the model and provider reported on every response, and it moves before user-visible quality does. A step change means an upstream is degrading; a slow drift means the provider mix or your load has shifted. Pair it with cost per request by served model so the economic consequence surfaces at the same time.
  • How do you justify allow_fallbacks:false to a team that measures itself on availability?
    Reframe it as choosing which failure you take. For traffic that writes to a system of record, a plausible answer from an unevaluated model is a data-integrity incident with a long tail, while an error is a retry. Quantify both: the availability you give up is measurable and bounded, the cost of a wrong write usually is not. Then recover availability by broadening the *approved* set rather than by loosening the fence.
  • What makes a model fallback list safe to add an entry to?
    Treat it as shipping a second model. The candidate must pass the same evaluation suite, the same prompts, the same output-schema validation and the same tool definitions as the primary, and the result must be recorded so the decision is reviewable. Re-run that gate when a listed model's generation is superseded or after any incident where the fallback actually fired, otherwise the list decays into untested models nobody owns.

saying these in an interview costs you the question

  • Treats more fallback as strictly better regardless of workload
  • Adds models to a fallback list without running them through evals
  • Assumes cost is unchanged when a backup provider or model serves
  • Never logs the serving model and provider, so substitution is invisible
  • Refuses to let any workload fail, even ones that write to a system of record

context