When is OpenRouter's unified API the wrong abstraction for a production app?
answer
- Breadth versus depth
- Common shape cannot express everything
- Frontier features arrive natively first
- Extra hop, extra failure domain
- Contracts and compliance attach to vendors
basics
~20 sWhen the app has settled on one vendor and depends on that vendor's deep features, contracts or compliance terms. A common-denominator gateway lags vendor launches, adds a hop and a dependency, and cannot express what the shared shape has no field for.
solid answer
~50 sThe gateway earns its keep when breadth is the requirement: evaluating many models, serving heterogeneous workloads, or keeping the option to move. It stops earning it once the app converges on one vendor and starts depending on things the common shape cannot express — newly launched capabilities that the gateway has not surfaced yet, fine-grained vendor-only controls, or response metadata that normalisation flattens. There are non-technical reasons too: you now have a second party in the request path and in your availability budget, and enterprise requirements such as data-processing agreements, retention terms, regional residency or a direct support relationship are usually negotiated with the model vendor, not through an intermediary. A reasonable posture is to keep the OpenAI-shaped call behind your own thin interface, so the gateway is a swappable implementation — and to move the one workload that has outgrown it to a native SDK rather than abandoning the abstraction wholesale.
go deeper
Be able to state the basic trade: one key and one request shape across many models, at the cost of only reaching features that all those models roughly share.
Give concrete losses — vendor features with no field in the common shape, lag before new capabilities are exposed, normalised metadata that hides detail — rather than a vague lock-in argument.
Bring operations in: the added hop in the latency and availability budget, distinguishing gateway failures from upstream ones in monitoring, and keeping provider access behind one internal interface.
Own the decision rule and its triggers. Say when a workload graduates to a native integration, who signs the data-processing and residency terms, and why the migration is per workload rather than a wholesale rewrite.
## What the abstraction actually buys A unified, OpenAI-shaped endpoint in front of many vendors gives you four concrete things: one credential instead of a dozen, one request shape so model choice becomes configuration, breadth for evaluation, and one integration to maintain instead of several. For early-stage products, for teams comparing models, and for applications that legitimately route different workloads to different models, that is a large win and it is why the pattern spread. ## The structural costs **Lowest common denominator.** A shared shape can only express what all vendors roughly share. Anything genuinely vendor-specific either gets a gateway-defined extension field or is simply unreachable. That is fine for chat, tools and streaming; it is limiting as soon as your value depends on a feature that lives outside the intersection. **Feature lag.** Vendors ship first into their own SDKs. The gateway then has to model the new capability, decide how it fits the unified shape, and expose it. That delay is unavoidable and it is worst exactly where you feel it most — at the frontier, where the new capability is why you wanted the model. **A hop and a dependency.** Every request now traverses an additional service. That is added latency, an added failure domain and an added attack surface, and your uptime is now a product of two availabilities rather than one. It also means your prompts — often the most sensitive payload your company sends anywhere — pass through a third party. **Flattened observability.** Normalisation collapses vendor-specific stop reasons and metadata into a common vocabulary. Good for portable code, worse for forensic debugging: the signal that would have explained a behaviour may not survive the mapping. **Commercial and compliance shape.** Enterprise buyers negotiate directly: data-processing terms, zero-retention commitments, regional data residency, capacity guarantees, security review, incident escalation. Those attach to a vendor relationship. An intermediary can offer its own terms, but it cannot hand you the vendor's, and for regulated workloads that is frequently the decisive argument. ## Signals that a workload has outgrown it - Your traffic is overwhelmingly one model, and has been for months. - You are writing workarounds for a capability the vendor supports natively. - Procurement, security or legal needs a contract with the model vendor specifically. - Latency budgets are tight enough that an extra network hop is material. - You need capacity commitments or throughput guarantees no reseller can give. - Debugging repeatedly stalls because the information you need was normalised away. ## Signals that it is still right - You are still choosing, and the cost of choosing wrongly is a rewrite. - Different workloads genuinely want different models — cheap and fast here, deep and expensive there. - You need optionality against a single vendor degrading, repricing or deprecating a model. - Your usage is small enough that per-vendor contracts are not worth the overhead. - You are shipping a product that lets *users* pick a model. ## The design that keeps both options open Do not let vendor choice be an architectural commitment. Put a narrow internal interface in front of generation — a function that takes messages, tools and a policy and returns a normalised result — and let both the gateway and a native SDK be implementations of it. That costs a little indirection and buys you the ability to move one workload at a time. It is also the only structure in which "we use the gateway" and "we call the vendor directly" can be true simultaneously, which is usually where mature systems land: the gateway for breadth, experimentation and long-tail models, a native integration for the one high-volume, deep-feature path that pays for its own code. ## Framing it in an interview The weak answer picks a side. The strong answer is a decision rule: use the unified API while breadth and optionality are worth more than depth, and migrate a specific workload — not the whole app — once its dependence on one vendor's capabilities, contracts or latency budget outweighs what a common shape can give it. Name the observable triggers, and name the interface that makes the migration a swap rather than a rewrite.
- How would you structure code so that leaving the gateway later is not a rewrite?Put a narrow internal interface in front of generation: messages, tools and a policy in, a normalised result out, with your own type for the answer. The gateway and a vendor's native SDK then become two implementations you can select per workload. The rule is that no application code imports a provider client directly, so switching one path is a wiring change rather than a diff across the codebase.
- What is the operational risk of putting a gateway in the request path?Your availability becomes the product of two services rather than one, and every call carries an extra hop of latency. You also inherit the gateway's incident surface and its own upstream failures. Budget for it explicitly: measure the added latency, decide whether a direct path is needed for the most latency-sensitive workload, and make sure your monitoring can distinguish a gateway failure from an upstream model failure.
- Does moving off the gateway mean moving everything?No, and treating it as all-or-nothing is the common mistake. Mature systems usually run both: a native integration for the one high-volume path that depends on a vendor's deep features and contractual terms, and the gateway for breadth, experimentation, and long-tail or user-selected models. Migrate per workload, triggered by an observable — traffic concentration, a needed feature, a compliance requirement — rather than by preference.
saying these in an interview costs you the question
- Claiming a gateway removes all vendor lock-in
- Ignoring the extra hop in latency and availability budgets
- Assuming vendor compliance terms pass through an intermediary
- Expecting brand-new vendor features to appear immediately
- Treating the choice as all-or-nothing for the whole system