How does OpenRouter map an OpenAI-shaped request onto a non-OpenAI vendor's API?
answer
- Gateway rewrites both directions
- System prompt does not always live in messages
- Some vendors demand fields OpenAI makes optional
- Tool calls become blocks and back
- Stop reasons folded into stop/length/tool_calls
basics
~20 sThe gateway rewrites the request into the vendor's own shape — hoisting system messages, converting tool calls and results into that vendor's block format, supplying required fields the vendor demands — then rewrites the reply back into OpenAI's choices/message/finish_reason structure.
solid answer
~50 sTake an Anthropic model as the example, since its Messages API is structurally different from OpenAI's. Your `messages` array may carry `role: "system"` entries, but Anthropic takes the system prompt as a separate top-level parameter, so the gateway lifts them out. Anthropic requires `max_tokens` on every call, so when you omit it the gateway must supply a value rather than pass the omission through. An assistant turn containing OpenAI-style `tool_calls`, and the `role: "tool"` message replying to it, are converted into the vendor's tool-use and tool-result content blocks and back again. On the way out, the vendor's stop reason is mapped into OpenAI's vocabulary — normal completion becomes `stop`, hitting the cap becomes `length`, wanting a tool becomes `tool_calls` — and the vendor's token counts are mapped into `usage`. The translation is faithful for mainstream chat, but anything vendor-specific that has no OpenAI equivalent has nowhere to land in the standard shape.
go deeper
It is enough to know the gateway does not forward your JSON as-is: it rewrites the request into the vendor's format and rewrites the answer back into the OpenAI-style choices and message you parse.
Name concrete transformations — system prompt moved out of the messages array, tool calls converted to the vendor's block format, stop reasons folded into stop, length and tool_calls — rather than saying it just translates.
Demonstrate debugging instinct: set max_tokens explicitly so no default is chosen for you, log the served model and the raw upstream stop reason, and recognise a translation artifact before rewriting your own tool loop.
Own the loss. Decide which vendor-specific capabilities the product genuinely needs, accept that a common shape cannot express them, and set the policy for when a workload leaves the gateway for a native integration.
## The gateway's real job Accepting an OpenAI-shaped body is easy. The work is in translating it into APIs that are not OpenAI-shaped at all, and translating the answers back so that a single client parser works everywhere. Anthropic's Messages API is the clearest example of a structurally different target, so it makes the best worked case, but the same class of transformation happens for every non-OpenAI vendor. ## Request-side transformations **System prompt placement.** OpenAI models a system prompt as a message in the array with `role: "system"`. Anthropic models it as a separate top-level parameter, outside the conversation turns. The gateway therefore extracts your system messages from `messages` and sets them where the vendor expects them. If you send several, they have to be consolidated, because the target has one slot. **Required fields.** Anthropic requires a `max_tokens` on every request; OpenAI treats it as optional. A request that omits it cannot simply be forwarded, so the gateway supplies a value. That is worth knowing operationally: if your output looks truncated on one vendor and not another with identical code, an implicit cap is the first thing to check, and setting `max_tokens` explicitly removes the ambiguity. **Content shape.** OpenAI's `content` is a string or a list of typed parts; other vendors use their own block structures for text and images. Multi-part user content gets rewritten into whatever block vocabulary the target speaks, and a model with no image support cannot receive image parts at all. **Tool calling.** This is the transformation with the most moving pieces. In the OpenAI shape you define `tools` as JSON-Schema functions, the assistant answers with `tool_calls`, each having an id, and you reply with a `role: "tool"` message carrying that `tool_call_id`. Vendors with block-structured messages instead put a tool-use block inside the assistant turn and expect the result as a tool-result block inside the next user turn, correlated by the vendor's own id. The gateway maps ids and roles in both directions so your loop stays in the OpenAI idiom. **Alternation and merging.** Some vendors are strict about turns alternating user/assistant and about the first turn being a user turn. Two consecutive user messages, or a conversation that opens with an assistant message, may have to be merged or adjusted to satisfy the target — a detail that occasionally makes a transcript reach the model slightly differently than you wrote it. ## Response-side transformations The reply is folded back into the familiar envelope: `choices[0].message` with `role` and `content`, `tool_calls` when tools were invoked, a `finish_reason`, and a `usage` object with prompt, completion and total token counts. The finish-reason mapping is the part worth memorising, because your control flow keys on it. A vendor's "model finished its turn" becomes `stop`; "hit the token cap" becomes `length`; "wants to call a tool" becomes `tool_calls`; a safety stop maps to the content-filter value. Alongside the normalised value, the gateway also exposes the upstream provider's raw stop reason, which is what you want in logs when the normalised value is too coarse to explain a behaviour. Streaming is normalised too: whatever event grammar the vendor uses upstream — and they differ substantially — comes out as OpenAI-style SSE chunks with `choices[].delta` fragments terminated by `data: [DONE]`. ## Where the translation is lossy A normaliser can only express what the common shape can express. Vendor features with no OpenAI counterpart either get an OpenRouter-specific extension field — the gateway defines a few unified extensions, such as a `reasoning` object for models that expose reasoning effort — or they are simply not reachable through the standard body. Fine-grained vendor-only controls, newly launched capabilities, and idiosyncratic response metadata are the usual casualties, and the lag between a vendor shipping a feature and a gateway exposing it is real. ## What to do with this knowledge Write code against the normalised shape, but log the raw stop reason and the served model so you can explain divergence. Set `max_tokens` explicitly instead of relying on a default you did not choose. Keep one system message rather than several, so consolidation is not making a decision for you. And when a tool loop misbehaves on one vendor only, suspect the id-and-role mapping, not your own loop, before rewriting application logic.
- You send no max_tokens and get a shorter answer than expected from one vendor. Why?Some vendors require a maximum output length on every request, so when your body omits it the gateway must supply one before forwarding. That implicit cap can truncate output on that vendor while an OpenAI model with no cap keeps going. Set `max_tokens` explicitly to a value your application actually wants, and check the finish reason: `length` confirms truncation rather than an early stop.
- Why should you log something beyond the normalised finish_reason?Because normalisation is lossy in exactly the place you debug. Several distinct upstream stop conditions collapse into `stop`, so the normalised value cannot always explain why generation ended. The gateway also surfaces the provider's raw stop reason and the model that actually served the request; logging both lets you correlate a behaviour change with a vendor-side condition instead of guessing from a coarse label.
- Does the tool-calling loop you write change when you switch to a vendor with block-structured messages?No — that is the point of the gateway. You keep defining JSON-Schema tools, keep reading `tool_calls` off the assistant message, and keep replying with a `role: "tool"` message carrying the matching `tool_call_id`. The conversion into that vendor's tool-use and tool-result blocks, including id correlation, happens inside the translation layer. What can still differ is how reliably each model decides to call tools at all.
saying these in an interview costs you the question
- Assuming the vendor receives your JSON body verbatim
- Expecting Anthropic to accept a system role in messages
- Treating finish_reason values as vendor-specific strings
- Believing normalisation exposes every vendor feature
- Ignoring that omitted fields may be filled in for you