skip to content

Migrating from a hosted OpenAI endpoint to vLLM, what still breaks despite API compatibility?

level: principalimportance: should knowfreq 38%

answer

  1. Protocol matches, platform does not
  2. Different checkpoint, different tokenizer
  3. Rejections become queueing
  4. Vendor-only surfaces have no equivalent
  5. Keep cross-cutting concerns in a gateway

basics

~20 s

The protocol matches; the platform does not. Token counts shift with a different tokenizer, length limits become hard rejections, provider rate limits and retry semantics disappear, vendor-only parameters and surfaces have no equivalent, and the model itself is a different model.

solid answer

~60 s

Route and schema compatibility means the SDK compiles and the first request succeeds — that is the easy 80%. What bites afterwards falls in three groups. **Model behaviour**: a different checkpoint means a different tokenizer, so token counts, cost accounting and prompt-budget assumptions all move, and output quality has to be re-validated on your own task rather than assumed. **Contract edges**: exceeding the server's configured context length is a rejected request rather than a silent accommodation; parameters the provider supported may be unsupported here; tool-call fidelity now depends on the checkpoint and its parser; `logprobs` depth is capped by server configuration. **Platform**: there is no billing, no organization quota, no managed failover, and no provider-issued rate-limit response — an overloaded vLLM queues, so a client tuned to retry on rate limits will instead pile on timeouts and make it worse. The principal move is to put auth, quotas, routing and provider fallback in a gateway so the swap really is a config change, then validate with replayed production traffic before cutting over.

go deeper

for a junior

Know that matching routes do not mean matching behaviour: it is a different model with a different tokenizer, and features the hosted platform provided around the API do not come with the compatibility layer.

for a middle

Name concrete divergences you would check — token counts and prompt budgets, hard rejection when the context limit is exceeded, optional parameters that may be unsupported, and tool-calling fidelity that now depends on the checkpoint.

for a senior

Focus on the operational shift: no provider rate limits means overload becomes queueing, retry policies must be re-tuned, capacity and availability are yours, and cutover needs replayed traffic and shadowing rather than a smoke test.

for a principal

Own the architecture that makes the move reversible: a gateway holding auth, quotas, chargeback, routing and hosted fallback, so migrating a workload either direction is a routing rule and no team rebuilds platform concerns inside application code.

## Why the easy part is genuinely easy The compatibility layer earns its keep. Routes, request schemas, response schemas and the streaming chunk format match, so the SDK, the streaming loop, the message-building code and the response parsing all survive the move untouched. Point `base_url` at the server, match the model name, and traffic flows. Any migration plan that budgets significant engineering for *that* is mis-scoped. The work is in everything the API shape does not describe. ## Group one: it is a different model The most under-estimated fact is that you are not migrating an endpoint, you are migrating a model. The open-weight checkpoint has its own tokenizer, so identical text yields a different token count — often by ten percent or more in either direction. Everything downstream of a token count shifts with it: your cost model, your "how many documents fit" retrieval budget, your truncation thresholds, your per-request usage records. Anywhere your code computes a token count against the provider's tokenizer, that computation is now wrong. Quality is a separate question and cannot be inferred from benchmarks. Instruction-following style, verbosity, refusal behaviour, JSON reliability and multilingual coverage all move. Prompts tuned over months against one model are, in effect, overfitted to it; expect a prompt-tuning pass, and expect some prompts to need real rework rather than tweaks. ## Group two: the contract edges **Length limits become hard errors.** The server has a configured maximum context length. If a request's prompt plus its requested output exceeds it, the request is rejected rather than quietly accommodated. Code paths that grew a conversation until the provider complained now fail earlier and differently, so history-trimming logic must be driven by your own accounting rather than by an error you catch. **Parameter support is a per-server fact.** Some fields the hosted API accepts have no equivalent here, and some vLLM extras have no equivalent there. The depth of `logprobs` you may request is bounded by server configuration. Structured-output and tool-calling behaviour depend on what the checkpoint and its configured parser support, not on the protocol. Audit which optional fields your code actually sends; the long tail is where a request starts failing in production for one feature flag. **Error surfaces differ.** Status codes and error bodies are similar in shape but not identical in trigger, so any code branching on a specific error condition needs re-verification. ## Group three: the platform is gone This is the part that catches senior teams. The provider was not only serving a model; it was providing capacity management, quota enforcement, billing, availability and a rate-limit protocol. All of that is now yours. The sharpest instance is **overload behaviour**. A hosted provider rejects excess traffic, and mature clients respond with backoff — that feedback loop is what keeps the system stable. vLLM does not reject; it queues. Excess load turns into a growing time-to-first-token, client timeouts fire, the client retries, and the retries add work to a server that is already behind. A retry policy tuned for a provider's rejections becomes an amplifier. You need client-side concurrency limits, timeouts set against your actual saturation curve, and admission control somewhere in the path. Availability is the other one: a single process on a single GPU has none of the redundancy you were implicitly buying, and losing the node loses the endpoint. ## The principal-level move Design the swap so it is genuinely a configuration change, which means the concerns that used to be the provider's should not be scattered through application code. Put a gateway in front of both the self-hosted endpoint and any hosted one: it owns authentication and per-tenant keys, quotas and rate limits, request-size caps, usage records for chargeback, routing by model name, and fallback to a hosted provider when the self-hosted pool is unhealthy or saturated. With that layer in place, moving a workload between hosted and self-hosted is a routing rule, and you keep the option to move back — which matters, because the first migration rarely covers every workload. ## Validating before you cut over Do not cut over on a smoke test. Capture a representative corpus of real production requests and replay it against both endpoints. Compare on things you can measure automatically: the distribution of `finish_reason` values, the rate of responses that fail to parse against the schema your code expects, tool-call validity, output-length distribution, and refusal rate — alongside whatever task-quality metric you already trust. Then shadow live traffic before it takes user-facing load, and keep the hosted route configured for a fast rollback. The compatibility layer makes that rollback a config flip, which is exactly the property you should protect.

  • Why can a retry policy tuned against a hosted provider make a self-hosted server worse?
    Because the failure signal changes. A provider rejects excess traffic and the client backs off, which sheds load. vLLM queues instead, so overload appears as rising time-to-first-token; the client's timeout fires and it retries, adding work to a server already behind. You need client concurrency caps, timeouts set from the real saturation curve, and admission control rather than blind retries.
  • How would you prove the swap is safe before sending user traffic?
    Replay a captured corpus of real requests against both endpoints and diff on automatable signals: finish_reason distribution, schema-parse failure rate for structured responses, tool-call validity, output-length distribution and refusal rate, plus your task metric. Then shadow live traffic without serving it, and keep the hosted route configured so rollback is a routing change.
  • Which responsibilities should sit in a gateway rather than in vLLM or the application?
    Authentication and per-tenant credentials, rate and concurrency limits, request-size and max-output caps, usage records for chargeback, routing by model name, and fallback to a hosted provider. Keeping them there means switching a workload between hosted and self-hosted is a routing rule, and no application team re-implements quota logic per service.
  • Your cost dashboard is built on the provider's token counts. What has to change?
    Recompute against the served model's tokenizer, because a different vocabulary yields different counts for identical text. Source counts from the response's usage fields rather than from a client-side estimate, and rebuild the unit economics on GPU-hour cost divided by delivered tokens instead of a per-token price, since self-hosting charges for capacity whether or not it is used.

saying these in an interview costs you the question

  • Assumes API compatibility means identical model behaviour
  • Reuses provider token counts for the new tokenizer
  • Keeps retry-on-rate-limit logic unchanged
  • Expects the server to trim prompts that exceed the limit
  • Cuts over on a smoke test with no replay or shadow
  • Thinks per-tenant quotas come with the endpoint

context