Is swapping base_url to an OpenAI-compatible provider like DeepSeek a sound portability strategy?
answer
- Protocol claim, not behaviour claim
- Transport ports, judgement does not
- Capabilities as explicit data
- Evaluation is the real migration
- Tokenisation differs, so cost differs
basics
~20 sPartly. Wire compatibility genuinely portes transport, streaming and error plumbing, which is most of the integration code. It does not port behaviour, capabilities or cost profile, so a base_url swap is a good starting point and a poor abstraction boundary.
solid answer
~50 sCompatibility is a **protocol** claim, not a **behaviour** claim. Because the request and response shapes match, your SDK, streaming loop, retry wrapper and logging survive the swap — genuinely valuable, and the reason a compatible provider is cheap to trial. What does not survive is everything above the wire: prompt sensitivity, tool-selection accuracy, JSON adherence, latency distribution, capability coverage such as strict schema output or embeddings, and token accounting that feeds your cost model. So the design rule I use is: share the transport, but make capabilities explicit. Keep a per-provider capability record the call site consults, keep the model id and base URL configured together, and keep a per-provider evaluation suite you actually re-run before shifting traffic. The failure I plan against is a single typed client that makes two unequal providers look identical, so a capability gap ships as an assumed equivalence.
go deeper
Know that a compatible API lets you reuse the same SDK, but that the model behind it is a different model, so answers and quality can change even though the code does not.
Be able to list what ports (transport, streaming, retries, response parsing) and what does not (prompt behaviour, capability coverage, latency, tokenisation) and give one concrete example of each.
Show the operational plan: capability record consulted at call sites, config that binds base URL, key and model together, re-tuned timeouts and error classification, and a shadow period before traffic moves.
Own the strategic call — whether multi-provider optionality has a concrete driver at your volume, what the carrying cost is in prompts and evaluations, and where the capability matrix lives so a gap cannot ship as an assumed equivalence.
## Separate the two claims When a provider says it is OpenAI-compatible, it is making a narrow, checkable claim: the HTTP path, request body, auth header, streaming framing and response envelope match. It is not claiming that the same prompt yields comparable output, that the same parameter has the same effect, or that the same feature exists. Almost every disappointment with portability comes from hearing the second claim in the first one. ## What compatibility genuinely buys The honest value is large. Client libraries, connection pooling, timeout and retry policy, streaming consumption, request/response logging, trace propagation and the shape of your test doubles are all reusable. That is the bulk of the integration surface, and it means the cost of *trialling* another provider drops from a sprint to an afternoon. For a team evaluating price or capacity alternatives, that is a real strategic asset, and it is why compatible APIs proliferated. It also lowers exit cost in a meaningful sense: you are not locked into a vendor SDK whose abstractions leak into your domain code. ## What it does not buy **Behaviour.** Prompts are tuned against a model, not a protocol. Instruction-following, verbosity, refusal behaviour, formatting habits and reasoning depth all change. A prompt suite that scores well on one model can degrade noticeably on another without a single line of your code changing. **Capability coverage.** Endpoint availability differs — embeddings, images, audio, stateful sessions. Guarantee strength differs — strict schema output versus coarse JSON mode. Parameter support differs, sometimes silently, where a value is accepted and has no effect. **Operational shape.** Latency distributions, throughput ceilings, throttling behaviour under load, error taxonomies including provider-specific statuses, and maintenance windows are all provider properties. Your timeouts, circuit breakers and alerts encode assumptions about them. **Cost model.** Even when both bill per token, tokenisation differs, so the same text is a different number of tokens. Caching and discount mechanics differ. A cost projection built on one provider's counters is an estimate, not a translation. ## The abstraction question The tempting move is a thick internal `LLMClient` that hides which provider is behind it. In practice that abstraction fails where it matters: it either exposes the lowest common denominator, discarding capabilities you paid for, or it exposes a union of features that silently no-op on providers lacking them. Both are worse than being explicit. The shape that holds up is thinner than people expect: share transport and observability, but treat capability as data. A small per-provider record — supports schema-strict output, supports fill-in-the-middle, supports prefix continuation, supports embeddings, ignores particular sampling parameters — consulted at the call site, with an explicit degraded path when a capability is absent. The point is not the record itself; it is that the gap becomes visible in code review instead of in production output quality. Pair that with configuration where base URL, credential and model id travel as one unit. Splitting them is how you end up sending one provider's model id to another's endpoint. ## Evaluation is the real portability tax If you take one thing from this: the code migration is hours, the evaluation migration is weeks. Before shifting production traffic you need a task-representative evaluation set, run against the candidate, scored on the dimensions you actually care about — answer quality, tool-selection accuracy, schema adherence rate, latency percentiles and cost per successful task rather than cost per million tokens. Without that, "it works" means "the happy path returned 200". A shadow or canary period is worth more than any amount of reasoning about protocol equivalence: mirror a slice of real traffic, compare outputs, and only then move the dial. ## When multi-provider is actually worth it Be honest that portability has carrying cost — two sets of keys, two error taxonomies, two evaluation runs, two sets of prompts drifting apart. It pays when you have a concrete driver: meaningful cost arbitrage at your volume, a real availability requirement that a single vendor cannot meet, a regulatory or regional constraint, or genuine leverage in negotiation. It does not pay as insurance against a vendor risk nobody has quantified. The mature position is therefore: exploit compatibility to keep switching *cheap*, but do not pretend the switch is *free*, and do not let a shared client type be the place where that pretence lives.
- Where would you draw the abstraction boundary between your code and the provider?Below the prompt, above the HTTP client. Share transport, retries, streaming and observability; keep prompts, capability assumptions and model choice provider-explicit. A thick client that hides which provider is in use either flattens to the lowest common denominator or silently no-ops features, and both mistakes surface as quality regressions rather than errors.
- What would you measure before shifting production traffic to a compatible alternative?Task-level outcomes on a representative evaluation set: answer quality, tool-selection accuracy, schema-adherence rate, latency percentiles, and cost per successful task rather than per million tokens. Then run a shadow or canary slice of real traffic and compare outputs before moving the dial. Protocol equivalence tells you nothing about any of these.
- Both providers bill per token. Why is cost still not directly comparable?Tokenisation differs, so identical text is a different token count on each side. Caching, discount and batching mechanics differ too, and retries caused by weaker output adherence add hidden volume. Compare cost per successfully completed task measured on your own traffic, not headline rates per million tokens.
saying these in an interview costs you the question
- Treating protocol compatibility as behavioural equivalence
- Building a thick client that hides which provider is used
- Migrating on a happy-path smoke test alone
- Comparing headline per-token prices across tokenisers
- Keeping multi-provider optionality with no concrete driver