skip to content

Porting an OpenAI service to DeepSeek: which calls break or silently change?

level: seniorimportance: must knowfreq 48%

answer

  1. Audit capabilities, not just URLs
  2. Chat ports; other endpoints do not
  3. Identical code, weaker guarantee
  4. Accepted-and-ignored beats rejected for danger
  5. One unfamiliar 4xx in the taxonomy

basics

~20 s

Chat completions port cleanly. Everything outside them does not: DeepSeek exposes no embeddings, image, audio or assistants endpoints. Schema-strict JSON output is unavailable, some parameters are accepted but ignored, and billing errors surface as an unfamiliar HTTP status.

solid answer

~50 s

Do the migration as a capability audit, not a URL change. **Endpoint coverage:** DeepSeek's platform is essentially chat completions plus account endpoints, so any embeddings, image, audio or assistants/threads code has no target and must keep a second provider. **Guarantee downgrades:** JSON output exists only in the coarse object mode, so schema-strict structured output does not carry over, and tool-calling reliability should be re-evaluated on your own prompts rather than assumed. **Silently ignored parameters** are the worst class — certain sampling parameters are accepted for some models and have no effect, so nothing in the response tells you a knob went dead. **Errors:** alongside the familiar 401 and 429, DeepSeek returns 402 for insufficient account balance, which OpenAI-shaped clients surface as a generic status error and most retry wrappers will not classify correctly. **Auth:** an unset `api_key` falls back to `OPENAI_API_KEY`. Model ids change too — `deepseek-chat` and `deepseek-reasoner`.

go deeper

for a junior

Know that only the chat side ports over, that model ids change to deepseek-chat or deepseek-reasoner, and that you must pass the DeepSeek key explicitly rather than relying on the SDK default.

for a middle

Explain the three categories — ports unchanged, ports weaker, no target — and give a concrete example of each, especially the missing embeddings endpoint and the coarser JSON output mode.

for a senior

Demonstrate the operational sweep: status-code taxonomy including the payment error, retry classification, timeout re-tuning, behavioural smoke tests for parameters you depend on, and re-running evaluations rather than trusting prompts to transfer.

for a principal

Own the decision of whether to stay dual-provider deliberately, where the capability matrix lives, and how you prevent a compatible SDK from disguising a capability gap as an equivalence in future changes.

## Frame it as a capability audit The migration temptation is to change `base_url`, run the happy path, see text come back and declare it done. That works for the one call you tested and hides everything else. The disciplined version is an inventory: list every distinct call your service makes, and classify each as *ports unchanged*, *ports with a weaker guarantee*, or *has no target*. ## Calls with no target DeepSeek's platform is a chat-completions provider. Its surface is the chat endpoint, a legacy completions endpoint behind the beta base URL, a model listing and an account balance endpoint. There is no embeddings API, no image generation, no speech, and no stateful assistants/threads/runs machinery. The consequence for a RAG service is immediate and often forgotten: your retrieval side still needs an embedding model from somewhere else. Two providers in one pipeline is fine, but it must be a decision rather than a surprise discovered when the first `client.embeddings.create` call 404s in staging. The same goes for anything built on stateful server-side sessions — that architecture does not transfer, and the conversation state comes back to your own store. ## Guarantees that get weaker while the code stays identical This is the category that produces incidents. The field names are the same, so the code compiles and runs, but the promise behind it is smaller. Structured output is the clearest case: object-mode JSON is available, schema-constrained strict output is not. Code that relied on the strong guarantee now needs prompt-level shape specification plus your own validation and a bounded retry. Tool calling is the second case. The request and response shape are OpenAI-like — a `tools` array of JSON Schema function definitions, `tool_calls` on the assistant message, replies with `role: "tool"` and a matching `tool_call_id`. What you cannot port is *reliability*: how often the model picks the right tool, how it behaves with many tools in scope, whether it loops. Re-run your agent evaluation suite; do not port a conclusion. ## Parameters that are accepted and ignored HTTP gives you two honest outcomes — accept and honour, or reject with an error. "Accept and ignore" is neither, and it is present here: depending on the model you select, some sampling parameters are taken without complaint and have no effect on generation. Nothing in the response says so. You discover it when outputs that used to be deterministic-ish under a low temperature stop being so, and you spend a day looking at your prompt. The defence is to read the provider's parameter-support notes per model before relying on a knob, and to build one behavioural smoke test per parameter you actually depend on — for example, does the same request at two very different temperature settings produce meaningfully different outputs? That is cheap and it catches the silent case that no status code will. ## Errors and retries Your retry policy was tuned against another provider's status taxonomy. Recheck it. 401 for a bad key and 429 for rate limiting behave as expected. The one that catches people is **402, insufficient balance**: DeepSeek is prepaid, and when credit runs out the API returns a payment status rather than a rate-limit or auth error. A generic OpenAI-shaped client raises an untyped status error for it, so a naive wrapper that retries on "any 4xx that isn't 400" will hammer a permanently-failing endpoint, and a wrapper that treats non-429 4xx as fatal will page someone at 3am for what is really a billing task. Classify 402 explicitly: do not retry it, alert on it, and monitor the balance ahead of time. Also re-tune timeouts. Response latency profiles differ between providers and between models, and a timeout inherited from elsewhere may cut off long generations that would have completed. ## Auth and model ids Two small ones that cause a disproportionate number of failed first attempts. The OpenAI SDK's default credential is the `OPENAI_API_KEY` environment variable, so a client built without an explicit `api_key` will send the wrong key to DeepSeek and 401. And OpenAI model ids do not exist here — the identifiers are `deepseek-chat` and `deepseek-reasoner`. Model id belongs in configuration alongside the base URL and the key, so the three always move together. ## A workable migration order Inventory the calls; move the pure chat paths first and confirm streaming, tool calls and error handling end to end; keep the other provider for embeddings and any non-text modality; replace strict structured output with object mode plus validation; re-run evaluations rather than trusting prompt quality to transfer; and add explicit handling and alerting for the payment status before you send production traffic. Then decide whether the abstraction you now have is worth keeping permanently dual-provider.

  • Your RAG pipeline moves to DeepSeek. What breaks first?
    The retrieval half. DeepSeek has no embeddings endpoint, so document and query embedding must stay with another provider or a self-hosted model. The generation half ports cleanly. Plan for a two-provider pipeline with separate keys, separate rate-limit behaviour and separate failure handling, rather than discovering the gap when the first embeddings call 404s.
  • How should a retry wrapper treat a 402 response?
    As fatal and alertable, never retryable. It means the prepaid balance is exhausted, so every retry fails identically and only adds load and latency. Classify it separately from 429, page a human or trigger a top-up workflow, and ideally monitor the balance proactively so you find out before traffic starts failing.
  • Why is an accepted-but-ignored parameter more dangerous than a rejected one?
    A rejection fails loudly at the first call, in development, with a message naming the field. An ignored parameter returns 200 and plausible text, so the request looks healthy while the behaviour you configured never takes effect. You only notice through degraded output quality much later, and the investigation usually starts by blaming the prompt.
  • Do tool definitions written for OpenAI work unchanged?
    The wire shape ports — a `tools` array of JSON Schema functions, `tool_calls` on the assistant message, replies with `role: "tool"` and the matching `tool_call_id`. What does not port is behaviour: selection accuracy, handling of many tools, and loop tendencies are model properties. Re-run your agent evaluations on the new model before trusting the migration.

saying these in an interview costs you the question

  • Assuming every OpenAI endpoint has a DeepSeek equivalent
  • Treating a compiling client as a verified migration
  • Reusing the retry policy without checking status codes
  • Retrying a payment-required error indefinitely
  • Porting prompt-quality conclusions without re-running evaluations

context