skip to content

DeepSeek

DeepSeek matters for two reasons: an OpenAI-compatible API at a fraction of the usual price, and open-weight reasoning models you can host yourself. Expect questions about handling reasoning tokens and about the cost arithmetic.

on this pageshow

explore

questions

19

How do you call DeepSeek's API using the OpenAI SDK, and what must you change?

level: juniorimportance: must knowfreq 78%

answer

  1. Same SDK, different endpoint
  2. Three constructor-level changes
  3. base_url, key, model id
  4. OPENAI_API_KEY fallback bites you
  5. /v1 is a path, not a version

basics

~10 s

Point the OpenAI client at base_url https://api.deepseek.com, pass your DeepSeek key as api_key, and use a DeepSeek model id such as deepseek-chat. The request shape, streaming and response objects stay unchanged.

solid answer

~40 s

DeepSeek serves an OpenAI-shaped `/chat/completions` endpoint, so you keep the official OpenAI SDK and change three things: the base URL, the key, and the model id. In Python that is `OpenAI(api_key=os.environ["DEEPSEEK_API_KEY"], base_url="https://api.deepseek.com")`, then `model="deepseek-chat"`. Both `https://api.deepseek.com` and `https://api.deepseek.com/v1` work — the `/v1` is only a path segment for clients that expect one, and says nothing about model version. Auth is a normal `Authorization: Bearer` header the SDK builds for you. The gotcha is the key: if you omit `api_key`, the SDK falls back to the `OPENAI_API_KEY` environment variable and you will send an OpenAI key to DeepSeek and get a 401. Everything else — `messages` with system/user/assistant roles, `stream=True` SSE deltas, `choices[0].message.content`, `finish_reason`, `usage` — behaves as you already know it.

code

python · 17 lines
python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEEPSEEK_API_KEY"],
    base_url="https://api.deepseek.com",
)

resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Summarise TCP in one sentence."},
    ],
    stream=False,
)
print(resp.choices[0].message.content)

go deeper

for a junior

Be able to name the three changes out loud: base_url, api_key, model id. Show that you would pass the DeepSeek key explicitly rather than relying on the environment default.

for a middle

Explain what compatibility actually covers — the wire protocol of the chat endpoint — and why streaming, retries and response parsing therefore keep working while other SDK namespaces do not.

for a senior

Show how you would structure the client factory so provider choice, credential and model id are configured together, and how you would smoke-test auth, streaming and a bad key after the swap.

for a principal

Own the question of whether a shared client is an asset or a trap: it preserves transport and observability code, but it hides capability differences behind a single type, so the capability matrix has to live somewhere explicit.

## What "OpenAI-compatible" actually means here An OpenAI-compatible API is one that speaks the same HTTP wire protocol as OpenAI's Chat Completions endpoint: the same path (`/chat/completions`), the same JSON request body (a `model` string plus a `messages` array of `{role, content}` objects), the same bearer-token auth header, the same server-sent-events framing when you stream, and the same response envelope (`id`, `object`, `created`, `model`, `choices[]`, `usage`). Because the OpenAI SDKs are thin clients over that protocol, any server that reproduces it can be driven by them. DeepSeek deliberately does this, which is why it is the standard worked example of provider portability. ## The three things you change **1. `base_url`.** The OpenAI SDK builds every request by appending a path to the base URL, so redirecting the client is a one-line change. DeepSeek accepts `https://api.deepseek.com` and, for clients that insist on a version segment, `https://api.deepseek.com/v1`. The `v1` is purely a URL alias — it does not select a model generation. **2. `api_key`.** DeepSeek issues its own keys from its platform console. The SDK sends whatever you give it as `Authorization: Bearer <key>`. If you construct the client without an explicit `api_key`, the SDK reads the `OPENAI_API_KEY` environment variable, so a half-migrated service will happily authenticate against DeepSeek with an OpenAI key and fail with 401. Read `DEEPSEEK_API_KEY` explicitly. **3. `model`.** OpenAI model ids do not exist on DeepSeek. The two you use are `deepseek-chat` (the general conversational model) and `deepseek-reasoner` (the reasoning line). Passing a foreign model id is a 400/422, not a silent fallback. ## What you do not change Message roles, multi-turn history construction, `stream=True` and the incremental `choices[0].delta.content` chunks, `max_tokens`, `stop`, `tools` for function calling, and the `usage` accounting object are all in the familiar shape. Existing retry wrappers, request-id logging, timeouts and connection pooling in the SDK continue to work, because they live below the protocol layer. This is the real value of compatibility: your transport, observability and error-handling code survives the swap untouched. ## The response object A non-streaming call returns the standard chat completion: `resp.choices[0].message.content` for the text, `resp.choices[0].finish_reason` for why generation stopped, and `resp.usage` with prompt, completion and total token counts. DeepSeek adds some vendor-specific fields to `usage`; the SDK's typed model may not surface them as attributes, so read them from the raw dictionary if you need them. ## Where the compatibility stops Compatibility is at the level of the chat endpoint, not the whole OpenAI platform. DeepSeek's surface is essentially chat completions plus a couple of account endpoints; there is no embeddings, image, audio or assistants equivalent, so those SDK namespaces will 404. Structured-output strict mode with a JSON Schema is not offered — only the coarser JSON object mode. Some sampling parameters are accepted and ignored rather than rejected, which is the dangerous class of difference because nothing tells you. And DeepSeek's beta features (fill-in-the-middle and prefix continuation) live behind a different base URL entirely. ## How to structure the code Because only the constructor arguments differ, the cleanest pattern is one factory that returns a configured client per provider, with the model id carried alongside it rather than hard-coded at each call site. Do not scatter `base_url` strings through the codebase, and do not assume that because the client type is the same the capabilities are. Treat the shared SDK as transport, and keep an explicit note of which provider-specific behaviours your prompts depend on. ## Practical checks After the swap, verify three things end to end: a non-streaming call returns content; a streaming call yields deltas and terminates; and an intentionally bad key returns 401 through your error path rather than crashing. Those three cover almost every wiring mistake in a base-URL migration.

  • If you construct the OpenAI client without passing api_key, where does the key come from?
    From the `OPENAI_API_KEY` environment variable — the SDK's default. On a DeepSeek base URL that means you silently send an OpenAI key to DeepSeek and get a 401 authentication failure. Always pass `api_key` explicitly (for example from `DEEPSEEK_API_KEY`) so the credential and the endpoint are chosen in the same place.
  • Does the /v1 in https://api.deepseek.com/v1 refer to a model version?
    No. It is only a URL path alias so clients that assume a version segment keep working; `https://api.deepseek.com` is equally valid. Which model generation you get is decided entirely by the `model` field — `deepseek-chat` or `deepseek-reasoner` — not by the URL.
  • Does streaming work the same way as with OpenAI?
    Yes. Setting `stream=True` returns SSE chunks in the same shape, and you accumulate `choices[0].delta.content` until the stream ends. Iteration code written against OpenAI needs no changes, which is why streaming is usually the first thing that keeps working after a base-URL swap.

It is like keeping the same phone and swapping the SIM: the handset, contacts and dialling habits are unchanged, but the number you call from and the network's feature list are not.

saying these in an interview costs you the question

  • Thinking you need a separate DeepSeek-specific SDK
  • Assuming every OpenAI endpoint exists because chat does
  • Leaving api_key unset and relying on OPENAI_API_KEY
  • Reading /v1 as the model version
  • Expecting an unknown model id to fall back silently

context

open as a page

In DeepSeek's API, what do the deepseek-chat and deepseek-reasoner names select?

level: juniorimportance: must knowfreq 72%

basics

~10 s

deepseek-chat selects DeepSeek's general-purpose V3 chat line; deepseek-reasoner selects the R1-style reasoning line that thinks before answering. Both are moving aliases pointing at whatever generation is current, not frozen model versions.

open as a page

In DeepSeek's API, what is reasoning_content in a deepseek-reasoner response?

level: juniorimportance: must knowfreq 70%

basics

~20 s

deepseek-reasoner puts its chain of thought in message.reasoning_content and the user-facing answer in message.content. Both come back in one ordinary chat-completion response. Show users content; treat reasoning_content as diagnostic text you log, collapse, or discard.

open as a page

What does DeepSeek's response_format json_object guarantee, and what does it not?

level: middleimportance: must knowfreq 60%

basics

~20 s

DeepSeek's JSON output mode constrains the model to emit syntactically valid JSON, nothing more. It does not validate against a schema you supply, so field names, types and required keys remain your responsibility to prompt for and then verify.

open as a page

In DeepSeek's API response, what do prompt_cache_hit_tokens and prompt_cache_miss_tokens mean?

level: middleimportance: must knowfreq 62%

basics

~20 s

DeepSeek splits a request's input tokens into two billed buckets: prompt_cache_hit_tokens, served from its context cache at a steeply discounted input rate, and prompt_cache_miss_tokens, billed at the full input rate. The two always sum to prompt_tokens.

open as a page

In DeepSeek's API, why must you strip reasoning_content before the next turn?

level: middleimportance: must knowfreq 62%

basics

~20 s

DeepSeek rejects a request whose messages array contains reasoning_content, returning a 400. Append only the assistant's content to the conversation. The model re-derives fresh reasoning each turn from the visible messages, so nothing is lost by dropping the previous trace.

open as a page

Porting an OpenAI service to DeepSeek: which calls break or silently change?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Chat completions port cleanly. Everything outside them does not: DeepSeek exposes no embeddings, image, audio or assistants endpoints. Schema-strict JSON output is unavailable, some parameters are accepted but ignored, and billing errors surface as an unfamiliar HTTP status.

open as a page

What must you change in a DeepSeek API request to enable context caching?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Nothing. DeepSeek's context caching is on by default for every account on both deepseek-chat and deepseek-reasoner. There is no cache parameter, marker or header to send, and DeepSeek charges nothing to store the cached content.

open as a page

What does MIT licensing of DeepSeek-V3 and R1 open weights permit?

level: middleimportance: should knowfreq 42%

basics

~20 s

It permits commercial use, modification, redistribution and derivative models with no user or revenue threshold and no naming requirement — you only keep the copyright and licence notice. It covers the weights, not your use of the hosted API.

open as a page

Is DeepSeek-R1-Distill-Qwen-14B a smaller DeepSeek-R1, and what is it actually?

level: middleimportance: should knowfreq 47%

basics

~20 s

No. The R1 distills are third-party base models — Qwen or Llama checkpoints of 1.5B to 70B — fine-tuned on reasoning traces produced by DeepSeek-R1. They are dense models with a different lineage, not shrunken R1 weights.

open as a page

Which sampling parameters have no effect on DeepSeek's deepseek-reasoner model?

level: middleimportance: should knowfreq 48%

basics

~20 s

DeepSeek documents temperature, top_p, presence_penalty and frequency_penalty as unsupported on deepseek-reasoner: the request succeeds but the values do nothing. logprobs and top_logprobs are rejected with an error instead of being ignored. Verify against the docs for your model release.

open as a page

Which DeepSeek features require the /beta base URL, and what do they do?

level: seniorimportance: should knowfreq 34%

basics

~10 s

Fill-in-the-middle completion and assistant-prefix continuation are opt-in beta features on DeepSeek. Point the client at https://api.deepseek.com/beta to use them; on the standard base URL the requests are rejected or the extra fields ignored.

open as a page

When is DeepSeek's R1 reasoning line the wrong choice versus the V3 chat line?

level: seniorimportance: should knowfreq 53%

basics

~20 s

Whenever the task has no deliberation to do: short classification, extraction, reformatting, routing, and any path with a tight latency budget or many small agent turns. Deliberation is generated text, so it costs time and output tokens on every call.

open as a page

How does the DeepSeek API behave under load, given it publishes no rate limits?

level: seniorimportance: should knowfreq 44%

basics

~20 s

DeepSeek does not enforce a per-key request or token limit. Under heavy traffic it queues your request and holds the HTTP connection open, sending filler — blank lines when not streaming, SSE comment lines when streaming — so client timeouts, not 429s, are the usual failure mode.

open as a page

When streaming deepseek-reasoner, how do reasoning_content and content deltas arrive?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Chunks first carry delta.reasoning_content while the model thinks, then switch to delta.content for the answer. Accumulate the two into separate buffers and treat the first content delta as the phase transition; usage totals only arrive at the end of the stream.

open as a page

Your deepseek-reasoner answers cut off mid-sentence — how do reasoning tokens explain it?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The chain of thought is generated output. It consumes the same token budget as the answer and is counted in completion_tokens, so a long think can exhaust max_tokens before the answer finishes. Check finish_reason for length and size the budget for thinking plus answer.

open as a page

How do you keep a product stable when deepseek-chat moves to a new generation?

level: principalimportance: should knowfreq 33%

basics

~20 s

Treat the alias as a floating dependency: hold a regression evaluation suite, run it continuously against production-shaped prompts, and alert on drift. Where behaviour must actually be frozen, serve a pinned revision of the open weights yourself — the only real version lock available.

open as a page

For a DeepSeek workload, how do you model per-request cost and prove caching paid off?

level: principalimportance: should knowfreq 34%

basics

~20 s

Model each call as a three-term sum: cache-hit input tokens, cache-miss input tokens and completion tokens, each at its own rate. Caching discounts only the first term, so an output-heavy or reasoner workload can show a 90% hit rate and barely move the bill.

open as a page

Is swapping base_url to an OpenAI-compatible provider like DeepSeek a sound portability strategy?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Partly. Wire compatibility genuinely portes transport, streaming and error plumbing, which is most of the integration code. It does not port behaviour, capabilities or cost profile, so a base_url swap is a good starting point and a poor abstraction boundary.

open as a page