Which DeepSeek features require the /beta base URL, and what do they do?
answer
- Two extensions, one separate base URL
- Not OpenAI-shaped, hence beta
- Generate between prompt and suffix
- Force how the answer starts
- Pair with stop sequences
basics
~10 sFill-in-the-middle completion and assistant-prefix continuation are opt-in beta features on DeepSeek. Point the client at https://api.deepseek.com/beta to use them; on the standard base URL the requests are rejected or the extra fields ignored.
solid answer
~50 sDeepSeek keeps two non-OpenAI extensions behind a separate base URL, `https://api.deepseek.com/beta`, as of mid-2026. **Fill-in-the-middle (FIM)** uses the legacy completions endpoint with both `prompt` and `suffix`, so the model generates the text that belongs *between* them — the shape code-completion tooling needs, where the cursor has code on both sides. **Prefix continuation** works on the chat endpoint: you append a final assistant message carrying `"prefix": true`, and the model continues from that partial text instead of starting a fresh reply. It is the reliable way to force an output to begin a certain way — a `{`, a ```python fence, a required opening clause — usually paired with `stop` to end it cleanly. Both are genuine departures from the OpenAI surface: a stock OpenAI endpoint has no `prefix` flag on messages, and a trailing assistant message means something different there. Gate them behind a capability flag, because they will not survive a switch back.
code
python · 15 linesimport os
from openai import OpenAI
beta = OpenAI(
api_key=os.environ["DEEPSEEK_API_KEY"],
base_url="https://api.deepseek.com/beta",
)
fim = beta.completions.create(
model="deepseek-chat",
prompt="def fib(n):\n if n < 2:\n return n\n",
suffix="\n\nprint(fib(10))\n",
max_tokens=64,
)
print(fim.choices[0].text)go deeper
Recall that DeepSeek has extras beyond the OpenAI-shaped API and that they need a different base URL ending in /beta. Naming fill-in-the-middle and prefix continuation is enough at this level.
Explain what each one constrains — generating between a prompt and a suffix versus continuing from a partial assistant message — and why a plain chat call cannot express either.
Show that you would gate them behind a capability check, pair prefix with a stop sequence, handle the continuation-only response, and keep a documented degraded path for providers without them.
Own the architectural line: the beta base URL exists so the compatible surface stays honest, and adopting a beta extension is a portability commitment that needs an explicit owner and a fallback plan.
## Why there is a second base URL at all DeepSeek's main base URL is deliberately kept OpenAI-shaped so that a swapped client behaves predictably. Features that have no OpenAI equivalent would break that promise, so they are served from `https://api.deepseek.com/beta` instead. The separation is doing real design work: it means the compatible surface stays clean, and it means your code has to opt in explicitly rather than accidentally depend on a non-portable extension. The practical consequence is that you may need two clients in the same process — a default one on `https://api.deepseek.com`, and a beta one — differing only in base URL. ## Fill-in-the-middle FIM is not a chat feature. It uses the older text completions endpoint, and the request carries two pieces of context: `prompt`, the text before the insertion point, and `suffix`, the text after it. The model generates only what belongs between them. This matters because chat completion is inherently a continuation format — it can extend the end of something, but it has no way to express "there is already code below the cursor and your output must join up with it". That makes FIM the right primitive for editor integrations, patch generation and any templated document where the surrounding text is fixed. It has practical limits: the generated span is short by design, and you should set a `stop` sequence or a modest token cap so the model does not run past the join. Reading the result is different too — a completions response exposes `choices[0].text`, not `choices[0].message.content`. ## Prefix continuation Prefix continuation solves a different problem: constraining how a chat answer starts. You build the usual `messages` array, then append one more message with `role: "assistant"`, the partial text you want the answer to begin with, and `"prefix": true`. The model treats that text as already-emitted output and continues from it. The uses are concrete. Forcing a leading `{` makes JSON extraction more reliable. Forcing a ```python fence, with `stop` set to the closing fence, gets you a clean code block with no surrounding commentary. Forcing an opening clause pins the model into a stance so it cannot open with a hedge or a preamble. One subtlety: the returned `content` contains the continuation, so you generally concatenate your prefix with the response yourself to reconstruct the full text. Decide that once, in a helper, rather than at every call site. ## What makes these non-portable The `prefix` field is a DeepSeek extension to a message object. Sending it to a plain OpenAI endpoint gets you either an unknown-field error or silent ignoring — and silent ignoring is the worse outcome, because your prompt was engineered on the assumption that the constraint holds. A trailing assistant message without the flag is simply prior conversation turn, so the model starts a new reply rather than continuing. FIM is non-portable for a different reason: it depends on `suffix` being honoured on the completions endpoint, which is not a given elsewhere. Even where a `suffix` parameter exists on some other provider, its handling and its token accounting are that provider's business, and you should not assume equivalence. ## Engineering around them Treat both as capabilities, not as defaults. A reasonable shape is a small provider-capability record — `supports_fim`, `supports_prefix` — consulted before the call, with a documented degraded path when the capability is absent. For prefix, the degraded path is a strong system instruction plus post-hoc stripping of any preamble; for FIM, it is a chat prompt that shows both sides of the insertion point and asks for only the middle span, accepting lower reliability. Also keep in mind that these are *beta* surfaces. Beta means the shape can change and the reliability bar is lower than the stable endpoint. Anything on a critical path deserves a fallback and an alert if the beta call starts failing, rather than a silent degradation in output quality that nobody notices for a week. ## Interview framing The strong answer connects the mechanism to the motive: DeepSeek serves an OpenAI-compatible API, so anything that is *not* OpenAI-shaped has to live somewhere the compatibility promise does not apply. Naming the beta base URL, describing what each feature actually constrains, and then saying how you would keep the dependency contained is the full arc.
- Why can't a normal chat call do what fill-in-the-middle does?Chat completion only extends the end of a conversation; it has no field for text that must follow the generated span. FIM takes `prompt` and `suffix` together, so the model conditions on both sides of the insertion point and produces something that joins up with the code below the cursor. You can approximate it by showing both halves in a chat prompt, but nothing constrains the model to stop at the join.
- What happens if you send the prefix flag to a provider that does not support it?Either the request is rejected for an unknown field, or the flag is ignored and the trailing assistant message is read as an ordinary prior turn — so the model starts a fresh reply. The second case is the dangerous one, because the call succeeds while the constraint you designed the prompt around has quietly vanished. Gate it behind an explicit capability check.
- How do you reconstruct the full text when using prefix continuation?The response content is the continuation only, so concatenate your prefix with it. Put that in one helper rather than at every call site, and remember the `stop` sequence is consumed rather than returned — if you forced an opening code fence and stopped on the closing one, you have to re-add both delimiters yourself when rendering.
saying these in an interview costs you the question
- Calling FIM or prefix on the standard base URL
- Thinking prefix works on the chat endpoint everywhere
- Confusing FIM with a longer chat prompt
- Forgetting the response holds only the continuation
- Treating beta endpoints as production-stable