skip to content

When do you need DashScope's native API instead of its OpenAI-compatible endpoint?

level: seniorimportance: should knowfreq 34%

answer

  1. Two envelopes, overlapping capability
  2. One surface streams cumulatively by default
  3. Portability is what you trade away
  4. Submit-and-poll has no chat-shaped form
  5. input and parameters, not flat fields

basics

~20 s

Reach for DashScope's native /api/v1 surface when the OpenAI shape has no room for what you need — asynchronous task submission for long-running generation, endpoints that are not chat-shaped, and native-only request knobs. Everything conversational is simpler in compatible mode.

solid answer

~50 s

Model Studio exposes the same models two ways. The OpenAI-compatible surface at `/compatible-mode/v1` buys you ecosystem compatibility: existing OpenAI clients, frameworks and gateways work unchanged. The native surface at `/api/v1/services/...` is Alibaba's own shape — `input` holding the messages, `parameters` holding the generation settings — and it is where features that do not fit a Chat Completions body live, including asynchronous task submission for long-running generation and service endpoints that are not chat-shaped at all. There is also a streaming difference worth knowing: the native surface's `incremental_output` parameter controls whether each event carries only the new text or the full text so far, whereas compatible mode always emits deltas. My default is compatible mode for anything conversational, dropping to native only for the specific capability that forces it — and isolating that call behind an interface so the rest of the codebase stays portable.

code

python · 14 lines
python
import dashscope
from dashscope import Generation

dashscope.base_http_api_url = "https://dashscope-intl.aliyuncs.com/api/v1"

responses = Generation.call(
    model="qwen-plus",
    messages=[{"role": "user", "content": "Explain incremental_output."}],
    result_format="message",
    stream=True,
    incremental_output=True,
)
for r in responses:
    print(r.output.choices[0].message.content, end="")

go deeper

for a junior

Know that Model Studio has two surfaces — the OpenAI-compatible one under /compatible-mode/v1 and Alibaba's native one under /api/v1 — and that the compatible one is the normal choice for chat.

for a middle

Describe the native body shape with messages under input and settings under parameters, and explain incremental_output: native streaming events are cumulative by default, unlike compatible mode's deltas.

for a senior

Name what actually forces native — asynchronous submit-and-poll generation, non-chat-shaped endpoints, native-only envelope parameters — and state the portability and tooling cost you accept when you go there.

for a principal

Own the boundary decision: default the organisation to the compatible shape so gateways, tracing and provider swaps stay cheap, and require any native usage to sit behind an internal interface with the same observability as everything else.

## Two doors to the same models Alibaba Cloud Model Studio serves Qwen through two request surfaces on the same hosts: - **native DashScope**, under `/api/v1/services/...` — for text generation, the path is `/api/v1/services/aigc/text-generation/generation` - **OpenAI-compatible**, under `/compatible-mode/v1` — Chat Completions request and response shape They are not two products. They are two envelopes around overlapping capability, and choosing between them is a portability-versus-capability decision you should be able to argue explicitly. ## The native request shape The native body is structured differently from Chat Completions. Rather than a flat object with `model`, `messages` and generation parameters as siblings, it nests: `model` at the top, the conversation inside **`input`**, and the generation settings inside **`parameters`**. Vendor flags that ride in `extra_body` in compatible mode are ordinary members of `parameters` here. Alibaba's own `dashscope` Python SDK targets this surface — `Generation.call(model=..., messages=..., result_format="message", ...)` — and reads the key from `DASHSCOPE_API_KEY` automatically, which is one small ergonomic win over the OpenAI client. That SDK also carries the region setting as module state: assigning `dashscope.base_http_api_url` to the Singapore native base URL switches deployments, because the default targets the Beijing host. ## The streaming difference people get wrong In compatible mode, streaming behaves as OpenAI users expect: every chunk is a **delta**, and you concatenate them. On the native surface, streaming events by default carry the **full text generated so far**, not just the increment — you set **`incremental_output=True`** to get delta semantics. Get this wrong and the symptom is unmistakable: output that repeats itself with exponentially growing duplication, because you concatenated cumulative snapshots. It is a five-minute bug once you have seen it and a bewildering one the first time. A related native knob is **`result_format`**, which selects whether the response is returned in a plain text-completion shape or a message shape with roles; message shape is what you want for chat and is the modern default. ## What actually forces you to native Be honest about this list, because "native is more powerful" is a vague claim that does not survive a follow-up question. The real drivers: 1. **Asynchronous task submission.** Some generation work — long-running media generation in particular — is submitted as a task and polled for completion rather than answered inline. A Chat Completions body has no place to express that lifecycle, so it lives on the native surface, driven by DashScope-specific request headers and a task-status endpoint. 2. **Endpoints that are not chat-shaped.** Model Studio hosts services beyond conversational text generation. Anything whose request or response does not fit a `messages` array plus a completion has no compatible-mode representation. 3. **Native-only parameters.** `incremental_output` and `result_format` are the clearest examples: they describe the native envelope itself and have no analogue in the OpenAI shape. Everything else — ordinary chat, streaming chat, tool calling, JSON output, the thinking flag — is reachable from compatible mode, with vendor parameters passed through `extra_body`. ## What you give up by going native Portability, and more of it than teams expect. The compatible surface is the reason a Qwen swap is a base-URL change: LLM frameworks, observability proxies, gateway products and internal abstractions overwhelmingly speak Chat Completions. Write directly against `input`/`parameters` and every one of those integrations needs a bespoke adapter, and a future provider migration becomes a rewrite of the call layer rather than a config change. You also give up a shared mental model. On a team where most services speak the OpenAI shape, the one service using the native envelope is the one nobody wants to touch. ## The recommendation to give in an interview Default to compatible mode. Use native for the specific capability that has no compatible-mode expression, and when you do, **put it behind your own interface** — a function that takes your domain types and returns your domain types, with the native body construction entirely inside it. Then the blast radius of the vendor-specific envelope is one file, and the rest of the codebase keeps its portability. A useful tie-breaker when both surfaces can do the job: pick the one your operational tooling already understands. If your token accounting, tracing and rate-limit middleware are built around Chat Completions responses, a native call is invisible to all of it until you write the adapter — and unmeasured traffic is a worse problem than a slightly less ergonomic request body.

  • A colleague's native DashScope stream renders text that repeats and grows. What is the bug?
    They are concatenating cumulative snapshots. The native surface's streaming events carry the full text so far unless `incremental_output=True` is set, so appending each event duplicates everything already printed. Either set the flag and keep appending, or leave it off and replace the buffer with each event rather than adding to it.
  • If you must use the native surface for one feature, how do you limit the damage to portability?
    Isolate it behind your own interface: one function that accepts domain types, builds the native `input`/`parameters` body inside itself, and returns domain types. Nothing outside that file learns the envelope. Also make sure the call still emits the same tracing and token-accounting signals as the rest of your traffic, or it becomes invisible to operations.
  • Which surface would you choose for a service that may later move to a different provider?
    Compatible mode, without much hesitation. Its whole value is that the migration becomes a base-URL and model-id change rather than a rewrite of the call layer, and it keeps you inside the ecosystem of gateways, frameworks and observability tooling that already speaks Chat Completions. Native is a deliberate exception for a capability that has no other expression.

saying these in an interview costs you the question

  • Claims native is simply the more powerful surface without naming a capability
  • Concatenates native stream events without setting incremental_output
  • Expects messages as a top-level field in the native request body
  • Ignores the portability cost of writing against the vendor envelope
  • Thinks the two surfaces serve different models

context