skip to content

Qwen

Alibaba's Qwen family is the broadest open-weight line-up — many sizes, MoE variants and specialist models — reachable either as a hosted API or as weights you run yourself. It is the usual answer when the requirement is multilingual and self-hosted.

on this pageshow

explore

questions

15

Which base_url and API key let the OpenAI SDK call Alibaba's DashScope Qwen models?

level: juniorimportance: must knowfreq 66%

answer

  1. Two changes, not a rewrite
  2. One field selects the region
  3. Bearer token, ordinary header
  4. The env var the OpenAI client ignores
  5. compatible-mode path plus DASHSCOPE_API_KEY

basics

~10 s

Point the OpenAI client at Model Studio's compatible-mode base URL — https://dashscope-intl.aliyuncs.com/compatible-mode/v1 for Singapore, https://dashscope.aliyuncs.com/compatible-mode/v1 for Beijing — and pass your DASHSCOPE_API_KEY as api_key. Model ids are Qwen names such as qwen-plus.

solid answer

~40 s

Alibaba Cloud Model Studio serves Qwen through DashScope, and DashScope offers an OpenAI-compatible surface at `/compatible-mode/v1`. So you keep the OpenAI SDK and change exactly two things: `base_url` becomes `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` (Singapore) or `https://dashscope.aliyuncs.com/compatible-mode/v1` (Beijing), and `api_key` becomes the Model Studio key, conventionally held in the `DASHSCOPE_API_KEY` environment variable. It travels as a normal `Authorization: Bearer` header. Watch the env-var trap: the OpenAI client falls back to `OPENAI_API_KEY`, never to `DASHSCOPE_API_KEY`, so you must pass the key explicitly or you will silently send the wrong credential and get a 401. After that, `chat.completions.create` works as usual with a Qwen model id; Qwen-only knobs ride in `extra_body`. This describes Model Studio as of mid-2026.

code

python · 13 lines
python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DASHSCOPE_API_KEY"],
    base_url="https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
)

resp = client.chat.completions.create(
    model="qwen-plus",
    messages=[{"role": "user", "content": "Summarise DashScope in one line."}],
)
print(resp.choices[0].message.content)

go deeper

for a junior

Be able to say the two things you change — base_url to the compatible-mode URL and api_key to the Model Studio key — and to write the four-line client setup from memory without reaching for a tutorial.

for a middle

Explain why the compatible surface exists at all, what the -intl infix selects, and why the key must be passed explicitly rather than left to the OpenAI SDK's environment fallback.

for a senior

Show you would make base_url and model id configuration rather than literals, keep the key in a secret manager, and pin snapshot ids so a vendor-side alias move cannot change production behaviour unannounced.

for a principal

Own the abstraction decision: whether the codebase talks to one OpenAI-shaped client with a swappable base URL, or to a provider interface, and what that choice costs when a vendor-specific parameter like a thinking flag has to leak through.

## What you are actually calling Alibaba Cloud **Model Studio** is the managed platform that hosts Qwen models (it is the same product marketed as Bailian inside China). **DashScope** is the API service behind it — the hostname you send requests to and the name of Alibaba's own Python SDK. Model Studio exposes two different request surfaces for the same models: - a **native DashScope surface** under `/api/v1/...`, with Alibaba's own request body shape, and - an **OpenAI-compatible surface** under `/compatible-mode/v1`, which speaks the Chat Completions request and response shape. The compatible surface exists so that code, libraries and frameworks written against OpenAI can talk to Qwen with a configuration change rather than a rewrite. That is why the answer to "how do I call Qwen from Python?" is usually "with the OpenAI SDK". ## The two things you change Construct the OpenAI client with `base_url` set to the compatible-mode URL for the region you signed up in, and `api_key` set to your Model Studio key. Nothing else in the call site changes: the same `client.chat.completions.create(model=..., messages=[...])` call, the same `response.choices[0].message.content`, the same `stream=True` iteration producing delta chunks. The two base URLs are: - Singapore / international: `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` - Beijing / mainland China: `https://dashscope.aliyuncs.com/compatible-mode/v1` Note the `-intl` infix — that single token is the region selector, and it is the most common copy-paste mistake when following a tutorial written for the other region. ## Auth on the wire A Model Studio API key is created in the console and is sent as an ordinary bearer token: `Authorization: Bearer sk-...`. There is no signature, no timestamp, no Alibaba Cloud AccessKey/Secret pair involved in the LLM call path — which is why generic OpenAI clients work unmodified. The convention is to export the key as `DASHSCOPE_API_KEY`, because Alibaba's own `dashscope` Python SDK reads that variable automatically. **The OpenAI SDK does not.** Its zero-argument fallback is `OPENAI_API_KEY`. If both variables happen to be set in a developer's shell, an `OpenAI(base_url=...)` constructed without `api_key` will happily send an OpenAI key to DashScope and fail authentication in a way that looks like a broken key rather than a broken wiring. Always pass `api_key=os.environ["DASHSCOPE_API_KEY"]` explicitly. ## Model ids The `model` field takes Qwen model ids: the commercial hosted tiers (`qwen-plus`, `qwen-turbo`, `qwen-max`) and hosted builds of the open-weight line. Ids come in two flavours — moving aliases such as `-latest`, and dated snapshot ids that pin a specific build. For anything whose output you regression-test, pin the snapshot and upgrade deliberately; aliases drift under you. Which tier to pick for which workload is a model-selection question; the API mechanics here are the same whichever id you send. ## Where the compatibility ends "OpenAI-compatible" means the envelope matches, not that every parameter exists. Streaming, tool calling and JSON output are supported for the models that advertise them; usage totals can be requested on the final stream chunk via `stream_options={"include_usage": True}`. Qwen-specific parameters have no place in the OpenAI request model, so the SDK's `extra_body` escape hatch carries them — `enable_thinking` and `thinking_budget` for hybrid-thinking models, `enable_search` for the built-in web search. Conversely, some OpenAI parameters are not honoured on every Qwen model, and DashScope may reject them with a parameter error rather than ignore them, so treat per-model support as something to verify rather than assume. ## The failure modes you will actually hit - **401 / invalid API key** — wrong region's key, or the OpenAI env-var fallback described above. - **Model not found** — the id is real but is not served in the region you called; catalogues differ between Beijing and Singapore. - **429** — you crossed a per-account throughput quota; back off and retry with jitter. ## Practical hygiene Build the client in one place, take `base_url` from configuration rather than hard-coding it, keep the key in a secret manager, and pin snapshot model ids in anything with tests. Doing those four things makes switching region, or switching between the compatible surface and the native one, a config change instead of a code change.

  • You set base_url correctly but still get a 401 — what is the first thing you check?
    Whether the key was actually passed. The OpenAI client's default is `OPENAI_API_KEY`; it never falls back to `DASHSCOPE_API_KEY`, so an unset `api_key` argument sends whatever OpenAI key is in the shell. After that, check the key's region: a Beijing-issued Model Studio key is unknown to the Singapore host and vice versa.
  • How do you pass a Qwen-only parameter that the OpenAI request model has no field for?
    Through the SDK's `extra_body` dict, which is merged into the JSON body verbatim — for example `extra_body={"enable_thinking": True}`. Adding it as an unknown keyword argument raises a client-side error, and putting it in the messages array just sends it to the model as text.
  • Should you use a `-latest` alias or a dated snapshot model id in production?
    A dated snapshot, for anything whose output you assert on. Aliases move when Alibaba ships a new build, which silently changes behaviour, formatting and sometimes token cost mid-deployment. Use the alias in exploration, pin the snapshot in services, and upgrade as a reviewed change with your evaluation suite as the gate.

saying these in an interview costs you the question

  • Thinks the OpenAI SDK reads DASHSCOPE_API_KEY automatically
  • Assumes one global DashScope URL with no region in it
  • Believes every OpenAI parameter is honoured by Qwen
  • Signs requests with an Alibaba Cloud AccessKey pair instead of a bearer key
  • Puts Qwen-only flags as top-level SDK keyword arguments

context

open as a page

Which Qwen variant do you pick for coding, vision or embedding work?

level: juniorimportance: must knowfreq 62%

basics

~10 s

Qwen ships task-specific lines beside its general chat models: Qwen3-Coder for code and agentic coding, Qwen3-VL for images and video, Qwen3-Embedding and Qwen3-Reranker for retrieval, and the Omni/Audio line for speech input.

open as a page

How do you serve Qwen weights behind an OpenAI-compatible endpoint with vLLM?

level: juniorimportance: must knowfreq 60%

basics

~20 s

vllm serve Qwen/Qwen3-8B pulls the weights from Hugging Face and starts an OpenAI-shaped HTTP server on port 8000. Point any OpenAI client at http://localhost:8000/v1, pass any non-empty API key, and set model to the served name.

open as a page

In the Qwen model name Qwen3-30B-A3B-Instruct-2507, what does each part mean?

level: middleimportance: must knowfreq 54%

basics

~20 s

Qwen3 is the generation, 30B the total parameter count, A3B the roughly 3B parameters active per token in a Mixture-of-Experts layout, Instruct the post-trained chat checkpoint (versus Base or Thinking), and 2507 the July 2025 refresh of that checkpoint.

open as a page

What does Qwen's ChatML chat template wrap around each message, and why does it matter?

level: middleimportance: must knowfreq 55%

basics

~20 s

Qwen uses ChatML: every message is rendered as <|im_start|>role, a newline, the content, then <|im_end|>. Generation starts from a trailing <|im_start|>assistant header, and <|im_end|> is the stop token. Wrong formatting means the model rambles or never stops.

open as a page

How much GPU memory does a self-hosted Qwen 32B model need beyond its weights?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Weights are bytes-per-parameter times parameters: about 64 GB for 32B at BF16, near 20 GB at 4-bit. On top sit the KV cache, which grows with context length and concurrency, plus activation and graph buffers — often 10 GB or more.

open as a page

On DashScope, why must a Qwen3 request with enable_thinking use streaming?

level: middleimportance: should knowfreq 40%

basics

~20 s

DashScope surfaces Qwen3 reasoning only over a stream: with enable_thinking set, a non-streaming call is rejected rather than silently downgraded. Pass the flag through extra_body, set stream=True, and read the reasoning from delta.reasoning_content before the normal delta.content arrives.

open as a page

Why does an Alibaba DashScope API key fail against the other region's endpoint?

level: middleimportance: should knowfreq 44%

basics

~20 s

Model Studio runs as two independent deployments — Beijing on dashscope.aliyuncs.com and Singapore on dashscope-intl.aliyuncs.com — with separate consoles, separate API keys and separate model catalogues. A key minted in one region is simply unknown to the other.

open as a page

How do Qwen3's hybrid thinking mode and the separate Instruct/Thinking checkpoints differ?

level: middleimportance: should knowfreq 46%

basics

~20 s

Qwen3 originally shipped hybrid checkpoints that could answer with or without an extended reasoning trace. The later refreshes dropped that and publish two specialised checkpoints instead — an Instruct model that never reasons at length and a Thinking model that always does — so the choice moves from runtime to deployment.

open as a page

When do you need DashScope's native API instead of its OpenAI-compatible endpoint?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Reach for DashScope's native /api/v1 surface when the OpenAI shape has no room for what you need — asynchronous task submission for long-running generation, endpoints that are not chat-shaped, and native-only request knobs. Everything conversational is simpler in compatible mode.

open as a page

Are all Qwen model weights Apache-2.0, and which releases are not?

level: seniorimportance: should knowfreq 40%

basics

~20 s

No. Most of the Qwen3 open-weight releases are Apache-2.0, but the Qwen2.5 generation put some tiers under a separate Qwen licence and one under a research-only licence, and the top hosted Max tier has no downloadable weights at all. Check each checkpoint's own licence file.

open as a page

Self-hosted Qwen returns tool calls as plain text in vLLM — which flags fix it?

level: seniorimportance: should knowfreq 42%

basics

~10 s

Start vLLM with --enable-auto-tool-choice and --tool-call-parser hermes. Qwen emits function calls as JSON inside <tool_call> tags in the assistant text; without a parser the server passes that through as content and leaves tool_calls null.

open as a page

When do you enable YaRN rope_scaling on a self-hosted Qwen model, and how?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Only when inputs genuinely exceed the checkpoint's native window. You add a yarn rope_scaling block naming the factor and the original max position to config.json, or pass the same override on the serving command line, then raise the served max context accordingly.

open as a page

How do you choose between Alibaba's hosted Qwen API and running Qwen weights yourself?

level: principalimportance: should knowfreq 38%

basics

~20 s

Hosted Model Studio wins on bursty or modest volume, fast model turnover and zero GPU operations. Self-hosting wins when prompts must stay inside your own network, when you need weights you control, or when sustained volume makes fixed GPU cost cheaper than per-token billing.

open as a page

When is a portfolio of small Qwen variants better than one large Qwen model?

level: principalimportance: should knowfreq 32%

basics

~20 s

When traffic splits into distinct, high-volume tasks with different modality or latency needs, several small specialists cost less and respond faster than one flagship. When traffic is mixed, low-volume or unpredictable, one capable general model wins on operational simplicity.

open as a page