skip to content

On DashScope, why must a Qwen3 request with enable_thinking use streaming?

level: middleimportance: should knowfreq 40%

answer

  1. One mode needs one transport
  2. No field for it in the OpenAI schema
  3. Two different delta fields
  4. The default is not the same everywhere
  5. Rejected, not downgraded

basics

~20 s

DashScope surfaces Qwen3 reasoning only over a stream: with enable_thinking set, a non-streaming call is rejected rather than silently downgraded. Pass the flag through extra_body, set stream=True, and read the reasoning from delta.reasoning_content before the normal delta.content arrives.

solid answer

~50 s

Qwen3's hybrid models can answer in a thinking mode that emits a reasoning trace before the answer, and on Model Studio that mode is documented as streaming-only: send `enable_thinking` with `stream=False` and DashScope returns a parameter error instead of quietly falling back. Mechanically, the flag has no field in the OpenAI request model, so with the OpenAI SDK it rides in `extra_body`, alongside `thinking_budget` if you want to cap how many tokens the model may spend reasoning. On the way back, each chunk's `delta` carries `reasoning_content` while the model is thinking and `content` once it starts answering, so a consumer that reads only `delta.content` will see an eerie silence followed by the answer. Defaults differ by model — the hybrid commercial ids are off by default, hosted open-weight Qwen3 builds are on — so set the flag explicitly rather than relying on it. This describes Model Studio as of mid-2026.

code

python · 16 lines
python
stream = client.chat.completions.create(
    model="qwen-plus",
    messages=[{"role": "user", "content": "Is 9.11 larger than 9.9?"}],
    stream=True,
    extra_body={"enable_thinking": True, "thinking_budget": 2000},
)

reasoning, answer = [], []
for chunk in stream:
    delta = chunk.choices[0].delta
    if getattr(delta, "reasoning_content", None):
        reasoning.append(delta.reasoning_content)
    elif delta.content:
        answer.append(delta.content)

print("".join(answer))

go deeper

for a junior

Know that thinking mode on Model Studio needs stream=True, that the flag goes in extra_body, and that the reasoning text arrives in a different field from the answer.

for a middle

Explain all three mechanics — streaming-only, extra_body placement, reasoning_content versus content — and note that the flag's default depends on which Qwen model id you called.

for a senior

Show the operational consequences: reasoning tokens cost money and latency, thinking_budget is the control, batch-shaped pipelines have to accumulate the stream, and only the answer belongs in conversation history.

for a principal

Own the policy question of when reasoning mode is worth its variable latency and token bill at all, and how the trace is treated — surfaced, logged, or dropped — given that it is model scratch work you may not want to retain or show.

## What the flag controls Qwen3 introduced hybrid models that can run in two modes: a fast direct mode, and a **thinking mode** in which the model first produces an internal reasoning trace and then the user-facing answer. On Alibaba Cloud Model Studio the switch is the request parameter `enable_thinking`, and the companion `thinking_budget` caps how many tokens the reasoning phase may consume before the model is pushed to answer. What thinking mode *is* as a modelling technique is general LLM material; what matters at this API surface is the three concrete mechanics below. ## Mechanic 1: it is streaming-only The headline constraint is that DashScope serves thinking output over a stream. A request that sets `enable_thinking` true with `stream` false is **rejected with a parameter error**, not silently downgraded to a non-thinking answer. This is the single most common first failure, and it is a good failure: a silent downgrade would give you a plausible answer produced by a different mode than the one you asked for, and you would never know. The consequence is architectural rather than cosmetic. Any code path that wants Qwen3 reasoning must be able to consume Server-Sent Events — which rules out the convenient one-shot request in a synchronous job, a serverless function with a short response budget, or a client library configured for simple request/response. If your pipeline is batch-shaped, you either accumulate the stream server-side and hand the joined result onward, or you do not use thinking mode. ## Mechanic 2: the flag has nowhere to live in the OpenAI request model Because `enable_thinking` is Qwen-specific, it does not exist as a field in the OpenAI Chat Completions schema. With the OpenAI SDK against DashScope's compatible mode you therefore pass it through **`extra_body`**, a dict merged verbatim into the JSON body — `extra_body={"enable_thinking": True, "thinking_budget": 2000}`. Passing it as an unknown keyword argument fails client-side before a request is ever made; putting it in a system message merely tells the model about a word. On the native DashScope surface the same flag lives inside the request's `parameters` object instead. ## Mechanic 3: reasoning arrives in a separate field During the thinking phase, each streamed chunk's `delta` carries **`reasoning_content`** and an empty `content`. When the model switches to answering, `reasoning_content` stops and `content` starts. A consumer written for ordinary streaming — one that appends `delta.content` and ignores everything else — will therefore render nothing at all for the entire reasoning phase, then produce the answer in a burst. Users read that as a hang. The fix is to branch on which field is populated: display or discard `reasoning_content` deliberately, and accumulate `content` as the answer. That also means reasoning tokens are real tokens. They are generated, they take time, and they are billed and counted against the response budget — which is exactly what `thinking_budget` exists to bound when latency or cost matters more than the last increment of quality. ## Defaults differ by model id The default value of `enable_thinking` is not uniform across the catalogue. The hybrid commercial ids default to thinking **off**, while hosted builds of the open-weight Qwen3 models default to thinking **on**. So the same code, pointed at a different model id, can flip between fast answers and a stream that begins with a long reasoning phase — and if that second model is called non-streaming, it fails outright. Set the flag explicitly on every request. Relying on a default here is how a model-id change becomes a production incident. ## Do not feed the trace back When you build the next turn's `messages`, send the **answer**, not the reasoning trace. Reasoning content is scratch work: replaying it wastes context, changes the distribution the model conditions on, and can confuse the next turn. Store it separately if you want it for debugging or evaluation. ## A checklist for wiring it up - Set `enable_thinking` explicitly, per request, from configuration — never rely on the model's default. - Always pair it with `stream=True`; make that invariant impossible to violate in your client wrapper. - Branch on `reasoning_content` versus `content` in the chunk handler, and decide deliberately whether the trace is shown, logged or dropped. - Set `thinking_budget` when you have a latency SLO, and measure the quality cost of tightening it. - Keep only the answer in conversation history.

  • What does thinking_budget do, and what happens when the budget runs out?
    It caps how many tokens the reasoning phase may consume before the model is pushed to produce its answer, giving you a latency and cost bound on the variable part of the response. Setting it low trades answer quality on hard prompts for predictable response times; the right value is measured against your own evaluation set, not guessed.
  • Should the reasoning trace be appended to the conversation history for the next turn?
    No — send the answer only. The trace is scratch work: replaying it burns context window, shifts the distribution the next turn conditions on, and tends to make follow-ups worse rather than better. Keep it out of messages, and persist it separately if you want it for debugging or offline evaluation.
  • Your UI shows nothing for ten seconds and then the whole answer at once. What is wrong in the client?
    The chunk handler is appending only `delta.content`, so the entire reasoning phase renders as silence. Branch on the delta fields: while `reasoning_content` is populated, either stream it into a collapsed thinking panel or at minimum drive a progress indicator, then switch to accumulating `content` when the answer begins.

saying these in an interview costs you the question

  • Expects a non-streaming thinking request to silently return a normal answer
  • Reads only delta.content and reports the model as hung
  • Assumes enable_thinking defaults the same way for every Qwen model
  • Passes enable_thinking as a top-level OpenAI SDK argument
  • Appends the reasoning trace to the next turn's messages

context