Self-hosted Qwen returns tool calls as plain text in vLLM — which flags fix it?
answer
- structure has to be recovered from text
- two flags, enabled together
- JSON inside tags in the assistant turn
- the parser name matches the syntax family
- coder variants encode calls differently
basics
~10 sStart vLLM with --enable-auto-tool-choice and --tool-call-parser hermes. Qwen emits function calls as JSON inside <tool_call> tags in the assistant text; without a parser the server passes that through as content and leaves tool_calls null.
solid answer
~50 sQwen2.5 and Qwen3 do not emit OpenAI-style structured tool calls natively — their chat template trains them to write a JSON object wrapped in `<tool_call>` … `</tool_call>` tags inside the ordinary assistant turn, in the Hermes convention. vLLM will happily hand that back verbatim in `message.content` unless you tell it to extract it, so the server must be started with `--enable-auto-tool-choice --tool-call-parser hermes`. With those flags the parser strips the tags, populates `message.tool_calls` with an id, function name and JSON arguments, sets `finish_reason` to `tool_calls`, and does the same incrementally for streamed responses. Two related gotchas: the checkpoint's chat template must actually render the `tools` array (Qwen's shipped template does, and serving frameworks ship replacement templates for checkpoints whose template does not), and Qwen's coder-specialised releases use a different tool-call encoding that needs its own matching parser rather than the Hermes one.
code
bash · 4 linesvllm serve Qwen/Qwen3-8B \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--port 8000go deeper
Know that a self-hosted server has to be told to extract tool calls — the model writes them as text, and without the right startup flags they arrive as ordinary content.
Explain the pair of flags and what the parser actually does to the response: strips the tag syntax, fills the tool-calls array, and changes the finish reason.
Show the diagnosis: distinguish a parsing failure from a template failure from malformed arguments, and check the streaming path separately from the non-streaming one.
Own the contract risk — tool calling on self-hosted weights depends on three independently versioned pieces that fail silently, so make an end-to-end tool round trip a release gate rather than a manual check.
## The gap between the model and the API contract The OpenAI-shaped API says an assistant message carries a structured `tool_calls` array. Open-weight models do not produce structured fields — they produce tokens. Everything between those two facts is the serving layer's job, and it is not automatic. Qwen's instruct models are trained in the Hermes-style convention: when the model decides to call a function, it writes, inside the ordinary assistant turn, something like <tool_call> {"name": "get_weather", "arguments": {"city": "Lima"}} </tool_call> and results come back into the conversation rendered by the template into its own tool-response tags. If the server does no extraction, that text is simply the assistant's content. Your client sees `tool_calls: null`, `finish_reason: "stop"`, and a body of text containing angle-bracket tags — which is precisely the bug report "tool calling doesn't work on my self-hosted Qwen". ## The two flags vllm serve Qwen/Qwen3-8B \ --enable-auto-tool-choice \ --tool-call-parser hermes `--enable-auto-tool-choice` allows the model to decide for itself whether to call a tool — the OpenAI `tool_choice: "auto"` behaviour — rather than rejecting or ignoring the request's tool settings. `--tool-call-parser` names the extraction strategy: which surface syntax to look for and how to map it into the response fields. `hermes` is the parser matching the `<tool_call>`-tag convention that Qwen's own template uses. Neither flag alone is enough; they are enabled as a pair. With them set, the server does four things: strips the tags out of `content`, emits `message.tool_calls[]` with a generated `id`, `function.name` and `function.arguments` as a JSON *string*, sets `finish_reason` to `tool_calls`, and applies the same extraction to streaming deltas so partial arguments accumulate the way OpenAI clients expect. ## The template half of the problem Parsing is only the response side. The request side needs the `tools` array rendered into the prompt so the model knows what it may call. That is the chat template's job, and Qwen's shipped template handles it — the tool schemas go into the system turn and `role: "tool"` messages are rendered back into the conversation. When a checkpoint's template lacks tool support, serving frameworks ship replacement Jinja templates you point at with a `--chat-template` argument. The failure signature here differs from the parsing one: the model never attempts a call at all, because it was never told the tools exist. ## Variant-specific parsers A parser is tied to a surface syntax, not to a vendor. Qwen's coder-specialised releases were trained to emit function calls in a different, more XML-ish encoding than the Hermes JSON-in-tags form, so pointing the Hermes parser at them yields either nothing extracted or malformed arguments. Those checkpoints need the parser that matches their own format. The general rule when self-hosting: check the model card for the tool-call format, then select the parser named for it — and validate with one real round trip rather than assuming. ## Debugging it in production The diagnosis path is short because the failure modes are distinguishable: - `content` contains tag syntax and `tool_calls` is null → parser not enabled or wrong parser for this checkpoint. - Model answers in prose and never attempts a call → tools were not rendered into the prompt (template gap), or the request omitted `tools`. - Arguments arrive as invalid JSON, or truncated → often a `max_tokens` cut mid-call, or a small model genuinely producing malformed JSON; validate arguments before executing and be ready to re-prompt. - Works non-streaming, breaks streaming → an incremental-parsing edge case; test both paths before shipping. ## Why this bites people Because the API is OpenAI-shaped, teams assume the semantics are complete out of the box. They are not: tool calling on a self-hosted open-weight model is an agreement between three independently configurable things — the model's trained syntax, the chat template that advertises the tools, and the server-side parser that recovers structure from text. Any one of them mismatched degrades silently into plain text rather than raising an error, which is why an end-to-end tool round trip belongs in your deployment smoke test.
- How do you tell a parser problem apart from a chat-template problem?Look at what the model produced. If `content` contains the raw tool-call tag syntax, the model tried to call a tool and the server failed to extract it — a parser issue. If the model answers in prose and never emits any call syntax, it was never told the tools exist, which points at the template not rendering the `tools` array (or the request omitting them).
- What changes about tool-call parsing when the client streams the response?The parser must work incrementally: it sees the tag opening and the JSON arguments arrive token by token, and has to emit partial `tool_calls` deltas that the client concatenates. That path has its own edge cases, so a deployment that passes a non-streaming tool test can still break under streaming. Test both, and have the client tolerate arguments that are only valid JSON once complete.
- Should you trust the arguments object the parser returns?No. The parser guarantees extraction, not validity — the model can emit arguments that miss required fields, use wrong types, or invent an enum value, and the arguments come back as a JSON string. Parse it, validate against the same schema you advertised, and treat a failure as a re-prompt with the validation error rather than as an executable call.
saying these in an interview costs you the question
- Assuming an OpenAI-shaped server does tool parsing out of the box
- Enabling the parser but not auto tool choice, or the reverse
- Using one tool-call parser for every open-weight checkpoint
- Blaming the model when tool syntax shows up as message content
- Executing parsed arguments without validating them against the schema