skip to content

In vLLM, which flags turn on OpenAI-style tool calling for a chat model?

level: middleimportance: must knowfreq 55%

answer

  1. Tool syntax is per model family
  2. Two flags, always together
  3. Something must translate text into JSON
  4. Parser name must match the checkpoint
  5. Wrong parser fails silently, not loudly

basics

~20 s

Two flags together: --enable-auto-tool-choice and --tool-call-parser with the parser matching the served model's family. The parser converts the model's own tool-call text into the tool_calls field of the response; without it, calls arrive as ordinary content.

solid answer

~50 s

vLLM cannot turn tool calling on generically, because every model family emits tool calls in its own syntax — special tokens, tagged blocks, or Python-like call expressions. So you pass `--enable-auto-tool-choice` to accept `tool_choice: "auto"` from clients, plus `--tool-call-parser <name>` naming a parser that matches the checkpoint (`hermes`, `mistral`, `llama3_json`, `pythonic` and others, with `--tool-parser-plugin` for a custom one). The parser extracts the model's emission into `choices[0].message.tool_calls` with `function.name` and `function.arguments` as a JSON string, and sets `finish_reason` to `"tool_calls"`. The model's chat template must also render the request's `tools` — if it ignores them, the model was never told the tools exist. Pick the wrong parser and nothing errors: the call text stays in `content` and `tool_calls` is empty. When you need guaranteed structure rather than model-authored calls, use `response_format` with a JSON schema, which constrains decoding instead of parsing after the fact.

code

bash · 3 lines
bash
vllm serve NousResearch/Hermes-3-Llama-3.1-8B \
  --enable-auto-tool-choice \
  --tool-call-parser hermes

go deeper

for a junior

Know that tool calling has to be enabled on the server, that it needs a parser matched to the model, and that a successful call comes back in message.tool_calls with finish_reason set to tool_calls.

for a middle

Name --enable-auto-tool-choice and --tool-call-parser together, explain that each model family emits calls in its own syntax so the parser translates it into the OpenAI shape, and note that function.arguments is a JSON string.

for a senior

Show the debugging path: a mismatched parser returns 200 with empty tool_calls, a template that drops the tools argument never informs the model, and streaming extraction is a separate code path worth testing on its own before rollout.

for a principal

Weigh model-authored tool calls against schema-constrained decoding for reliability, and set the policy for how a checkpoint upgrade is validated — the correct parser is a per-checkpoint fact, so tool-call fidelity belongs in the release gate, not in production discovery.

## Why this is not just a switch The OpenAI-shaped API pretends tool calling is a property of the protocol: you send `tools`, and you get back `tool_calls`. Underneath, it is entirely a property of the *model*. A tool-calling checkpoint has been fine-tuned to emit a call in one specific textual form — one family wraps a JSON object in special tool tokens, another emits a tagged block, another writes what looks like a Python function call. There is no shared wire format. A server that wants to return a structured `tool_calls` array therefore needs two things it cannot infer: permission to accept `tool_choice: "auto"`, and a parser that knows this family's syntax. That is exactly the pair of flags. `--enable-auto-tool-choice` opts the server into automatic tool selection; `--tool-call-parser <name>` selects the extractor. They go together — enabling one without the other is a configuration error, and clients sending `tools` with `tool_choice: "auto"` to a server without them get a rejection rather than a silent degradation. ## What the parser actually does On a non-streaming request the parser runs over the finished generation, finds the tool-call region, and rewrites the response so that `choices[0].message.content` holds any surrounding prose (often null) and `choices[0].message.tool_calls` holds one entry per call, each with an `id`, a `function.name`, and `function.arguments`. Note that `arguments` is a **JSON string**, not a nested object — that is the OpenAI shape, and clients must parse it themselves. `finish_reason` becomes `"tool_calls"`, which is how a client loop knows to execute the tool rather than show the text to a user. Streaming is a separate, harder code path in the same parser: it must decide, incrementally, that the tokens arriving are the beginning of a call, and emit deltas carrying `tool_calls[].index`, then the name, then fragments of the arguments string. Some parsers are more robust here than others, and streaming tool calls are a common place to find rough edges when a model family is newly supported. ## The chat template is half the feature A parser handles the model's *output*. The model's *input* — how the available tools are described to it — comes from the chat template, which receives the request's `tools` as a rendering argument and is responsible for laying them out in whatever form the checkpoint was tuned on. If a template ignores the `tools` argument, the tools never reach the model at all, and you will see the model answering in prose while your parser dutifully finds nothing to extract. When debugging a tool-calling setup, verify the rendered prompt first; the same `/tokenize` route that shows a chat prompt shows the tool block. ## The silent failure mode Mismatching the parser to the model is the classic mistake, and it does not raise. The server starts, requests succeed with a 200, the model happily emits its call in its own syntax, the parser looks for a syntax that never appears, and `tool_calls` stays empty while the raw call text sits in `content`. Downstream, an agent loop sees a normal assistant message and either loops or answers nonsense. Always test one tool-calling request end-to-end after any model change and assert on `finish_reason == "tool_calls"` rather than on the response merely arriving. The parser name must track the checkpoint, not the vendor. Two models from the same organization can use different formats across generations, and the set of parsers grows every release as families are added, so the correct value is a per-checkpoint fact you look up rather than remember. ## When parsing is the wrong tool If what you actually want is a guaranteed structure — a filled-in record rather than a genuine decision by the model to call something — do not rely on free-text parsing at all. Send `response_format` with a JSON schema so decoding is constrained to that grammar. Structure then comes from the decoder rather than from a post-hoc extractor, so it cannot be malformed, at the cost of the model no longer choosing whether to call. Use tool calling when the choice matters; use constrained decoding when the shape matters. ## Reasoning models Checkpoints that emit a separate thinking block complicate extraction, because the reasoning text may itself contain something that looks like a call. vLLM exposes `--reasoning-parser` to split that block into a `reasoning_content` field, keeping it out of the answer the tool-call parser inspects. If you serve a reasoning model with tools, configure both parsers, and expect the combination to need a real end-to-end test before you trust it.

  • How does a client tell that the model chose to call a tool rather than answer?
    `finish_reason` on the choice becomes `"tool_calls"`, and `message.tool_calls` is populated with one entry per call carrying an `id`, `function.name`, and `function.arguments` as a JSON string. The agent loop executes the tool, appends a message with role `tool` and the matching `tool_call_id`, and calls the endpoint again.
  • Tool calling returns 200 but tool_calls is always empty. Where do you look first?
    Two places, in order. Confirm the chat template actually renders the request's `tools` — check the prompt through `/tokenize`; if the tools are absent, the model never knew about them. If they are present, the parser is mismatched to the model family: the call text will be sitting in `content` in the model's native syntax, which tells you which parser you should have picked.
  • When would you use a JSON schema in response_format instead of tool calling?
    When you need a guaranteed shape rather than a decision. Constrained decoding forces the output to satisfy the schema, so it cannot be malformed and no extractor can miss it. Tool calling is the right choice when the model must decide whether and which function to invoke; schema-constrained output is right when you already know you want one structured record back.

saying these in an interview costs you the question

  • Thinks any model supports tool calling once tools are sent
  • Believes vLLM auto-detects the model's tool format
  • Expects function.arguments to be a parsed object
  • Sets the parser without --enable-auto-tool-choice
  • Ignores that the chat template must render the tools

context