skip to content

How does xAI's API expose reasoning effort and reasoning traces on Grok?

level: middleimportance: should knowfreq 38%

answer

  1. Coarse dial, not a token budget
  2. Only the lighter reasoning line takes it
  3. Trace arrives beside the answer, not inside it
  4. Thinking tokens bill as output
  5. Never replay the trace as history

basics

~20 s

On xAI's lighter Grok reasoning models you set reasoning_effort to "low" or "high" to trade depth against latency and cost. The trace comes back in the message as reasoning_content, and thinking tokens are billed as completion tokens with a reasoning_tokens breakdown in usage.

solid answer

~50 s

xAI splits reasoning control across model lines. The smaller reasoning models — the grok-3-mini line — accept a `reasoning_effort` parameter with values `"low"` and `"high"`, which is a coarse budget dial: high spends more internal tokens and wall-clock time before answering. The flagship grok-4 family does not take the parameter at all, because it reasons as part of normal operation and has no reduced mode to select. Where a trace is exposed, it arrives as `reasoning_content` on the assistant message, separate from `content`, so your rendering code must decide deliberately whether to show it. Billing-wise, thinking tokens are output tokens: they are charged as completion tokens and reported under `usage.completion_tokens_details.reasoning_tokens`. Two consequences follow — a "short" answer can be expensive, and you should not feed `reasoning_content` back as conversation history on the next turn; send only the final `content`.

code

python · 15 lines
python
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["XAI_API_KEY"], base_url="https://api.x.ai/v1")

response = client.chat.completions.create(
    model="grok-3-mini",
    messages=[{"role": "user", "content": "Is 8191 prime? Show the check."}],
    reasoning_effort="high",
)

message = response.choices[0].message
print(getattr(message, "reasoning_content", None))
print(message.content)
print(response.usage.completion_tokens_details.reasoning_tokens)

go deeper

for a junior

Know that some Grok models take a reasoning_effort setting of low or high, and that any thinking text comes back separately from the actual answer.

for a middle

Explain the split across model lines, name reasoning_content as the trace field, and say that thinking tokens bill as completion tokens with a reasoning_tokens breakdown in usage.

for a senior

Show the operating discipline: route by task instead of one global effort setting, chart reasoning tokens as a real cost line, size output budgets for trace plus answer, and retune timeouts.

for a principal

Own the tradeoff — what accuracy the extra effort actually buys on your evaluation set, how requests are routed between model lines, and how effort control stays portable when every vendor names it differently.

## Two different things called reasoning When people say a model "reasons", two separable mechanisms are in play, and the API exposes them differently. 1. **Internal thinking tokens** the model generates before its visible answer. They cost money and time, and they are the thing an effort dial controls. 2. **A trace of that thinking** returned to the caller, which is a presentation and debugging concern, not a capability. xAI's surface reflects this split. ## Effort control The smaller Grok reasoning line accepts **`reasoning_effort`** with the values `"low"` and `"high"`. It is intentionally coarse — not a token count, not a millisecond budget, but a hint about how hard to think. `low` returns faster and cheaper and is right for classification, extraction, routing, and anything where the answer is looked up rather than derived. `high` is for multi-step arithmetic, constraint satisfaction, careful code reasoning, and cases where you would otherwise have prompted the model to work step by step. The flagship **grok-4 family does not accept the parameter**. That is not an oversight: the model reasons as part of its normal generation, so there is no lower gear to shift into, and xAI documents the field as unsupported there. This is the practical trap — code that sets `reasoning_effort` unconditionally works on one Grok line and fails on another. Bind the parameter to the model in configuration, not to a global default. ## Reading the trace Where a trace is exposed, it arrives on the assistant message as **`reasoning_content`**, alongside the ordinary `content`. Treat them very differently: - **`content`** is the answer. It is what you render, what you parse, and the only thing you append to conversation history. - **`reasoning_content`** is diagnostic. It is useful in logs when you are investigating a wrong answer, and it is occasionally shown in a collapsed "thinking" panel in developer tools. It is not a contract: its shape, verbosity and even presence vary with model line and generation, so never parse it for structured data. Crucially, **do not feed the trace back**. On the next turn, send the prior `content` only. Replaying thinking as history inflates input tokens on every subsequent turn, and the model was not trained to consume its own trace as context — you get cost and confusion for no benefit. ## What it costs Thinking tokens are **completion tokens**. They are billed at the output rate and reported in the usage breakdown as `usage.completion_tokens_details.reasoning_tokens`. Two operational consequences: - **Short answers can be expensive.** A three-word reply that took a thousand internal tokens costs like a thousand-token reply. If you monitor only visible output length, your cost model is wrong. Chart reasoning tokens explicitly. - **Latency follows the same curve.** Time-to-first-visible-token stretches with effort, because thinking happens before the answer streams. Streaming still helps the perceived experience, but the initial silence is longer, and timeouts tuned on a non-reasoning model will misfire. Budget your output cap with the trace in mind: a limit sized only for the visible answer leaves no room for the thinking that precedes it. ## Choosing settings for a workload A useful default is to route by task rather than pick one global setting. Cheap, high-volume, shallow calls — routing, tagging, extraction — go to the small model at `low`. Genuinely hard single calls go to `high`, or to the flagship line if the problem needs its capability, accepting that you lose the dial there. Then measure: run your evaluation set at each setting and compare accuracy against cost and p95 latency. Effort is one of the few knobs where the tradeoff is legible enough to tune empirically, and teams routinely find that `high` buys nothing on a task whose difficulty is retrieval rather than deduction. ## Portability note Every major provider now has some version of this idea, and none of them agree on the field names, the value sets, or whether the trace is visible. Keep effort and trace handling behind your own small abstraction so switching providers is a mapping change rather than a rewrite of every call site.

  • Why should you not append reasoning_content to the message history on the next turn?
    Because it is diagnostic output, not conversational context. Replaying it re-bills every prior thinking token as input on each subsequent turn, and the model is not trained to consume its own trace as history, so accuracy does not improve. Send only the final `content`. If you need the reasoning later — for audit or debugging — store it out of band in your own logs rather than inside the prompt.
  • Your p95 latency doubled after enabling high effort. How do you decide whether to keep it?
    Measure the accuracy it actually buys. Run your evaluation set at both settings and compare correctness against p95 latency and reasoning-token cost. If the gain is inside the noise — common when the task's difficulty is retrieval or formatting rather than deduction — revert to low. If it is real, consider routing: send only the hard classes of request at high effort rather than paying the latency on all traffic.
  • How do reasoning tokens affect how you size the output limit?
    They consume the same output budget as the visible answer, so a cap sized only for the reply can leave the model no room to think, or truncate the response after the thinking is already paid for. Size the limit for trace plus answer, watch `completion_tokens_details.reasoning_tokens` to see how much of it the thinking actually takes, and treat unexpectedly truncated answers on a reasoning model as a budgeting bug first.

saying these in an interview costs you the question

  • Setting reasoning_effort on every Grok model regardless of line
  • Assuming thinking tokens are free because they are internal
  • Parsing reasoning_content as if it were structured output
  • Sending the trace back as conversation history
  • Judging cost by visible answer length alone

context