How do you enable xAI's Live Search in a Grok chat completions request?
answer
- A request parameter, not a prompt trick
- search_parameters on the request body
- mode: on, auto, or off
- Sources: web, x, news, rss
- Billed per source, not just tokens
basics
~20 sSend a search_parameters object on the chat completions request. Its mode field turns search on, off, or leaves the decision to the model; sources selects web, x, news or rss. The response carries a citations array, and each source used is billed on top of tokens.
solid answer
~50 sLive Search is an xAI request-level feature, not a prompting trick: you add a `search_parameters` object to the chat completions body. `mode` takes `"on"` (always search), `"auto"` (let Grok decide) or `"off"`; `sources` is a list of typed source objects — `web`, `x`, `news`, `rss` — each of which can be narrowed further, for example by excluding sites or restricting to given X handles. `from_date` and `to_date` bound the freshness window, and `max_search_results` caps how many sources a single call may pull. The response comes back with a `citations` array of the URLs used, and the `usage` object reports `num_sources_used`. That last field matters commercially: sources are billed per source in addition to input and output tokens, so an unbounded `mode: "on"` deployment has a cost axis that token accounting alone will not show you. With the OpenAI SDK, `search_parameters` must be passed through `extra_body`.
code
python · 21 linesimport os
from openai import OpenAI
client = OpenAI(api_key=os.environ["XAI_API_KEY"], base_url="https://api.x.ai/v1")
response = client.chat.completions.create(
model="grok-4",
messages=[{"role": "user", "content": "What changed in EU AI regulation this month?"}],
extra_body={
"search_parameters": {
"mode": "on",
"sources": [{"type": "web"}, {"type": "news"}],
"max_search_results": 8,
"return_citations": True,
}
},
)
print(response.choices[0].message.content)
print(getattr(response, "citations", None))
print(response.usage.model_dump().get("num_sources_used"))go deeper
Know that xAI turns live retrieval on with a search_parameters object on the request, that mode chooses on, auto or off, and that citations come back with the answer.
Explain the field set — mode, sources, date window, max_search_results — and be able to say that with the OpenAI SDK it travels in extra_body because it is not an OpenAI parameter.
Demonstrate the operational view: per-source billing tracked via num_sources_used, wider latency variance, non-determinism in evaluations, and treating retrieved pages as untrusted input.
Own the policy — which routes may search at all, what source allowlist and freshness window the product needs, and how per-source cost is budgeted and attributed when it is invisible to token dashboards.
## What Live Search is Most providers make retrieval your problem: you fetch documents, you stuff them into the prompt. xAI's differentiator is that live retrieval is a **first-class request parameter**. You ask for an answer and, in the same call, tell the server it may go and read the live web, X, news feeds or specified RSS links before answering. There is no separate search endpoint to orchestrate and no second round trip in your code. ## The request shape Search is off unless you opt in by sending a `search_parameters` object alongside the usual `model` and `messages`. The fields that matter: - **`mode`** — `"on"` searches on every request; `"auto"` lets the model judge whether the question needs fresh information; `"off"` disables it. `auto` is the sane default for a mixed workload, because most conversational turns do not need the web and every avoided search is money not spent. - **`sources`** — a list of objects each with a `type`: `web`, `x`, `news`, or `rss`. Leaving it out uses xAI's defaults; setting it explicitly is how you keep a support bot on documentation sites, or a social-listening tool on X. Individual source types accept their own narrowing options, such as excluding particular websites, restricting to named X handles, or listing RSS links. - **`from_date` / `to_date`** — an ISO date window that bounds how old retrieved material may be. Essential for "what happened this week" workloads and for reproducibility when you re-run an evaluation. - **`max_search_results`** — the ceiling on sources consulted for the call. This is your primary cost lever. - **`return_citations`** — controls whether the URLs behind the answer come back. Because none of this exists in OpenAI's schema, a typed OpenAI client cannot take it as a named argument: in Python it goes into `extra_body`, and over raw HTTP it is just another key in the JSON body. ## The response shape Two additions matter. First, a **`citations`** array of the URLs the model actually consulted — surface these in your UI, both because users trust cited answers more and because they are how you debug a wrong answer sourced from a bad page. Second, the **`usage`** object gains **`num_sources_used`**, the count of sources this request is billed for. ## The billing model This is the part candidates miss. Live Search is priced **per source consulted**, separately from input and output tokens. Two requests with identical token counts can differ several-fold in cost because one of them read a dozen pages. Practical consequences: - Log `num_sources_used` per request next to your token counts, and alert on its moving average. Token dashboards alone will silently under-report Live Search spend. - Prefer `mode: "auto"` over `"on"` unless the workload is definitionally about current events. - Set `max_search_results` deliberately rather than accepting the default; it is a hard ceiling on the per-request blast radius. - Retrieved content also enters the prompt, so searching inflates input tokens too — the cost is per-source plus the tokens those sources contribute. ## Operational caveats **Latency.** A searching request is slower by however long retrieval takes, and the variance is much wider than a pure generation call. Timeouts tuned for non-search traffic will start firing; size them per route, and stream so the user sees progress. **Non-determinism.** The live web changes. The same prompt at `temperature=0` can return different answers on different days because different pages were read. For evaluation suites, either disable search or pin `from_date`/`to_date` and accept that you are testing the pipeline, not a fixed corpus. **Trust and injection.** Retrieved pages are untrusted input. A page can contain text engineered to steer the model. Keep the model's tool permissions minimal on searching routes, never let a searched answer flow straight into a privileged action, and show citations so a human can sanity-check the provenance. **Do not confuse it with client-side tools.** Live Search is executed server-side by xAI; you neither implement nor see a tool-call round trip for it. If you also define your own function tools on the same request, those still come back as ordinary tool calls for your code to execute — the two mechanisms coexist but are not the same thing.
- Your Grok bill jumped while token usage stayed flat — what is the likely cause and how do you confirm it?Live Search sources, which are billed per source on top of tokens. Confirm it by reading `num_sources_used` from the `usage` object on each response and charting it alongside token counts; a rising source count with flat tokens is the signature. The fixes are to move `mode` from `"on"` to `"auto"`, lower `max_search_results`, and restrict `sources` so the model is not fanning out across source types it does not need.
- How does Live Search interact with an evaluation suite that expects deterministic outputs?It breaks determinism, because the corpus being read changes underneath you. Either disable search for evaluation runs so you are testing the model and prompt in isolation, or pin `from_date` and `to_date` to a fixed window and accept that you are measuring the retrieval-plus-generation pipeline rather than a stable answer. Recording the returned `citations` per run at least tells you when a regression came from a changed source rather than a changed prompt.
- What is the security posture for a route that lets Grok read arbitrary web pages?Treat every retrieved page as untrusted input capable of carrying instructions aimed at the model. Narrow `sources` to domains you are willing to trust, never let a searched answer trigger a privileged action without human confirmation, keep client-side tools on that route minimal, and display citations so provenance is visible. Bounding the freshness window and result count also limits how much attacker-controlled text can reach the context.
saying these in an interview costs you the question
- Thinking you enable search by asking for it in the prompt
- Assuming search is free because tokens look normal
- Leaving mode on 'on' for all traffic
- Treating retrieved pages as trusted content
- Expecting deterministic answers with live search active