skip to content

xAI

xAI serves its Grok models behind an OpenAI-compatible API, so it is another drop-in provider. Its distinguishing feature is live web and X search offered as a first-class request parameter.

on this pageshow

questions

5

How do you call xAI's Grok models using the OpenAI SDK instead of OpenAI's?

level: juniorimportance: must knowfreq 70%

answer

  1. Same SDK, different endpoint
  2. Three knobs: URL, key, model
  3. api.x.ai, bearer key, grok- ids
  4. Non-OpenAI fields ride in extra_body
  5. Compatible in shape, not in limits

basics

~20 s

xAI exposes an OpenAI-compatible chat completions API. Keep the OpenAI SDK, point base_url at https://api.x.ai/v1, authenticate with your xAI key (conventionally XAI_API_KEY) as the bearer token, and pass a Grok model id such as grok-4.

solid answer

~40 s

xAI's API is deliberately shaped like OpenAI's, so you do not swap client libraries — you change three things. Set `base_url="https://api.x.ai/v1"`, pass your xAI key instead of an OpenAI key (people usually read it from `XAI_API_KEY`; it travels as an `Authorization: Bearer` header), and use a Grok model id such as `grok-4`. Everything else keeps the familiar shape: a `messages` array with system/user/assistant roles, `temperature`, `top_p`, `max_tokens`, `stream=True` for incremental deltas, and `tools`/`tool_choice`/`tool_calls` for function calling. The important caveat is that compatible does not mean identical: xAI-only request fields such as `search_parameters` have to be smuggled through `extra_body`, some OpenAI parameters are rejected by certain Grok models, and the rate limits, error bodies and model lifecycle are xAI's own.

code

python · 16 lines
python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["XAI_API_KEY"],
    base_url="https://api.x.ai/v1",
)

response = client.chat.completions.create(
    model="grok-4",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Summarise the CAP theorem in two sentences."},
    ],
)
print(response.choices[0].message.content)

go deeper

for a junior

Be able to name the three changes — base URL https://api.x.ai/v1, an xAI key as the bearer token, and a grok- model id — and say that the messages array and streaming work exactly as with OpenAI.

for a middle

Explain what actually travels on the wire: the bearer header, the chat completions body, and why non-OpenAI fields need a passthrough like extra_body rather than a named SDK argument.

for a senior

Show that you treat compatibility as per-model capability, not a vendor-wide promise: verify supported parameters, handle xAI's own error and limit behaviour, and keep provider config in one place.

for a principal

Own the abstraction decision — whether a single OpenAI-shaped client across vendors is worth the untyped edges, and what a capability matrix plus config-driven model ids buys you when a model line is retired.

## What xAI serves xAI hosts the Grok model family behind a REST API rooted at `https://api.x.ai/v1`. Its primary text surface is a chat completions endpoint whose request and response JSON is intentionally modelled on OpenAI's, which is why the whole ecosystem of OpenAI clients, gateways and framework adapters can talk to Grok with a configuration change rather than a code change. ## The three things you change **1. Base URL.** Every OpenAI-shaped client exposes a base URL knob (`base_url` in the Python and Node SDKs, `OPENAI_BASE_URL` as an environment variable, a `baseURL` field in most frameworks). Set it to `https://api.x.ai/v1` and the SDK will POST to xAI's `/chat/completions` instead of OpenAI's. **2. Credentials.** xAI issues its own API keys from its console. The key travels in the standard `Authorization: Bearer <key>` header, so the SDK's `api_key` argument works unchanged; by convention it is read from `XAI_API_KEY` rather than `OPENAI_API_KEY`. Keep the two variables distinct — a service that reads `OPENAI_API_KEY` while pointing at xAI is the classic misconfiguration, and it fails with an authentication error that people waste time blaming on the base URL. **3. Model id.** Model names are xAI's: the `grok-` prefixed ids for the current generation (for example the `grok-4` family, faster variants for cheap high-throughput work, and code-specialised variants). Never hardcode a model id in call sites — put it in configuration, because vendors retire and rename model ids far faster than they change API shapes. ## What carries over unchanged Because the wire format matches, the following behave as you expect: the `messages` array with `system`, `user`, `assistant` and `tool` roles; sampling knobs like `temperature` and `top_p`; `max_tokens`; `stream=True`, which yields incremental delta chunks over Server-Sent Events that you accumulate exactly as with OpenAI; function calling via a `tools` array of JSON-Schema function definitions, a `tool_choice` selector, and `tool_calls` on the assistant message answered by follow-up messages carrying `tool_call_id`; structured output via `response_format`; and a `usage` object reporting prompt and completion token counts. ## Where compatibility stops This is the part interviews actually probe. - **xAI-only request fields.** The differentiating feature — live web and X search — is requested with a `search_parameters` object that is not part of OpenAI's schema. The typed SDKs will not accept it as a named argument, so it goes in `extra_body` (Python) or the equivalent passthrough, and over raw HTTP it is simply another JSON key. - **xAI-only response fields.** A searched answer comes back with extra top-level data such as a `citations` array and a source count in `usage`. The typed OpenAI models tolerate unknown fields but will not autocomplete them, so read them defensively. - **Per-model parameter gaps.** Compatibility is per model, not per vendor: some OpenAI parameters are documented as unsupported on particular Grok models and cause the request to fail rather than being ignored. Treat "which parameters does this model accept" as a capability matrix you verify, not an assumption. - **Operations are xAI's.** Rate limits, quota tiers, error codes and message bodies, billing and model deprecation schedules all come from xAI. Retry and backoff logic written against OpenAI's limits is a starting point, not a spec. - **A second compatibility mode exists.** xAI also documents an Anthropic-SDK-compatible way in, which is useful if your codebase already standardised on that client — but the same rule applies: the shape is borrowed, the semantics and limits are xAI's. ## Practical setup Keep provider configuration as a triple — base URL, key, model id — resolved at startup, and log the resolved base URL (never the key) so a misrouted deployment is obvious. If you front several vendors, put the triple behind one small factory rather than sprinkling `base_url` overrides through the code, and keep a per-model list of parameters you are allowed to send. That single indirection is what makes "try Grok on this workload" a config flag instead of a project.

  • You pointed the OpenAI SDK at xAI and every call returns an authentication error — what do you check first?
    Which key the client actually resolved. The commonest cause is that the SDK fell back to `OPENAI_API_KEY` from the environment while `base_url` pointed at xAI, so an OpenAI key is being presented to xAI. Pass the key explicitly from `XAI_API_KEY` rather than relying on implicit environment pickup, and log the resolved base URL and key prefix (never the whole key) at startup.
  • How do you send xAI's Live Search fields when the typed OpenAI client rejects unknown arguments?
    Use the client's passthrough: `extra_body={"search_parameters": {...}}` in the Python SDK, the equivalent body-merge option in other clients, or just build the JSON yourself over HTTP. The field is sent verbatim to xAI. Because it is untyped, validate the object in your own code — a typo in a passthrough field is silently accepted by the SDK and only surfaces as a server-side error or, worse, as a feature that never activated.
  • Does OpenAI compatibility mean you can reuse the same retry and rate-limit code across both providers?
    Only the mechanism, not the numbers. The status codes and the general backoff strategy transfer, but request-per-minute and token-per-minute ceilings, tier progression and the exact error payloads are xAI's own and differ per model. Make limits and retry budgets per-provider configuration, and drive backoff off the response's retry signalling rather than constants tuned against a different vendor.

saying these in an interview costs you the question

  • Assuming an OpenAI API key authenticates against xAI
  • Believing every OpenAI parameter works on every Grok model
  • Passing search_parameters as a normal SDK argument
  • Hardcoding Grok model ids throughout call sites
  • Thinking compatibility means shared rate limits or pricing

context

open as a page

How does xAI's API expose reasoning effort and reasoning traces on Grok?

level: middleimportance: should knowfreq 38%

basics

~20 s

On xAI's lighter Grok reasoning models you set reasoning_effort to "low" or "high" to trade depth against latency and cost. The trace comes back in the message as reasoning_content, and thinking tokens are billed as completion tokens with a reasoning_tokens breakdown in usage.

open as a page

Which OpenAI chat-completion parameters does xAI's Grok 4 family reject?

level: seniorimportance: should knowfreq 42%

basics

~20 s

xAI documents presence_penalty, frequency_penalty, stop and reasoning_effort as unsupported on the grok-4 family: sending them fails the request rather than being ignored. Compatibility is per model, so a shared parameter struct must be filtered before dispatch.

open as a page

For a production Grok service, when is xAI's native SDK worth it over OpenAI-compatible calls?

level: principalimportance: should knowfreq 32%

basics

~20 s

Use the OpenAI-compatible endpoint when portability across vendors is the priority and you mostly need plain chat. Reach for xai-sdk when xAI-specific capabilities — live search, deferred sampling, typed helpers — are central enough that untyped passthrough fields become a liability.

open as a page