skip to content

Which OpenAI chat-completion parameters does xAI's Grok 4 family reject?

level: seniorimportance: should knowfreq 42%

answer

  1. Compatibility is per model, not per vendor
  2. Reasoning line drops decode-time knobs
  3. Penalties and stop sequences are out
  4. It errors, it does not ignore
  5. Filter parameters by model id

basics

~20 s

xAI documents presence_penalty, frequency_penalty, stop and reasoning_effort as unsupported on the grok-4 family: sending them fails the request rather than being ignored. Compatibility is per model, so a shared parameter struct must be filtered before dispatch.

solid answer

~50 s

OpenAI compatibility at xAI is a per-model promise, not a vendor-wide one. On the grok-4 family, xAI documents `presence_penalty`, `frequency_penalty`, `stop` and `reasoning_effort` as unsupported, and a request that includes them errors out instead of silently dropping them — which is arguably the friendlier behaviour, but it breaks any client that always emits a full parameter set. Older Grok lines accept the penalties, and the smaller reasoning models are the ones that take `reasoning_effort`, so you cannot reason about "xAI" as a single capability surface. The practical response is a capability matrix keyed by model id, applied where you build the request: strip fields the target model does not accept rather than sending defaults, keep the matrix in configuration so a new model line is a config change, and make sure your error handling distinguishes a 4xx invalid-parameter failure — which retrying will never fix — from a rate-limit response, which retrying will.

go deeper

for a junior

Know that OpenAI compatibility does not guarantee every parameter is accepted, and that xAI's flagship Grok line refuses some of them outright.

for a middle

Name the excluded fields on the grok-4 family — presence_penalty, frequency_penalty, stop, reasoning_effort — and explain that the request fails rather than ignoring them.

for a senior

Show the production response: a per-model capability matrix, request builders that omit unset fields, and error handling that separates permanent parameter errors from retryable rate limits.

for a principal

Own the abstraction: how a multi-provider layer represents per-model capability, how new model lines are onboarded without a deploy, and where you accept vendor-specific fields instead of a lowest common denominator.

## The trap: compatible shape, per-model semantics "OpenAI-compatible" describes the envelope — a `/chat/completions` path, a `messages` array, a bearer key. It does not promise that every field OpenAI defines is honoured by every model behind that envelope. xAI is explicit about this: certain parameters are documented as unsupported on particular Grok model lines, and the flagship reasoning line is the one with the longest exclusion list. ## What grok-4 does not accept As documented by xAI, the grok-4 family does not support: - **`presence_penalty`** and **`frequency_penalty`** — the repetition-discouraging logit adjustments. - **`stop`** — client-supplied stop sequences. - **`reasoning_effort`** — the effort dial, which is meaningless here because the model reasons as part of its normal operation rather than on request. The important detail is the failure mode: these are **rejected**, not ignored. A request carrying them comes back as a client error. Silent ignoring would hide the mismatch and quietly change behaviour; an explicit error tells you immediately that your abstraction over-assumed. But it also means that any layer which always sends a fully populated parameter object — many gateways, ORM-ish wrappers and framework integrations do exactly this, filling unset knobs with defaults — will fail 100% of requests against that model until it is taught to omit them. ## Why the gaps exist Stop sequences and repetition penalties are decode-time interventions on the token stream. On a model that produces an internal reasoning pass before its visible answer, both are awkward: a stop sequence can cut the model off inside its own thinking, and penalties applied across a long internal trace distort it in ways the vendor cannot make safe generically. Rather than implement a subtly wrong version, xAI declines the parameter. `reasoning_effort` is excluded for the opposite reason — the model has no low-effort mode to select. ## What to do about it in production **Keep a capability matrix.** A small map from model id to the set of parameters that model accepts, resolved at request-build time. Sending is opt-in per model, not opt-out. Store it in configuration, not code, so onboarding a new model line does not need a deploy. **Never send defaults.** The root cause of the failure is usually a request builder that serialises every field of a settings object, turning "the user did not set a stop sequence" into `stop: []` on the wire. Omit unset fields entirely. **Classify errors correctly.** An invalid-parameter 4xx is permanent: retrying with backoff burns budget and latency for a result that cannot change. Your client should retry rate-limit and transient server failures, and fail fast — with the offending field named in the log — on parameter validation errors. Teams that lump all non-2xx responses into one retry path turn a five-minute fix into a mystery outage. **Replace the capability, not the parameter.** If you genuinely need the effect of `stop`, do it client-side: stream and cut the output when your delimiter appears, or ask for structured output so the boundary is syntactic rather than lexical. If you were using penalties to fight repetition, the honest answer on a reasoning model is usually a prompt change, since the model's decode strategy is not yours to tune. **Test the matrix, not the vendor.** A tiny contract test per model that fires the parameters you intend to send catches drift when xAI adds or removes support, which happens with new model generations. This matters more than usual here because the exclusions are attached to model families, and model families turn over fast. ## The general lesson This question is really about how you model multi-provider capability. The naive design assumes a lowest-common-denominator parameter set works everywhere; the mature design treats each model as a capability profile and makes the request builder ask what this specific model supports. That design also survives the inverse case — vendor-specific fields such as xAI's search parameters that no other provider understands — because both directions are just entries in the same matrix.

  • Your gateway sends a fully populated settings object to every model. What is the minimal fix?
    Stop serialising unset fields. Make the request builder emit only parameters the caller explicitly set, then intersect that with a per-model allowlist before dispatch. Defaults like an empty `stop` array are the usual culprit: the caller never asked for stop sequences, but the struct materialises the field and the model rejects it. This is a serialisation change, not a business-logic change, and it fixes the whole class of vendor mismatches at once.
  • You need stop-sequence behaviour against a model that rejects the stop parameter. How do you get it?
    Move the cut client-side. Stream the response and terminate consumption when your delimiter appears in the accumulated text, or restructure the task so the boundary is syntactic — ask for JSON and parse it, so the object's end is the stop condition. Both give you a hard boundary you control; neither depends on the server honouring a decode-time parameter, which also makes the code portable across providers that differ on this.
  • How should retry logic treat an invalid-parameter error differently from a 429?
    An invalid-parameter response is deterministic: the identical request will fail identically forever, so retrying only wastes latency and quota. Fail fast, log the model id and the rejected field, and surface it as a configuration bug. A 429 is transient and should be retried with exponential backoff and jitter, honouring whatever retry signalling the response carries. Collapsing both into one retry path is how a config error becomes an outage that looks like throttling.

saying these in an interview costs you the question

  • Assuming any OpenAI parameter works on any Grok model
  • Expecting unsupported parameters to be silently ignored
  • Retrying a 400 invalid-parameter error with backoff
  • Serialising unset defaults like empty stop arrays
  • Treating one vendor as a single capability surface

context