In OpenAI Chat Completions, how does max_completion_tokens differ from max_tokens?
answer
- Output side only, never the prompt
- New name, old name deprecated
- Hidden thinking counts against it
- Watch the finish reason field
- Truncation arrives as a 200
basics
~20 smax_completion_tokens is the current parameter and caps everything the model generates, including the invisible reasoning tokens on reasoning models. max_tokens is the older, deprecated name for the output cap and is not accepted by the newer reasoning models.
solid answer
~40 sBoth cap generated output, not the prompt — neither bounds the input side. `max_tokens` was the original Chat Completions parameter and is now deprecated in favour of `max_completion_tokens`, which counts **all** tokens the model produces: visible completion text plus the hidden reasoning tokens that o-series reasoning models emit before answering. That distinction is why the rename happened, and why sending `max_tokens` to a reasoning model is rejected rather than translated. Operationally: your prompt plus this budget must fit the model's context window, so an over-large cap on a long conversation turns into a context-length 400. When the cap is hit, generation stops mid-sentence and `choices[0].finish_reason` comes back as `"length"` — check it, because the body still looks like a normal successful response, and truncated JSON is the classic downstream bug.
go deeper
Know that this parameter caps how much the model writes, not how much you send, and that the newer name is max_completion_tokens.
Explain the deprecation, that reasoning tokens count against the newer parameter, and that hitting the cap surfaces as finish_reason length on a 200 response.
Show how you size the budget against the context window and observed usage percentiles, and how your code branches on finish_reason instead of blindly parsing output.
Own it as a cost, latency and reliability control across a model fleet — per-model budgets, reasoning headroom, and the alerting that catches a rising rate of length truncations.
## One budget, two names Every Chat Completions request may bound how much the model writes. Historically the parameter was `max_tokens`. OpenAI deprecated it in favour of `max_completion_tokens`. The new name is not cosmetic: it exists because reasoning models generate tokens you never see. ## What max_completion_tokens counts `max_completion_tokens` is an upper bound on **all** tokens generated for the response — the assistant text you receive **plus** reasoning tokens on models that reason internally before answering. Those hidden tokens are billed as output and are reported separately in `usage.completion_tokens_details.reasoning_tokens`. The practical trap: set a tight budget on a reasoning model and the model can spend the entire allowance thinking, returning empty or near-empty content with `finish_reason` `"length"`. You paid, and you got nothing readable. On reasoning models the budget must be generous enough for reasoning *and* the answer. ## Why max_tokens was deprecated Under the old name it was ambiguous whether the cap covered invisible work. Rather than redefine it, OpenAI introduced the clearer name and marked the old one deprecated in Chat Completions; the o-series reasoning models do not accept `max_tokens` at all and reject the request. Non-reasoning models still honour `max_tokens` for compatibility, but new code should send `max_completion_tokens`. ## Neither one limits the prompt A persistent misconception is that this parameter protects you from a giant conversation. It does not. It bounds output only. The relationship you must respect is arithmetic: prompt tokens plus generated tokens must fit the model's context window. If your transcript is already near the ceiling, reserving a large output budget can make the request unsatisfiable and you get a 400 with a context-length error before a single token is generated. That is why history trimming and output budgeting are the same design problem. ## finish_reason is the field that matters When generation stops because the cap was reached, the HTTP status is still 200 and the body still has content — just truncated. `choices[0].finish_reason` is `"length"` instead of `"stop"`. Other values you will see are `"tool_calls"` when the model wants a tool run and `"content_filter"` when output was filtered. Production code should branch on this field explicitly: `"stop"` means the model finished on its own; `"length"` means you cut it off and the payload is probably invalid if you were expecting structured text. Retrying the same request unchanged usually reproduces the truncation — you need a bigger budget, a shorter prompt, or a continuation strategy. ## Sizing the budget in practice Three rules cover most services. First, set a cap deliberately rather than leaving it unset — an unbounded response can run away in latency and cost. Second, size it from what a good answer actually needs, plus headroom for reasoning if the model reasons, and validate it against the remaining context space. Third, log the distribution of `usage.completion_tokens` and set the cap near a high percentile rather than at the mean, so ordinary answers are never clipped. Because the cap also bounds worst-case generation time for a streamed response, it doubles as a latency lever. ## Interview-grade summary The strong answer names the deprecation, explains that the new parameter includes hidden reasoning tokens, states clearly that neither bounds the prompt, and then moves straight to `finish_reason` `"length"` as the thing you must actually handle in code. Candidates who stop at "it limits the response length" sound like they have only read the quickstart.
- You set a small budget on a reasoning model and get back empty content. What happened?The model spent the whole allowance on reasoning tokens before producing any visible answer, so the response ends with `finish_reason` `"length"` and little or no content, while `usage.completion_tokens_details.reasoning_tokens` shows where the budget went. Those tokens are billed as output. The fix is a materially larger `max_completion_tokens` on reasoning models, not a blind retry.
- How do you detect and handle a truncated response in code?Branch on `choices[0].finish_reason`. `"stop"` means the model completed naturally; `"length"` means it was cut off and any JSON or structured payload should be treated as invalid rather than parsed. Handle it by raising the budget, trimming the prompt, or asking for a continuation — never by silently passing the fragment downstream, which is how malformed data leaks into a pipeline.
- Why can raising max_completion_tokens cause a request to fail before generation starts?Because the prompt and the reserved output budget must both fit the model's context window. On a long conversation, a large reserved budget can push the total past the limit and the API returns a 400 context-length error immediately. Output budgeting and history trimming are one problem: as the transcript grows, the space you can reserve for the answer shrinks.
saying these in an interview costs you the question
- Says the parameter limits the size of the prompt
- Thinks hitting the cap returns a 4xx error
- Assumes max_tokens works on every current OpenAI model
- Ignores that reasoning tokens consume the budget and are billed
- Parses truncated JSON without checking finish_reason