skip to content

Chat Completions API

The endpoint everything else is built on: you send a list of role-tagged messages and get a completion back. Interviewers check that you understand the API is stateless — the whole conversation is resent each turn — and what temperature, top_p and max_tokens actually do to sampling.

on this pageshow

questions

6

In OpenAI Chat Completions, what do the system, user, and assistant roles do?

level: juniorimportance: must knowfreq 85%

answer

  1. Four roles, one array
  2. Who said what, nothing else
  3. Standing instructions live at the top
  4. You append the model's own reply
  5. tool closes the tool-call loop

basics

~20 s

Every entry in the messages array carries a role. system holds standing instructions the model should follow throughout, user holds what the human said, and assistant holds the model's own earlier replies that you replay as context.

solid answer

~50 s

A Chat Completions request is a `model` plus a `messages` array, and each message is `{"role": ..., "content": ...}`. The **system** message carries standing instructions — persona, output format, constraints — and is conventionally the first element; newer OpenAI models also accept `developer` for the same purpose. **user** messages are the human turns. **assistant** messages are the model's own prior answers, which *you* append back into the array so the next call can see them. There is also a **tool** role used to hand back the result of a tool the model asked for. Roles are not required to alternate strictly: a request may contain a single user message, or several user messages in a row. The role is the only signal telling the model who said what, and instruction-following is trained to weight system text above user text.

code

python · 12 lines
python
import os
from openai import OpenAI

client = OpenAI()
response = client.chat.completions.create(
    model=os.environ["OPENAI_MODEL"],
    messages=[
        {"role": "system", "content": "Answer in one short sentence."},
        {"role": "user", "content": "What is a message role?"},
    ],
)
print(response.choices[0].message.content)

go deeper

for a junior

Be able to name the roles and say what each holds, and state plainly that you append the model's own reply as an assistant message before the next call.

for a middle

Explain that the array is the entire conversational state, that content may be a string or an array of parts, and that consecutive same-role messages are legal.

for a senior

Show judgment about what belongs in the system message versus per-turn text, including its token cost on every request and its weakness as a trust boundary.

for a principal

Own the instruction hierarchy as a design concern: which policies live in the system layer, how user text can subvert them, and what must be enforced in code rather than in a prompt.

## The shape of a request A call to OpenAI's Chat Completions endpoint (`POST /v1/chat/completions`, or `client.chat.completions.create(...)` in the SDK) has two required parts: `model`, naming which model runs, and `messages`, an ordered array describing the conversation so far. Every element has at least `role` and `content`. The `role` string is the only thing telling the model who produced a given piece of text — there is no separate speaker field. ## system — standing instructions The `system` message is where you put things that should hold for the whole conversation: who the assistant is, what tone to use, what format to answer in, what it must refuse. It is placed first by convention. Newer OpenAI models also accept the `developer` role, introduced as the successor name for the same job — application-owner instructions ranked above end-user text. Two practical facts follow. First, the system message is re-sent on every call (see statelessness), so it is re-billed as input tokens every turn; keep it tight. Second, system instructions are strong but not a security boundary — a determined user message can still argue with them, so never rely on a system prompt alone to protect secrets or enforce authorization. ## user — the human turn `user` messages are what the person typed. `content` may be a plain string, or an array of content parts (`{"type": "text", ...}` plus image parts for vision-capable models) when you need multimodal input. Multiple consecutive user messages are legal; the API does not enforce alternation. ## assistant — the model's own replies This is the role people get wrong. The API does not remember what it said last turn. When a response comes back, you read `choices[0].message`, and if you want the next call to know about that answer, you append that message object — role `assistant` and its content — to your own `messages` array. An assistant message may also carry `tool_calls` instead of, or alongside, content when the model wants a tool run. You can even write an assistant message the model never produced, to seed the style of an answer or supply a few-shot example; the API accepts it. ## tool — results coming back in When the model emits tool calls, the loop is closed by appending messages with role `tool` whose content is the tool's output. That flow belongs to tool calling proper; the point here is simply that `tool` is the fourth role you will see in a `messages` array. ## Ordering and validity There is no hard rule that roles alternate. Common valid shapes: system plus one user message (a single-shot call); system plus user/assistant/user/... (a chat); user only, with the system message omitted entirely so the model uses its default behaviour. A request with an empty `messages` array is rejected with a 400. ## Common mistakes Putting per-turn instructions in the system message when they belong in the user turn, or the reverse — burying standing policy in one early user message where later text outweighs it. Forgetting to append the assistant reply, so the model loses its own last answer. Assuming the system message is protected from the user. And assuming a role mismatch will error: it usually will not, it will just produce a confused conversation that is hard to debug because nothing failed.

  • If you omit the system message entirely, what happens?
    Nothing breaks — the request is valid and the model answers with its default trained behaviour. You simply lose the lever that sets persona, format and constraints, so output style becomes less predictable across turns. Most production apps always send one, both for control and because a stable leading prefix is what OpenAI's automatic prompt caching keys on.
  • Can you put words in the assistant's mouth by adding an assistant message it never produced?
    Yes. The API takes the messages array at face value, so you can inject synthetic assistant turns as few-shot demonstrations or to steer format. It is a legitimate technique, but those tokens are billed as input on every subsequent call, and the model may treat a fabricated turn as established fact.
  • What does the developer role mean relative to system?
    It is the newer name for application-owner instructions on OpenAI's more recent models, sitting above end-user content in the instruction hierarchy. Functionally you use it the way you used system; existing system messages continue to work, so most codebases keep sending system unless they are targeting the newer role deliberately.

saying these in an interview costs you the question

  • Thinks the API remembers the assistant's previous reply automatically
  • Claims roles must strictly alternate user, assistant, user
  • Treats the system prompt as a security boundary users cannot override
  • Puts per-turn user input into the system message
  • Believes content must always be a plain string

context

open as a page

Why is OpenAI's Chat Completions API stateless, and how do you keep multi-turn context?

level: middleimportance: must knowfreq 78%

basics

~20 s

The endpoint stores nothing between calls: there is no conversation id and no server-side memory. Your application keeps the transcript and re-sends the whole messages array every turn, appending each new user message and each returned assistant message.

open as a page

In OpenAI Chat Completions, what happens if you set both temperature and top_p?

level: middleimportance: must knowfreq 68%

basics

~20 s

Nothing errors — both are applied to the same sampling step, and OpenAI's own guidance is to change one or the other, not both, because their combined effect is hard to predict. Defaults are temperature 1 and top_p 1.

open as a page

How do you consume a streamed OpenAI Chat Completions response chunk by chunk?

level: middleimportance: should knowfreq 62%

basics

~20 s

Set stream to true and iterate the Server-Sent Events response. Each event is a chat.completion.chunk whose choices[0].delta holds the newest fragment; concatenate those fragments yourself, and stop when the stream sends its final done marker.

open as a page

In OpenAI Chat Completions, how does max_completion_tokens differ from max_tokens?

level: seniorimportance: should knowfreq 52%

basics

~20 s

max_completion_tokens is the current parameter and caps everything the model generates, including the invisible reasoning tokens on reasoning models. max_tokens is the older, deprecated name for the output cap and is not accepted by the newer reasoning models.

open as a page

As an OpenAI chat outgrows the context window, how do you decide which history to resend?

level: principalimportance: should knowfreq 40%

basics

~20 s

There is no default — the API errors rather than trimming for you, so the policy is yours. Choose between a sliding window, rolling summarization, and retrieving only relevant past turns, budgeting tokens against the window and leaving room for the answer.

open as a page