skip to content

OpenAI

OpenAI is the reference API of the whole space — the request shape most other vendors ended up imitating. Interviews walk the surface: chat completions, tool calling, structured outputs, embeddings, and the operational side of rate-limit tiers and cost.

on this pageshow

explore

questions

page 1 of 2

In OpenAI Chat Completions, what do the system, user, and assistant roles do?

level: juniorimportance: must knowfreq 85%

answer

  1. Four roles, one array
  2. Who said what, nothing else
  3. Standing instructions live at the top
  4. You append the model's own reply
  5. tool closes the tool-call loop

basics

~20 s

Every entry in the messages array carries a role. system holds standing instructions the model should follow throughout, user holds what the human said, and assistant holds the model's own earlier replies that you replay as context.

solid answer

~50 s

A Chat Completions request is a `model` plus a `messages` array, and each message is `{"role": ..., "content": ...}`. The **system** message carries standing instructions — persona, output format, constraints — and is conventionally the first element; newer OpenAI models also accept `developer` for the same purpose. **user** messages are the human turns. **assistant** messages are the model's own prior answers, which *you* append back into the array so the next call can see them. There is also a **tool** role used to hand back the result of a tool the model asked for. Roles are not required to alternate strictly: a request may contain a single user message, or several user messages in a row. The role is the only signal telling the model who said what, and instruction-following is trained to weight system text above user text.

code

python · 12 lines
python
import os
from openai import OpenAI

client = OpenAI()
response = client.chat.completions.create(
    model=os.environ["OPENAI_MODEL"],
    messages=[
        {"role": "system", "content": "Answer in one short sentence."},
        {"role": "user", "content": "What is a message role?"},
    ],
)
print(response.choices[0].message.content)

go deeper

for a junior

Be able to name the roles and say what each holds, and state plainly that you append the model's own reply as an assistant message before the next call.

for a middle

Explain that the array is the entire conversational state, that content may be a string or an array of parts, and that consecutive same-role messages are legal.

for a senior

Show judgment about what belongs in the system message versus per-turn text, including its token cost on every request and its weakness as a trust boundary.

for a principal

Own the instruction hierarchy as a design concern: which policies live in the system layer, how user text can subvert them, and what must be enforced in code rather than in a prompt.

## The shape of a request A call to OpenAI's Chat Completions endpoint (`POST /v1/chat/completions`, or `client.chat.completions.create(...)` in the SDK) has two required parts: `model`, naming which model runs, and `messages`, an ordered array describing the conversation so far. Every element has at least `role` and `content`. The `role` string is the only thing telling the model who produced a given piece of text — there is no separate speaker field. ## system — standing instructions The `system` message is where you put things that should hold for the whole conversation: who the assistant is, what tone to use, what format to answer in, what it must refuse. It is placed first by convention. Newer OpenAI models also accept the `developer` role, introduced as the successor name for the same job — application-owner instructions ranked above end-user text. Two practical facts follow. First, the system message is re-sent on every call (see statelessness), so it is re-billed as input tokens every turn; keep it tight. Second, system instructions are strong but not a security boundary — a determined user message can still argue with them, so never rely on a system prompt alone to protect secrets or enforce authorization. ## user — the human turn `user` messages are what the person typed. `content` may be a plain string, or an array of content parts (`{"type": "text", ...}` plus image parts for vision-capable models) when you need multimodal input. Multiple consecutive user messages are legal; the API does not enforce alternation. ## assistant — the model's own replies This is the role people get wrong. The API does not remember what it said last turn. When a response comes back, you read `choices[0].message`, and if you want the next call to know about that answer, you append that message object — role `assistant` and its content — to your own `messages` array. An assistant message may also carry `tool_calls` instead of, or alongside, content when the model wants a tool run. You can even write an assistant message the model never produced, to seed the style of an answer or supply a few-shot example; the API accepts it. ## tool — results coming back in When the model emits tool calls, the loop is closed by appending messages with role `tool` whose content is the tool's output. That flow belongs to tool calling proper; the point here is simply that `tool` is the fourth role you will see in a `messages` array. ## Ordering and validity There is no hard rule that roles alternate. Common valid shapes: system plus one user message (a single-shot call); system plus user/assistant/user/... (a chat); user only, with the system message omitted entirely so the model uses its default behaviour. A request with an empty `messages` array is rejected with a 400. ## Common mistakes Putting per-turn instructions in the system message when they belong in the user turn, or the reverse — burying standing policy in one early user message where later text outweighs it. Forgetting to append the assistant reply, so the model loses its own last answer. Assuming the system message is protected from the user. And assuming a role mismatch will error: it usually will not, it will just produce a confused conversation that is hard to debug because nothing failed.

  • If you omit the system message entirely, what happens?
    Nothing breaks — the request is valid and the model answers with its default trained behaviour. You simply lose the lever that sets persona, format and constraints, so output style becomes less predictable across turns. Most production apps always send one, both for control and because a stable leading prefix is what OpenAI's automatic prompt caching keys on.
  • Can you put words in the assistant's mouth by adding an assistant message it never produced?
    Yes. The API takes the messages array at face value, so you can inject synthetic assistant turns as few-shot demonstrations or to steer format. It is a legitimate technique, but those tokens are billed as input on every subsequent call, and the model may treat a fabricated turn as established fact.
  • What does the developer role mean relative to system?
    It is the newer name for application-owner instructions on OpenAI's more recent models, sitting above end-user content in the instruction hierarchy. Functionally you use it the way you used system; existing system messages continue to work, so most codebases keep sending system unless they are targeting the newer role deliberately.

saying these in an interview costs you the question

  • Thinks the API remembers the assistant's previous reply automatically
  • Claims roles must strictly alternate user, assistant, user
  • Treats the system prompt as a security boundary users cannot override
  • Puts per-turn user input into the system message
  • Believes content must always be a plain string

context

open as a page

What does OpenAI's /v1/embeddings endpoint take as input and return?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A POST to /v1/embeddings sends a model plus input (a string or an array of strings) and returns a data array of vectors, each carrying its embedding and the index of the input it came from, plus a usage object counting input tokens.

open as a page

How is an OpenAI API call billed, and why do long chats cost more?

level: juniorimportance: must knowfreq 72%

basics

~20 s

OpenAI bills per token, quoted per million tokens, and output tokens cost several times more than input tokens. The API is stateless, so every turn resends the whole conversation as input — the input bill grows with each turn.

open as a page

In OpenAI's Responses API, what does previous_response_id actually do?

level: middleimportance: must knowfreq 66%

basics

~20 s

previous_response_id points at a stored earlier response, so the server replays that whole conversation as the prefix of the new request. You send only the new turn, but you are still billed for every replayed input token.

open as a page

How do you make documents searchable by OpenAI's hosted file_search tool?

level: middleimportance: must knowfreq 56%

basics

~20 s

Upload the files, add them to a vector store, and wait for indexing to finish, then reference that store's id from the file_search tool on your request. OpenAI handles chunking, embedding, and ranking; you pay for storage per day.

open as a page

Why is OpenAI's Chat Completions API stateless, and how do you keep multi-turn context?

level: middleimportance: must knowfreq 78%

basics

~20 s

The endpoint stores nothing between calls: there is no conversation id and no server-side memory. Your application keeps the transcript and re-sends the whole messages array every turn, appending each new user message and each returned assistant message.

open as a page

In OpenAI Chat Completions, what happens if you set both temperature and top_p?

level: middleimportance: must knowfreq 68%

basics

~20 s

Nothing errors — both are applied to the same sampling step, and OpenAI's own guidance is to change one or the other, not both, because their combined effect is hard to predict. Defaults are temperature 1 and top_p 1.

open as a page

What does the dimensions parameter do in OpenAI's Embeddings API?

level: middleimportance: must knowfreq 56%

basics

~20 s

Passing dimensions asks a text-embedding-3-small or text-embedding-3-large call to return a shorter vector — for example 256 instead of 3072. The models are trained so the leading coordinates carry the most information, so quality degrades gradually rather than collapsing.

open as a page

When would you pick text-embedding-3-large over text-embedding-3-small?

level: middleimportance: must knowfreq 68%

basics

~20 s

Pick text-embedding-3-large when retrieval quality is the bottleneck and the corpus is small enough that its higher per-token price and 3072-dimensional vectors are affordable. text-embedding-3-small is the sane default: far cheaper, 1536 dimensions, and close on benchmark quality.

open as a page

How do you handle an OpenAI API 429 rate_limit_exceeded error?

level: middleimportance: must knowfreq 80%

basics

~20 s

Retry it with exponential backoff plus random jitter and a capped attempt count, using the x-ratelimit-reset headers to time the wait. First check the error code: rate_limit_exceeded is transient and worth retrying, while insufficient_quota is a billing problem no retry will clear.

open as a page

Why can an OpenAI call return 429 on TPM while RPM is barely used?

level: middleimportance: must knowfreq 63%

basics

~20 s

OpenAI enforces several limits at once — requests per minute, tokens per minute, and daily caps — per model and per project. Exceeding any single one returns 429, so a handful of very large prompts can exhaust the token budget while the request count stays trivial.

open as a page

In the OpenAI API, how does response_format json_object differ from json_schema strict mode?

level: middleimportance: must knowfreq 72%

basics

~20 s

json_object only promises the reply parses as JSON — which keys appear and what types they hold is left to the prompt. json_schema with strict true constrains decoding to a schema you supply, so the returned object provably matches that shape.

open as a page

In OpenAI strict json_schema mode, what schema rules apply and how do you mark a field optional?

level: middleimportance: must knowfreq 58%

basics

~20 s

Strict mode accepts a restricted JSON Schema subset: every object must set additionalProperties to false and list every one of its properties in required. Optionality is expressed by making the type a union with null, never by leaving the key out of required.

open as a page

How do you define a tool in OpenAI's Chat Completions API and return its result?

level: middleimportance: must knowfreq 78%

basics

~20 s

You send a tools array whose entries describe a function with a JSON Schema. The model replies with message.tool_calls instead of prose; your code runs each call, then appends a role 'tool' message carrying the matching tool_call_id.

open as a page

What do OpenAI's tool_choice values auto, required and none do?

level: middleimportance: must knowfreq 62%

basics

~20 s

tool_choice auto lets the model decide and is the default when tools are present; required forces at least one tool call this turn; none forbids calling while still sending the definitions; and a named function object pins the call to that one function.

open as a page

In the OpenAI Python SDK, what does the chat completions parse() helper give you over create()?

level: juniorimportance: should knowfreq 62%

basics

~20 s

parse() accepts a Pydantic model as response_format, converts it into a strict JSON Schema for you, and hands back the reply already deserialised into an instance of that model on message.parsed — plus the separate refusal field to check first.

open as a page

In OpenAI's Assistants API, what states does a run pass through before finishing?

level: middleimportance: should knowfreq 48%

basics

~10 s

A run is created as queued, moves to in_progress, and normally ends completed. It can pause at requires_action while it waits for tool outputs, and can instead end as failed, cancelled, incomplete, or expired.

open as a page

How do you consume a streamed OpenAI Chat Completions response chunk by chunk?

level: middleimportance: should knowfreq 62%

basics

~20 s

Set stream to true and iterate the Server-Sent Events response. Each event is a chat.completion.chunk whose choices[0].delta holds the newest fragment; concatenate those fragments yourself, and stop when the stream sends its final done marker.

open as a page

Why does OpenAI recommend cosine similarity for comparing its embeddings?

level: middleimportance: should knowfreq 50%

basics

~20 s

OpenAI embeddings are returned normalised to unit length, so cosine similarity reduces to a plain dot product and gives exactly the same ranking as Euclidean distance. Cosine is recommended because it is the cheapest and least error-prone of the equivalent choices.

open as a page

How do you count an OpenAI request's tokens with tiktoken before sending?

level: middleimportance: should knowfreq 52%

basics

~20 s

Get the encoding for your model with tiktoken.encoding_for_model, then take len(encoding.encode(text)) for each message. Add the per-message framing overhead, because raw text encoding alone undercounts a chat request and ignores tool schemas and images entirely.

open as a page

How do parallel tool calls work in OpenAI's API, and when do you disable them?

level: middleimportance: should knowfreq 50%

basics

~20 s

One assistant message can carry several entries in tool_calls; run them concurrently, then append one tool message per tool_call_id before the next request. Set parallel_tool_calls to false when calls must be ordered or have side effects.

open as a page

When would you stream an OpenAI Assistants run instead of polling it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Stream when a human is waiting on the text, because polling only reveals the answer after the run finishes. Streaming does not remove the tool-output round trip, the one-active-run-per-thread rule, or the need to re-read the run after a dropped connection.

open as a page

In OpenAI Chat Completions, how does max_completion_tokens differ from max_tokens?

level: seniorimportance: should knowfreq 52%

basics

~20 s

max_completion_tokens is the current parameter and caps everything the model generates, including the invisible reasoning tokens on reasoning models. max_tokens is the older, deprecated name for the output cap and is not accepted by the newer reasoning models.

open as a page

How would you embed a large document corpus through OpenAI's Embeddings API?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Chunk documents under the model's 8192-token input limit, pack many chunks per request as an array (up to 2048 entries), reassemble results by the index field, and for very large corpora submit the work through the asynchronous Batch API instead of live calls.

open as a page

Your live feature keeps hitting OpenAI 429s at peak — how do you fix it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Stop relying on retries and add admission control: a shared token-bucket limiter sized to the tightest limit, so work waits in your queue instead of bouncing off the API. Then reduce demand — realistic completion caps, smaller models for cheap traffic, bulk work moved to the Batch API.

open as a page

With OpenAI strict Structured Outputs, when can a response still fail to parse against your schema?

level: seniorimportance: should knowfreq 47%

basics

~20 s

The guarantee covers completed generations only. A refusal returns a refusal string instead of content, a generation cut off by the token limit returns a valid prefix that is not valid JSON, and content filtering or transport errors return no usable body at all.

open as a page

How do you keep an OpenAI tool-calling agent loop from running forever?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Loop until a response has no tool_calls and finish_reason is stop, and bound it: a maximum step count, a token or cost budget, per-tool timeouts, and duplicate-call detection. Force a prose answer with tool_choice none on the final step.

open as a page

How do you reconstruct tool_calls from an OpenAI streaming chat completion?

level: seniorimportance: should knowfreq 40%

basics

~20 s

With stream enabled, tool calls arrive as fragments in delta.tool_calls. Each fragment has an index; the first for that index carries the id and function name, later ones carry pieces of the arguments string you concatenate before parsing.

open as a page

When is OpenAI's hosted conversation state the wrong choice for production?

level: principalimportance: should knowfreq 42%

basics

~20 s

Hosted state is wrong when storage is contractually forbidden, when transcripts must be your own for audit, evaluation or portability, or when you need explicit control over what the model sees each turn. It saves client complexity, never token cost.

open as a page

As an OpenAI chat outgrows the context window, how do you decide which history to resend?

level: principalimportance: should knowfreq 40%

basics

~20 s

There is no default — the API errors rather than trimming for you, so the policy is yours. Choose between a sliding window, rolling summarization, and retrieving only relevant past turns, budgeting tokens against the window and leaving room for the answer.

open as a page

showing 1–30 of 35