Why is OpenAI's Chat Completions API stateless, and how do you keep multi-turn context?
answer
- No session, no conversation id
- The array is the memory
- You append both sides
- Input tokens climb every turn
- Retries are safe because self-contained
basics
~20 sThe endpoint stores nothing between calls: there is no conversation id and no server-side memory. Your application keeps the transcript and re-sends the whole messages array every turn, appending each new user message and each returned assistant message.
solid answer
~50 sChat Completions is a pure function of the request body. Each `POST /v1/chat/completions` is independent — the server holds no session, and there is no conversation identifier to pass. Multi-turn behaviour is an illusion you create: keep the `messages` array in your own store, append the user's new turn, call the API, then append `choices[0].message` back into the array before the next call. Two consequences matter operationally. First, input tokens grow with every turn, so a long chat re-bills the whole history each time and eventually hits the model's context window. Second, because the request is self-contained, calls are trivially retryable and horizontally scalable — any worker can serve any turn. If you want the server to hold state instead, that is what OpenAI's newer Responses API surface is for; Chat Completions itself will never remember.
code
python · 14 linesimport os
from openai import OpenAI
client = OpenAI()
messages = [{"role": "system", "content": "Be concise."}]
for turn in ["My name is Ada.", "What is my name?"]:
messages.append({"role": "user", "content": turn})
reply = client.chat.completions.create(
model=os.environ["OPENAI_MODEL"],
messages=messages,
).choices[0].message
messages.append({"role": "assistant", "content": reply.content})
print(reply.content)go deeper
Be able to say the API remembers nothing and that you send the whole messages array each turn, appending both the user's message and the model's reply.
Explain the loop precisely, including appending the returned message object whole, and describe how input tokens grow with conversation length.
Discuss operating it: where the transcript is persisted, when you trim, how the context-length error surfaces, and why self-contained requests make retries and scaling easy.
Own the tradeoff between client-owned transcripts and server-held sessions — data retention, auditability, forking and replay against the engineering cost of managing history yourself.
## Stateless means exactly that When you call Chat Completions, the server reads `model`, `messages` and the sampling parameters, runs the model, returns a completion, and forgets. There is no session cookie, no thread id, no conversation identifier in the request. Send the identical body twice and you get two independent generations. Nothing you sent on the previous call influences the next one unless you send it again. ## The re-send loop The multi-turn pattern is a loop your code owns: 1. Maintain a list, usually starting with one system message. 2. Append `{"role": "user", "content": <what they typed>}`. 3. Call the API with the whole list. 4. Take `choices[0].message` from the response and append it verbatim — role `assistant`, plus any `tool_calls` it carries. 5. Repeat. Step 4 is the one people skip, and the symptom is a model that answers each question as if it had never spoken: it can see what the user said before, but not what it itself replied. ## What statelessness costs Because the full transcript is re-sent, input tokens grow roughly linearly per turn and the *cumulative* input across a conversation grows quadratically. A forty-turn chat can spend far more on re-reading its own history than on generating new text. The response's `usage.prompt_tokens` shows exactly what you paid for on that call, and watching it climb is the fastest way to see the effect. The second, harder limit is the model's context window: the entire array plus the reserved output budget must fit, and exceeding it returns a 400 with a context-length error rather than silently truncating. OpenAI's automatic prompt caching softens the cost when your request has a long, byte-identical leading prefix, reporting reused tokens in `usage.prompt_tokens_details.cached_tokens`. That is another reason to keep the system message and any long fixed preamble at the front and stable — an edit near the top invalidates the whole prefix. ## What statelessness buys The upside is real and worth saying out loud in an interview. Self-contained requests mean: any instance can serve any turn, so no sticky sessions; retries after a network error are safe to replay; you can fork a conversation by copying the array; you can edit or delete a past turn, which server-held state would not let you do; and you can replay an exact transcript for debugging or evaluation. It also means the transcript lives in *your* database, under your retention and privacy rules. ## Where the state actually lives In practice you persist the array per conversation — rows in a table, a document, or a cache key. Store the messages in the exact shape the API expects so rebuilding a request is a straight serialization, and store the token count you observed so you can decide when to trim. Never rely on the client to hold the authoritative history if the content matters; a browser can edit it. ## Common failure modes Forgetting to append the assistant turn (amnesia); appending a *summary* of the assistant turn instead of the message object, which loses tool calls; letting the array grow unbounded until requests start failing at the context limit; and mutating the earliest messages on every turn, which defeats prefix caching and quietly raises cost.
- What exactly do you append after a response — the text, or the whole message object?Append the whole `choices[0].message` object. Copying only `content` throws away `tool_calls` and any refusal field, which breaks the tool loop on the next turn because the model's request to call a tool disappears from the transcript. Serialize the object as-is into your stored array.
- How does statelessness change your retry logic compared with a stateful session API?It makes retries trivial: the request body is the complete input, so re-sending after a timeout or a server error cannot corrupt any server-side state. The only risk is duplicate generation and duplicate billing if the first call actually succeeded, so you retry with backoff and de-duplicate on your side by conversation turn.
- Why does trimming the oldest messages sometimes make your bill worse rather than better?Because it shifts the request's leading prefix. OpenAI's automatic prompt caching only reuses an exact matching prefix, so dropping or rewriting early messages invalidates the cache and subsequent requests pay full input price again. Trim from the middle, and keep the system message and any fixed preamble byte-identical at the front.
saying these in an interview costs you the question
- Thinks the API keeps a session it can resume by id
- Sends only the newest user message and wonders why context is lost
- Stores just the assistant's text and loses tool_calls
- Assumes old turns stop being billed once sent
- Believes the server truncates history automatically instead of erroring