skip to content

In OpenAI's Responses API, what does previous_response_id actually do?

level: middleimportance: must knowfreq 66%

answer

  1. server holds the transcript, not the bill
  2. one id points at the whole prefix
  3. transport savings, not token savings
  4. instructions do not ride along
  5. no id to chain when nothing is stored

basics

~20 s

previous_response_id points at a stored earlier response, so the server replays that whole conversation as the prefix of the new request. You send only the new turn, but you are still billed for every replayed input token.

solid answer

~50 s

The Responses API is the stateful side of OpenAI's platform. When `store` is true (the default), the server keeps the response and its items under an id like `resp_...`. Passing `previous_response_id` on the next call tells the server to prepend that response's conversation items to your new `input`, so the client only ships the newest turn instead of re-serialising the whole transcript. The important nuance is that this is a *transport* convenience, not a billing one: the model still sees the full conversation, so those replayed tokens are charged as input tokens on every turn (automatic prompt caching can discount the repeated prefix, but it does not remove it). Two other details bite people: `instructions` are not inherited across a chain — resend them each call — and with `store` false there is nothing to chain to, so you must send the history yourself.

code

python · 19 lines
python
from openai import OpenAI

client = OpenAI()

first = client.responses.create(
    model="gpt-5",
    instructions="You are terse.",
    input="Name one benefit of connection pooling.",
)
print(first.id, first.output_text)

second = client.responses.create(
    model="gpt-5",
    instructions="You are terse.",  # not inherited from the first call
    previous_response_id=first.id,
    input="And one drawback?",
)
print(second.output_text)
print(second.usage.input_tokens)  # includes the replayed first turn

go deeper

for a junior

Know that the Responses API can hold the conversation for you: the response comes back with an id, and passing that id as previous_response_id on the next call continues the chat without resending everything.

for a middle

Explain the mechanics: store defaults to true, the id addresses a stored item list, and the server prepends it to your new input. Say clearly that the replayed tokens are still billed as input tokens.

for a senior

Show the operational judgment — instructions and tools are not inherited, truncation decides what happens at the context limit, and zero-data-retention orgs must run store false and resend history themselves.

for a principal

Own the tradeoff: hosted state is a client-simplicity optimisation, not a cost or durability strategy. Argue for keeping your own transcript as the system of record so evals, audits, and provider migration stay possible.

## The problem this solves A stateless chat API makes the client the owner of conversation history: you keep an array of messages, append each new turn, and resend everything. That works, but it means every client that talks to the model needs its own transcript store, and mobile or edge clients have to ship a growing payload on every request. OpenAI's Responses API offers the other arrangement: the platform keeps the conversation for you. You POST to `/v1/responses` with `model` and `input`; you get back a response object with an id of the form `resp_...`, a `status` (`completed`, `incomplete`, `failed`, `in_progress`), an `output` array of items, and a `usage` object with `input_tokens` / `output_tokens` / `total_tokens`. ## store and the response id The request field `store` defaults to `true`. A stored response is retrievable later with `GET /v1/responses/{id}`, its inputs can be listed with `GET /v1/responses/{id}/input_items`, and it can be removed with a DELETE. Stored responses are subject to your organization's retention policy (30 days by default) and can be disabled org-wide — organizations under a zero-data-retention agreement cannot store at all. Setting `store: false` gives you back the stateless model: nothing is persisted, no id can be chained, and the client must resend the full conversation each turn. ## Chaining with previous_response_id `previous_response_id: "resp_abc"` tells the server: take the item list of that response — its inputs and its outputs — and place them ahead of the new `input`. The model therefore sees the full conversation even though the request body contained one sentence. What is *not* carried over is as important as what is: - **`instructions` are not inherited.** If your system-level guidance was passed as `instructions` on turn one, it is absent on turn two unless you pass it again. Guidance you put inside `input` as a message item *is* part of the stored items and does travel. - **Tool definitions are not inherited.** Send `tools` on every call that may need them. - **Model choice is per request.** You can chain a cheap model's response into an expensive model's call. An alternative addressing scheme exists: conversation objects. Instead of threading response ids one after another, you create a conversation and pass its id on each request, letting the server own the item list as a first-class object rather than as a chain of response ids. Either way the token economics are the same. ## The economics — the part interviewers probe Candidates routinely assume that because the server holds the history, they stop paying for it. They do not. The model is stateless underneath; every turn re-tokenises the whole replayed prefix and bills it as input tokens. Server-side state saves **bandwidth and client complexity**, not money. What does reduce the bill is automatic prompt caching on the stable prefix, which applies whether the history came from `previous_response_id` or from your own resent array. The second consequence is the context window. A chain grows without bound, and eventually the replayed prefix plus the new turn exceeds the model's window. The `truncation` parameter decides the outcome: `disabled` (the default) fails the request, while `auto` lets the server drop items from the middle of the conversation to make it fit — convenient, but it means the model silently loses turns you may have considered important. ## Reasoning models and stateless mode Reasoning models produce internal reasoning items alongside the visible message. When responses are stored, chaining preserves those items so the next turn can build on prior reasoning. If you run with `store: false`, they vanish — the workaround is to request `include: ["reasoning.encrypted_content"]` and pass the encrypted reasoning items back yourself on the following call, which keeps the benefit without server-side persistence. ## Practical guidance Use hosted state when the client is thin, the conversation is short-lived, and you have no compliance constraint on storage. Keep your own transcript anyway if you need evaluation datasets, audit logs, or the ability to replay a conversation against a different provider — the hosted copy is convenient, not a system of record. And always resend `instructions` and `tools`: forgetting them produces a maddening bug where the assistant "forgets its persona" on the second turn while the conversation history is plainly intact.

  • Your organization runs under zero data retention. How do you still get multi-turn behaviour?
    Send `store: false` and own the transcript client-side: keep the item list yourself and pass the whole conversation as `input` each turn. Nothing is persisted, so `previous_response_id` is unavailable. With reasoning models, request `include: ["reasoning.encrypted_content"]` and pass the returned encrypted reasoning items back on the next call so the model keeps the benefit of its earlier thinking without the server storing anything.
  • A chained conversation eventually exceeds the model's context window. What controls the outcome?
    The `truncation` parameter. The default, `disabled`, makes the request fail once the replayed prefix plus the new input exceeds the window. Setting it to `auto` lets the server drop items from the middle of the conversation so the request fits. `auto` keeps the app alive but silently loses turns, so anything the model must never forget belongs in `instructions` or a summarised item you re-send, not in the mid-conversation history.
  • How do you audit what the server actually replayed for a given turn?
    Retrieve the response with `GET /v1/responses/{id}` and list what fed it with `GET /v1/responses/{id}/input_items`. That returns the item list the server assembled, which is the authoritative answer to "what did the model see" — far more reliable than reconstructing it from your own logs, especially when `truncation: auto` may have dropped items.

It is like leaving your notes with the receptionist instead of carrying the folder: you walk in empty-handed, but the whole folder is still read aloud at every meeting, and you are charged by the page.

saying these in an interview costs you the question

  • Claiming server-side state means you stop paying for old turns
  • Assuming instructions from the first call persist across the chain
  • Thinking previous_response_id works when store is false
  • Believing chaining removes the context-window limit
  • Treating the stored response as a durable system of record

context