skip to content

When is OpenAI's hosted conversation state the wrong choice for production?

level: principalimportance: should knowfreq 42%

answer

  1. convenience layer, not a system of record
  2. who owns the transcript?
  3. some contracts forbid storing at all
  4. evaluation needs replayable history
  5. an id that means nothing elsewhere

basics

~20 s

Hosted state is wrong when storage is contractually forbidden, when transcripts must be your own for audit, evaluation or portability, or when you need explicit control over what the model sees each turn. It saves client complexity, never token cost.

solid answer

~50 s

Server-side conversation state buys one thing: the client stops carrying the transcript. Weigh that against what it costs. Zero-data-retention agreements and some regulated environments forbid storing content at the provider at all, which forces `store: false` and client-owned history. Even where storage is allowed, you usually need your own transcripts anyway — for evaluation sets, incident replay, audit, and deletion requests — at which point the hosted copy is a duplicate rather than a system of record. It also creates lock-in: a conversation addressed by a provider-issued id cannot be replayed against another model vendor. And it hands away context control, since what the model sees is assembled server-side and, with `truncation: auto`, may silently drop turns. The pragmatic pattern is to own the transcript and use hosted state as an optimisation you can switch off.

go deeper

for a junior

Know that you can either let the provider keep the conversation or keep it yourself, and that keeping it yourself means sending the whole history with each request.

for a middle

Explain the mechanics of turning it off — storage disabled, full history resent each turn — and state plainly that hosted state saves payload and plumbing, not tokens.

for a senior

Bring the operational cases: retention agreements, audit and deletion requests, replaying real conversations against a new prompt, and controlling what gets dropped at the context limit.

for a principal

Own the call and its blast radius. Set the transcript as the system of record on your side, treat hosted state as a switchable optimisation, and justify it against portability, compliance, and provider-deprecation risk.

## Framing the decision "Should the platform hold the conversation?" sounds like an API-ergonomics question and is actually an architecture question about where your system of record lives. Answer it once, deliberately, because retrofitting a transcript store into a product built on provider-held state is painful. ## What hosted state genuinely buys - **Thin clients.** A mobile or edge client sends one turn instead of a growing array. Payload size stops scaling with conversation length. - **Less plumbing.** No transcript table, no serialization of tool-call items, no bug where you forget to append a tool result. - **Reasoning continuity.** With reasoning models, stored responses preserve internal reasoning items across turns without you shipping them back and forth. These are real and they are why the feature exists. For a prototype, an internal tool, or a consumer chat with no compliance surface, hosted state is the right default. ## What it does not buy **It does not reduce your bill.** The model still reads the whole conversation each turn and you are charged those input tokens. Prompt caching discounts the repeated prefix whether the prefix came from the server or from you. Anyone who justifies hosted state on cost has the mechanism wrong. **It does not give you durability guarantees you would choose.** Stored responses live under the provider's retention policy — a limited window by default, adjustable at the org level, and deletable. That is a cache, not an archive. ## The five reasons to say no 1. **Contractual and regulatory.** Under a zero-data-retention arrangement, nothing may be persisted provider-side; storing is simply unavailable and the client must send the full history each turn. Data-residency commitments can push the same way. 2. **Audit and incident response.** When a user disputes what the assistant told them, you need the exact transcript on your side, with your timestamps and your identity context. Reconstructing it from provider-held objects is a dependency you do not want in an incident. 3. **Evaluation.** Building regression suites means harvesting real conversations, replaying them against a new model or prompt, and diffing. That requires the raw item list in your own store, in a shape you control. 4. **Portability.** A provider-issued conversation id means nothing anywhere else. Teams that intend to route across vendors, or simply keep the option open, must keep the canonical transcript in a neutral format — which they then have to send anyway. 5. **Context control.** Serious applications curate what the model sees: summarising old turns, dropping tool noise, pinning key facts, injecting fresh retrieval. Hosted state assembles the prefix for you, and the only lever at the limit is a truncation setting that either fails the request or silently discards middle items. If context engineering is where your quality comes from, own the assembly. ## The hybrid that usually wins Write every turn to your own store as the system of record. Then choose per surface: interactive clients may chain server-side for latency and payload size, while batch, evaluation and replay paths reconstruct the conversation explicitly from your data. Because the canonical copy exists, switching off hosted state is a configuration change rather than a migration. This also means a provider deprecation is survivable — a fact worth weighing given that the earlier stateful surface, with its threads and runs, is being retired in favour of the newer one, and any team that had made those objects their source of truth is now doing a data migration rather than a client refactor. ## Questions to ask in the design review - Where does a deletion request land, and can you satisfy it end to end? - If the provider had an outage, could you serve a user's history at all? - Can you replay last week's worst conversation against a candidate prompt today? - Which component decides what is dropped when the conversation outgrows the window, and is that decision visible? - If pricing or capability made a second vendor attractive next quarter, what would move? A candidate who answers those crisply is not arguing against hosted state; they are showing they know it is a convenience layer over a stateless model, and placing it accordingly.

  • An organization with zero data retention wants multi-turn chat on the Responses API. What changes?
    Storage is off, so `store` must be false and the response ids cannot be chained. Your application keeps the item list and sends the full conversation as input on every turn — exactly the stateless pattern. With reasoning models, request encrypted reasoning content in the response and pass those items back on the next call so the model retains its prior thinking without anything being persisted provider-side.
  • Someone argues hosted state will cut their token bill. How do you correct them?
    The model is stateless underneath: whether the history arrives in your request body or is replayed from server-side storage, it is tokenised and billed as input on every turn. What server-side state saves is upload bandwidth and client code. The actual cost lever is the repeated-prefix cache discount, which applies in both arrangements, plus curating the context so less of it is resent — and curation is easier when you own the transcript.
  • What is your migration story if the provider retires the stateful surface you built on?
    If your own store holds the canonical transcript, migration is a client refactor: map your item list onto the new request shape and redeploy. If provider-held objects were the source of truth, it is a data migration under a deadline — export every conversation before shutdown, reshape it, and reconcile. The retirement of the older threads-and-runs surface in favour of the newer stateful API made exactly that distinction expensive for teams on the wrong side of it.
  • How do you handle a conversation that outgrows the context window when you own the transcript?
    You get to choose the policy instead of accepting a default. Common approaches: summarise older turns into a compact item and keep the last few verbatim; drop tool-call noise while keeping decisions; pin invariant facts into the instructions so they can never be evicted. Each is a deliberate, testable rule, whereas an automatic truncation setting either rejects the request or silently removes middle items you may have needed.

saying these in an interview costs you the question

  • Claiming server-side state lowers token cost
  • Treating provider-stored conversations as the system of record
  • Ignoring that zero-data-retention orgs cannot store at all
  • Assuming a provider conversation id is portable to another vendor
  • Letting automatic truncation decide what the model forgets

context