As an OpenAI chat outgrows the context window, how do you decide which history to resend?
answer
- The API errors, it never trims
- Window, summary, retrieval
- Recent turns stay verbatim
- Reserve the answer's room first
- Stable prefix keeps caching alive
basics
~20 sThere is no default — the API errors rather than trimming for you, so the policy is yours. Choose between a sliding window, rolling summarization, and retrieving only relevant past turns, budgeting tokens against the window and leaving room for the answer.
solid answer
~50 sBecause Chat Completions is stateless and re-sends the whole `messages` array, an unbounded conversation eventually exceeds the model's context window and the request fails with a context-length error — nothing is dropped silently. So you pick a policy. A **sliding window** keeps the system message plus the last N turns: cheap, predictable, and it forgets facts from early on. **Rolling summarization** compresses old turns into a synthetic message: preserves gist at the cost of an extra model call and lossy detail. **Retrieval** stores turns externally and re-injects only the relevant ones: scales furthest, adds the most machinery, and can drop context that looked irrelevant. Most production systems combine them — pinned system message, a durable facts block, recent turns verbatim, older material summarized. Budget explicitly: history tokens plus reserved output must fit the window, and keep the leading prefix byte-stable so prefix caching still hits.
go deeper
Know that the conversation array grows every turn and that the request eventually fails on the context limit, so something has to be dropped or condensed.
Compare a sliding window against summarization concretely, and explain how you budget tokens so the prompt plus the reserved answer fits the window.
Design the layered composition — pinned instructions, durable facts, retrieved snippets, verbatim recent turns — and instrument token usage, truncation rate and cache hits.
Own the tradeoff across cost, recall quality and operational complexity, choose per product rather than per fashion, and insist the policy is validated on long-conversation task success before rollout.
## Why this is a decision and not a setting Every Chat Completions request carries the entire conversation, and the sum of what you send plus what you reserve for the answer must fit the model's context window. Exceed it and the API returns a 400 with a context-length error before generating anything. There is no server-side trimming and no flag that says drop the oldest — the policy is application code, which is exactly why this is asked as a judgment question. ## The options **Keep everything.** Correct until it is not. Fine for short-lived, bounded interactions such as a support form or a one-screen tool. It fails on any long-running assistant, and gets expensive well before it fails, because early turns are re-billed on every call. **Sliding window.** Always send the system message, then the most recent N turns or the most recent T tokens. Trivial to implement, constant cost per call, predictable latency. The cost is amnesia with a hard edge: the user's name, the constraint they gave in turn two, the file they mentioned — all silently gone. Mitigate by pinning a small durable facts block right after the system message. **Rolling summarization.** When the transcript crosses a threshold, ask a model to compress the oldest chunk into a short synthetic message and replace those turns with it. Keeps long-range gist and bounds growth. It costs an extra call at the moment of compression, adds latency at an unpredictable turn, and is lossy in a way you cannot audit — details vanish and the model later states them wrong with full confidence. Use a cheaper model for the summarizer, summarize in the background rather than in the user's critical path, and keep the raw transcript in your own store so nothing is truly lost. **Retrieval over history.** Embed and index past turns, then re-inject only the passages relevant to the current message. This is the only approach that scales to conversations far larger than any window, and the only one that can answer a question about something said months ago. It also brings a whole retrieval subsystem — indexing, freshness, relevance tuning — and fails in a nastier way, since irrelevant-looking but load-bearing context gets dropped and the model confabulates around the hole. ## The composition most teams land on A layered request: pinned system message; a small structured facts block the app maintains explicitly (user identity, active task, hard constraints); optionally retrieved snippets; a summary of older conversation; then the last several turns verbatim. Recency stays verbatim because that is where coreference lives — "do that again with the other one" is unresolvable from a summary. ## Budgeting and caching Make the budget explicit and enforce it before the call, not after the error. Reserve output tokens first, subtract the fixed prefix, and let history fill what remains. Two second-order effects deserve attention. Cost: OpenAI's automatic prompt caching discounts a long *identical leading prefix*, so an architecture that rewrites the top of the array every turn pays full price forever — put the volatile material at the end, keep the head byte-stable, and watch `usage.prompt_tokens_details.cached_tokens` to confirm. Quality: models attend unevenly across a very long context, so a bigger window is not automatically a better answer; more history can dilute the instruction that matters. ## Measuring it Instrument `usage.prompt_tokens` per turn, the rate of context-length errors, the rate of `finish_reason` `"length"`, and the cache-hit ratio. Then measure the thing that actually matters — task success across long conversations — with an evaluation set that includes questions whose answer was established twenty turns ago. Without that, trimming policy gets tuned on cost alone and quality regressions ship invisibly. ## What a strong answer sounds like Name the failure mode (hard error, not silent truncation), lay out the three strategies with honest tradeoffs, describe the layered composition, and connect it to cost, caching and evaluation. Recommending one strategy unconditionally is the weak answer; the right one depends on conversation length, how much long-range recall the product genuinely needs, and how much machinery the team can operate.
- Why not just switch to a model with a much larger context window and keep everything?It buys headroom, not a policy. Cost still scales with what you send on every turn, attention quality degrades across very long inputs so the answer is not automatically better, and latency grows with prompt size. A larger window raises the ceiling and delays the decision, but any conversation that runs long enough still needs a trimming or retrieval strategy.
- Where do you put a fact the assistant must never forget, like a user's stated constraint?In an application-maintained structured block near the top of the array — extracted deliberately, not left to survive by luck in the transcript. Summarization and sliding windows both lose details probabilistically, so anything load-bearing should be state your code owns and re-injects verbatim every turn, updated when the user changes it.
- How would you validate a new trimming policy before rolling it out?Replay stored conversations through both policies offline and compare on an evaluation set that deliberately asks about facts established many turns earlier, scoring task success rather than token count. Then ship behind a flag and watch prompt tokens per turn, context-length error rate and cached-token ratio alongside the quality metric, since a policy that only improves cost is a regression in disguise.
saying these in an interview costs you the question
- Assumes the API drops old messages automatically when full
- Says a bigger context window removes the need for any policy
- Summarizes recent turns and keeps the oldest verbatim
- Rewrites the top of the array each turn, killing prefix caching
- Treats summarization as lossless and never keeps the raw transcript