How do you design a multi-turn chat prompt so each turn reuses the cached prefix?
answer
- conversation as an append-only sequence
- never rewrite earlier turns
- breakpoint moves to the newest completed turn
- each turn reads and extends the prefix
- compaction is a deliberate reset
basics
~20 sTreat the conversation as append-only: never rewrite, reorder or drop earlier turns, and move the cache breakpoint to the end of each completed turn so the growing history becomes the stable prefix that the next request reuses.
solid answer
~50 sA conversation is naturally a growing prefix, so it caches beautifully — but only if you never touch what is already in it. Each request should be the previous request plus new messages at the end, with the breakpoint moved to the end of the newest completed turn. Turn *n+1* then reads the prefix that turn *n* wrote and extends it by one turn; this incremental pattern keeps almost the whole conversation reusable no matter how long it runs. What breaks it is any operation that mutates history: editing or regenerating an earlier message, summarizing older turns in place, re-sorting tool results, re-injecting freshly retrieved documents at the top of every turn, or a header that stamps the turn number or clock time. Compaction inevitably rewrites history — treat it as a deliberate, occasional reset rather than something you do every few turns.
go deeper
Know that a chat caches well because each request is the previous one plus new messages, and that changing an old message means everything after it has to be redone.
Be ready to explain the incremental pattern — the breakpoint moves to the end of the newest completed turn, so each request reads the previous prefix and extends it — and to name what mutates history.
Demonstrate diagnosis: reuse collapsing at a consistent turn points at a summarizer or retrieval refresh firing there; flat low reuse from turn two points at a variable header. Show you would place compaction at a task boundary rather than on a fixed timer.
Own the conflict between context hygiene and prefix stability. Set policy on when history may be rewritten — batched, at natural boundaries, never on a per-turn reflex — and decide whether product features like message editing justify the reuse they cost.
## The append-only invariant Prefix reuse asks for exactly one property from a conversation: request *n+1* must begin with the full token sequence of request *n*. If that holds, every turn reuses everything that came before it and only the newest exchange is new work. A chat transcript satisfies this naturally — you speak, the assistant speaks, you speak again — which is why multi-turn conversations are among the best-behaved caching workloads there are. The invariant is fragile in one direction only: anything that modifies history *behind* the newest message breaks the match at the point of modification, and the entire tail below it is lost. ## Moving the breakpoint each turn The pattern is incremental. On each request, place the breakpoint at the end of the last *completed* turn — the assistant's most recent finished reply — rather than leaving it pinned after the system prompt. The request reads the prefix stored by the previous turn and writes a slightly longer one covering the turn that just completed. Over a fifty-turn conversation this walks steadily down the transcript, and the reusable region grows to be nearly the whole prompt. Leaving the breakpoint fixed near the top is a common half-measure. It caches the system prompt and documents, which is worth something, but it recomputes the entire growing history on every turn — and in a long conversation, history eventually dwarfs the static header. ## What breaks the invariant - **Edit or regenerate.** A user editing their third message, or asking for a different answer to it, rewrites the transcript from that point. Everything after it is new. This is unavoidable when the product offers the feature; just know that the cost is proportional to how far back the edit reaches, and that branching from a recent point is cheap while branching from the first message is not. - **In-place summarization.** Replacing turns 1–10 with a summary paragraph changes the sequence at position one. Every subsequent token is new. - **Retrieval injected at the top.** A system that re-runs retrieval each turn and splices the results in above the history divergences the prompt at the splice point every single turn — reuse never gets past it. Retrieved material for the current turn belongs in the current turn's message, at the tail. - **Reordered or reformatted tool results.** Agent loops that collect several tool results and serialize them in a non-deterministic order, or that re-render prior tool output with fresh formatting, mutate history without meaning to. - **Per-turn headers.** "Turn 7 of 20", a clock time, or a token-budget banner regenerated at the top of every request is a variable block sitting in the most expensive possible position. ## Tool-calling turns An agentic turn is several messages: an assistant message requesting tools, the tool results, then the assistant's reply. Each of these appends, so the invariant holds — provided you append results in a stable order and do not go back to trim or rewrite earlier results after the fact. Note the tension with context management: dropping stale tool output from history is a legitimate context-hygiene move, and it also rewrites the prefix. Both things are true; make the choice deliberately rather than letting a cleanup routine run every turn by reflex. ## Compaction as a deliberate reset Sooner or later a long conversation must be compacted, and compaction is by definition a rewrite of history. Accept it as a planned reset: one expensive turn, after which a new prefix is established and grows again. The design implication is to make it rare and to place it where it hurts least — at a natural boundary such as the end of a completed sub-task rather than on a fixed every-N-turns timer that keeps clipping the prefix just as it becomes valuable. ## Branching is an underrated benefit Because reuse is prefix-based, two conversations that share an opening share their cached head. Forking a conversation at turn twelve to explore two options means both branches reuse the first twelve turns. The same applies to a support system where every conversation starts from the same scripted opening: that opening is a shared prefix across the entire population, not just within one session. ## Symptoms of a broken invariant Reuse that starts healthy and then collapses at a consistent point in the conversation usually means something rewrites history at that point — a summarizer firing at a turn threshold, or a retrieval refresh on the first tool-using turn. Reuse that is flat and low from turn two onwards usually means a variable block at the top: a timestamp, a turn counter, or per-turn retrieval spliced above the transcript.
- A user edits their third message in a twenty-turn conversation. What is the cost?Everything from that message onward is new: seventeen turns must be reprocessed, and the reusable region shrinks back to the system prompt plus the first two turns. Editing recent messages is cheap, editing early ones is not. If the product leans heavily on editing, that is an argument for keeping the static header substantial, so there is always a meaningful floor of reuse beneath any edit.
- Where should freshly retrieved documents go in a multi-turn agent?Inside the current turn's message, at the tail — never spliced above the transcript. Retrieval results change every turn, so placing them above the history makes the prompt diverge at that splice point on every request and history never becomes reusable. Putting them last means only this turn's retrieval is new work, and the whole conversation above it still hits.
- How do you reconcile trimming stale tool results from history with the append-only rule?Accept that they conflict and choose per case. Trimming is good context hygiene and it does rewrite the prefix from the trim point down. The workable compromise is to trim rarely and in batches at a natural boundary, rather than pruning a little on every turn, so you pay one reset instead of clipping the prefix continuously.
Think of it like appending to a log file rather than editing a document: readers can keep everything they have already processed as long as nobody goes back and changes an earlier line.
saying these in an interview costs you the question
- Leaves the breakpoint pinned after the system prompt forever
- Re-injects retrieval above the conversation history each turn
- Summarizes old turns in place every few turns
- Thinks each turn is cached independently of the others
- Adds a turn counter or clock stamp to every request header