skip to content

Why does system-prompt adherence decay over a long conversation?

level: middleimportance: must knowfreq 62%

answer

  1. it is still there, just quieter
  2. distance from the newest turn
  3. the transcript teaches the model
  4. one slip makes the next likelier
  5. tokens of history, not turn count

basics

~20 s

Adherence decays because the standing rule sits at the far end of a growing history: it competes with thousands of tokens of newer, more concrete dialogue, and the model's own earlier replies — including any that already broke the rule — become the strongest example of how this conversation behaves.

solid answer

~50 s

Three effects compound. **Position and competition**: the system prompt is the oldest block in the request, and as dialogue accumulates between it and the newest user turn, a one-line constraint competes for attention with far more recent, far more specific material. **Self-imitation**: the moment the model violates the rule once, that violation is in the transcript as an assistant turn, and an in-conversation pattern is a powerful signal about how to continue — drift is self-reinforcing rather than a fresh coin flip each turn. **Instruction load**: a system prompt carrying fifteen standing rules loses the low-salience ones first, typically tone and formatting before hard behavioural rules. A Spanish-only language tutor that starts slipping into English corrections around turn thirty is the classic shape: it is not that the rule was removed, it is that the rule stopped being the loudest thing in the window.

go deeper

for a junior

Know the phenomenon by name: in long chats, rules given at the start get followed less reliably later, even though the system prompt is still sent with every request. Be able to give one concrete example, such as a bot slowly abandoning a required output format.

for a middle

Explain the mechanics: the rule's shrinking share of a growing window, its distance from the newest turn, and the model imitating its own earlier replies once one violation is in the transcript. Mention that many stacked rules dilute each other.

for a senior

Show you can diagnose it in a running system — bucket adherence by conversation length or tokens of history, distinguish a real decay curve from sampling noise, and resist the reflex to fix it by making the system prompt longer and louder.

for a principal

Own the framing that unassisted adherence has a practical depth ceiling, and that the ceiling is a design input: it decides which constraints may live in a prompt at all and which must be enforced outside the model. Set the drift budget the product can tolerate.

## What drift actually looks like Instruction drift is the gradual erosion of system-prompt adherence as a conversation gets longer. Nothing dramatic happens at any single turn. A tutor told to reply only in Spanish answers in Spanish for twenty-five turns, then starts glossing a grammar point in English, then answers the next question mostly in English. A support assistant told to end every reply with a ticket reference does so reliably early and omits it by turn forty. A model told to keep answers under three sentences slowly returns to its default essay length. Note that the system prompt is almost always still there — chat APIs are stateless, so the client resends the whole conversation, system block included, on every request. Presence is not the problem; **salience** is. ## Why the rule loses ground **Distance and competition.** The system prompt is the oldest text in the request. Early in a chat it is a large share of everything the model is conditioned on. By turn fifty it may be one percent of the window, sitting far from the position where generation begins, competing with thousands of tokens of newer, more concrete, more topically relevant material. Attention is finite and is spread across the whole sequence; a short abstract rule is easy to out-shout with a long specific discussion. This is the same phenomenon people describe as context rot: quality of instruction-following falls as the occupied window grows, even well inside the nominal limit. **Self-imitation.** This is the effect people underestimate. Once the model produces one reply that breaks the rule, that reply is appended to the history as an assistant turn and becomes part of the conditioning for every subsequent turn. In-context examples are the strongest steering signal available to a language model, and the transcript is now an example set demonstrating the wrong behaviour. So drift is not independent per turn — it is a ratchet. Recovery rarely happens on its own; the conversation has to be pushed back. **Instruction load and salience ranking.** Rules are not equally sticky. A prompt with fifteen standing constraints does not decay uniformly: constraints that are reinforced by the content of the conversation survive longer, and constraints the dialogue never touches fade first. Tone, persona and formatting rules typically slip earliest because nothing in the exchange re-evidences them; a rule that the user's own messages keep implicitly reinforcing ("in grams, please") survives longer. Practically, ten rules do not get ten times the adherence of one — they dilute each other. ## Turns or tokens? Turn count is the convenient proxy, but token distance and content volume are the better predictors. Forty short turns totalling three thousand tokens drift far less than forty turns totalling ninety thousand tokens with pasted documents and long tool outputs in between. Two other factors matter: how much of the recent window is model-generated (self-imitation pressure), and whether anything has re-stated the constraint recently. Two conversations at the same turn number can be in completely different adherence regimes. ## Which failures matter Separate **persona and tone slip** — the assistant stops sounding like the character, gets more verbose, drops a stylistic convention — from **hard rule violation** — it discloses something it was told never to disclose, gives advice it was told never to give, or breaks an output format a downstream parser depends on. Both are drift; only one is usually an incident. Which of the two your product cares about determines how much you spend fighting drift and where you enforce the rule. ## How you notice You will not see this in a single-turn test set, because every case there starts at turn one where adherence is at its best. Drift is only visible when you replay long conversations and check the same constraint at several depths — for example asserting it at turn five, turn twenty and turn fifty — and look at the adherence curve rather than a pass/fail. In production, the same idea shows up as adherence rate bucketed by conversation length or by tokens of history; if the metric falls off a cliff past some depth, that is your practical ceiling for unassisted adherence. ## What it is not Drift is not the model deciding the user outranks the system prompt, and it is not the system prompt being stripped or expiring. It is also not randomness alone: temperature adds variance, but the trend across a long chat is directional, driven by dilution and by the model imitating its own recent output. Diagnosing it as "the model ignored my instruction" leads people to write a longer, angrier system prompt, which usually makes things worse — more instructions means more dilution.

  • Does resending the full system prompt on every request prevent drift?
    No. Chat APIs are stateless, so the system prompt is resent every turn anyway — its presence was never the issue. What changes over a long conversation is its share of the window and its distance from the generation point, plus the accumulating example of the model's own recent replies. A rule can be perfectly present and still be out-weighted by newer, more specific context.
  • Which constraints tend to survive longest in a long chat, and why?
    Constraints that the conversation itself keeps re-evidencing. If the user repeatedly asks for quantities and the assistant keeps giving grams, the pattern is reinforced in-context every few turns. Abstract rules the dialogue never touches — tone, a rarely triggered formatting rule, a confidentiality clause — have nothing refreshing them and fade first. Rule count matters too: fifteen standing rules dilute each other.
  • How does temperature interact with drift?
    Temperature adds per-turn variance, so a lower temperature makes any single violation less likely, but it does not remove the underlying trend — the rule still competes with a growing history, and once a violation lands it conditions later turns regardless of sampling settings. Treat low temperature as noise reduction, not as a fix for adherence decay.

It is like a house rule announced once at the start of a long party: nobody deleted the rule, but by hour four everyone is copying what the room is currently doing rather than what was said at the door.

saying these in an interview costs you the question

  • Claims the system prompt is dropped or expires after N turns
  • Says the model chose to obey the user over the system prompt
  • Believes a longer, more emphatic system prompt fixes decay
  • Thinks drift is pure sampling randomness with no trend
  • Measures adherence only with single-turn test prompts

context