skip to content

After 60 steps an agent ignores its original constraints — what failure is that?

level: seniorimportance: should knowfreq 45%

answer

  1. failure that grows with run length
  2. progress is real, destination is wrong
  3. oldest content, weakest influence
  4. summarizers drop terse constraints first
  5. enforce in the harness, not the prompt

basics

~20 s

Goal drift, also called instruction forgetting: constraints stated at the start lose influence as later observations dominate the context. The agent still works competently, just toward a subtly different objective than the one it was given.

solid answer

~60 s

Goal drift is the long-horizon failure where the original objective and its hard constraints stop steering the agent's choices, while every individual step still looks locally reasonable. It has two usual causes: dilution — the opening instruction sits far back behind tens of thousands of observation tokens and competes with everything since — and lossy carry-over, where summarizing or trimming older context to make room drops the constraints along with the noise. It is distinct from a stuck loop, because the agent is making progress; it is just progressing toward the wrong thing. Detect it by checking each step against the stated constraints rather than only grading the final answer, and by replaying a trace and asking at step N what the agent believes the goal and its restrictions are — the step where the answer changes is where drift began. The mitigation is re-anchoring: keep the goal and hard constraints in the most recent turn, or in a durable artifact the agent re-reads, so they are never the oldest thing in the window.

go deeper

for a junior

Know the term and its signature: the agent keeps working sensibly but stops honouring instructions given at the very start of a long run, because those instructions are now buried in a large context.

for a middle

Explain the mechanics — dilution across a long window, weaker recall of middle-of-context material, and constraints lost when older turns are summarized — and say why restating the goal near the latest turn helps.

for a senior

Demonstrate detection you have actually run: per-step constraint assertions, replaying a trace to find the step where the recalled goal changed, and success rates stratified by trajectory length.

for a principal

Argue for enforcement outside the model — any constraint expressible as a rule belongs in the tool layer — and for bounding horizon length by design, since the drift depth keeps moving as models improve and will surprise anyone relying on memory alone.

## What drift looks like A sixty-step run opens with an instruction containing a goal and three constraints — only use approved sources, never write to production, always attach the case identifier. By step forty-five the agent is doing careful, competent work that violates one of them. Nothing crashed. Every individual action reads as sensible in isolation. Only when you compare the trajectory against the opening instruction does the divergence appear. This is goal drift, and its close relative instruction forgetting. It belongs in the failure taxonomy as its own row because its signature — steady visible progress, locally plausible actions, wrong destination — is unlike anything else on the list. ## Why it happens **Dilution.** By step forty-five the context is dominated by tool observations, intermediate reasoning and partial results. The opening instruction is a few hundred tokens among tens of thousands, and it is the oldest content in the window. Attention over very long contexts degrades unevenly — material in the middle of a long window is recalled less reliably than material at either end, a pattern usually called lost-in-the-middle, and the general degradation of quality as context fills is called context rot. A constraint stated once at the start is a prime victim of both. **Lossy carry-over.** Long runs cannot keep everything, so the harness summarizes or trims older turns to make room. Whatever is dropped stops existing from the model's point of view. Summarizers optimize for narrative continuity — what has been done so far — and constraints are exactly the kind of terse, non-narrative content that gets compressed away. **Local optimization.** Each step is chosen against the immediate observation. If a step-forty tool result suggests an obvious next action that happens to violate a step-one constraint, the immediate signal is loud and the constraint is faint. ## Distinguishing it from neighbours Against a **stuck loop**: a loop shows no state change; drift shows plenty. If the trace is advancing, it is not a loop. Against **wrong-tool selection**: that is one bad choice with the goal intact; drift is a sustained pattern of choices that are individually defensible under a mutated objective. Against **false success**: drift usually still produces real work — the wrong work. False success produces no work and says otherwise. A run can end with both: drifted output confidently declared complete. A useful discriminator is the temporal profile. Drift is a function of run length. Plot per-step constraint compliance and it degrades with depth; a wrong-tool pick is uncorrelated with step index. ## Detecting it **Per-step constraint checks.** Where constraints are mechanically expressible — no writes to this system, only these sources, this identifier always present — assert them on every action. This finds the exact step where compliance broke, which final-answer grading cannot. **Trace replay probes.** Take a stored trajectory, and at step N ask the model what it believes the goal and its restrictions are. Walk N forward. The step where the recalled constraints change is where drift began, and it typically sits just after a summarization boundary or a large observation. **Length-stratified outcomes.** Score task success bucketed by trajectory length. A success rate that falls off sharply past a certain depth, with no matching rise in errors or timeouts, is the fingerprint of drift rather than of hard tasks. ## Mitigating it The organizing principle is that constraints must never be the oldest, most compressed thing in the window. - **Re-anchor.** Restate the goal and hard constraints in or adjacent to the most recent turn, so they always sit in the recency-favoured region rather than the middle. - **Externalize.** Write the objective and constraints into a durable artifact the agent re-reads at each step. Because the artifact is fetched fresh, it never gets summarized away. - **Protect them from compression.** Whatever the carry-over strategy, treat the constraint block as non-summarizable, verbatim content. - **Enforce outside the model.** Anything expressible as a rule belongs in the harness — a tool that rejects writes to production enforces that constraint at step sixty as reliably as at step one, no matter what the model has forgotten. - **Shorten the horizon.** Splitting a sixty-step run into bounded segments, each carrying its own restated objective, reduces the depth over which any single instruction has to survive. ## The honest caveat Drift is a moving target. Models have improved substantially at holding instructions over long contexts, and better carry-over strategies help further, so the depth at which drift appears keeps receding — which is precisely why it keeps surprising teams: it does not vanish, it just moves out to a horizon nobody has tested yet. The durable answer is not "the model is good enough now" but enforcement outside the model for anything that actually matters.

  • How do you tell goal drift apart from a stuck loop when reading a trace?
    Look at whether state advances. A loop repeats an action or oscillates between two with no new information, so the trace grows while the world does not. Drift shows genuine progress — new files touched, new records read, real intermediate results — just aimed at a mutated objective. Loops are caught by repeated-state detection; drift is invisible to it, because every step is novel.
  • Which constraints should you enforce in the harness rather than rely on the agent to remember?
    Anything expressible as a rule and costly to violate: no writes to production, only these data sources, this tenant only, no spend above a threshold. A tool that refuses the action enforces it identically at step sixty and step one, independent of what has fallen out of context. Reserve prompt-level statement for genuinely judgement-shaped guidance that no rule can capture.
  • Why does summarizing older context sometimes make drift worse rather than better?
    Because summarizers optimize for narrative continuity — what has been done so far — and constraints are terse, non-narrative and easy to treat as boilerplate. The summary reads complete while the restrictions have quietly disappeared, and from the model's perspective anything dropped no longer exists. The remedy is to mark the objective and constraint block as verbatim, non-summarizable content, or to hold it in an artifact re-read each step.

saying these in an interview costs you the question

  • Confusing goal drift with a stuck loop
  • Assuming a longer context window eliminates drift
  • Stating constraints once at the start and never again
  • Grading only the final answer, so drift is invisible
  • Trusting the model to remember rules the harness could enforce

context