How do you keep an agent's replanning anchored to the original goal over a long run?
answer
- drift is cumulative, not a single error
- the goal recedes as history condenses
- repairs optimize what is in front of them
- immutable charter, mutable plan
- a change of goal is escalated, not silent
basics
~20 sTreat the goal and its hard constraints as an immutable artifact that is re-injected verbatim into every replanning cycle and never summarized away. Require each replan to restate the goal and justify how the new plan serves it, and check the final result against the original success criteria.
solid answer
~60 sEach replan is a fresh derivation conditioned on recent context. Over dozens of cycles, two forces pull it away from the goal: the original instruction recedes as the run's history is condensed to fit the window, and each local repair optimizes whatever objective is salient at that moment. A plant-maintenance agent told to **minimize downtime** starts choosing cheaper vendors with longer lead times by step 60 — no single repair was wrong, and the trajectory is. The defence is separating **immutable charter** from **mutable plan state**. The goal, hard constraints and success criteria live in an artifact that is never rewritten and is re-injected verbatim at the head of every replanning cycle; the plan and world state churn around it. Each replan restates the goal and justifies the new plan against it; a periodic audit holding only the charter and the current plan catches drift the executing agent cannot see. The cost is real: re-injecting the charter every cycle spends tokens, and rigid anchoring blocks legitimate goal revision. Goal changes should escalate, not happen silently.
go deeper
Know that over a long run an agent can gradually stop pursuing what it was originally asked to do, and that the goal needs to be repeated to it rather than assumed to persist from the first message.
Explain the two mechanisms — the goal attenuating as history is condensed, and each repair optimizing whatever pressure is locally in front of it — and the basic fix of re-injecting the goal verbatim on every replanning cycle.
Show the separation of immutable charter from mutable plan state, the restate-and-justify discipline with its audit trail, and an independent drift check that does not share the executing agent's context.
Own the boundary between plan revision and goal revision. Decide when an agent may re-scope versus escalate, what per-call overhead the anchoring controls justify at your run lengths, and how success is measured against criteria the agent cannot edit.
## Why replanning is where goals go to die A single plan derived from a goal usually reflects that goal well — the instruction is right there in the prompt. The danger is cumulative. Every replan derives a plan from *the current context*, and on a 90-step run that context is mostly recent observations, tool outputs and repair reasoning. The original instruction is a small, old fragment competing with a large, fresh one. Drift is not a single bad decision; it is a hundred locally-defensible ones pointing slightly downhill. Two mechanisms drive it. **Attenuation.** Long runs cannot keep everything, so history gets condensed. Whatever survives is what the next planner sees. If the goal statement is condensed along with everything else, its exact wording — including the constraint that distinguishes *minimize downtime* from *minimize cost* — degrades into a paraphrase, and paraphrases lose precisely the qualifiers that were doing the work. **Local-objective substitution.** Repairs respond to local pressure. A vendor is out of stock, so the repair reaches for the vendor that is available; that one is cheaper but slower. The next repair does the same. Nothing consults the objective function, because the objective function is not in front of the repair step — the failure is. Over sixty steps the agent has quietly re-optimized for the thing it kept being nudged by. ## The structural fix: immutable charter, mutable state Split the agent's state into two categories with different rules. The **charter** — the goal verbatim, hard constraints, success criteria, prohibitions — is written once, never rewritten by the agent, exempt from any condensation, and re-injected in full at the head of every planning and replanning call. It is small, it is stable, and it is the one thing the agent may not paraphrase. The **working state** — the plan, executed steps, observations, world facts — churns freely, gets condensed, gets rewritten by every repair. Externalizing the plan to a file or artifact makes this natural: the goal sits at the top of the document, unchanged, and the step list below it is edited. Every read of the plan re-reads the goal, which is a structural property rather than a hoped-for behaviour. ## Behavioural reinforcements **Restate-and-justify.** Require every replan to open by restating the goal and its binding constraints, then explain how the proposed plan serves them. This is cheap, and it forces the goal through the model's own attention at exactly the moment the plan is being rewritten. It also gives you an audit trail: the point where the restatement first drifts is the point where the run went wrong. **Independent drift audit.** Periodically — every N replans, or on any full replan — run a separate check that sees only the charter and the current plan, not the run's reasoning. Ask whether this plan still pursues this goal within these constraints. The executing agent cannot reliably audit itself here, because the reasoning that produced the drift is the same reasoning that would have to detect it; an auditor without that context is not similarly captured. **Terminal criteria check.** At completion, evaluate the result against the success criteria as originally written, not against the goal as the agent currently understands it. This catches the run where the agent solved a nearby problem well. ## The tension: anchoring versus legitimate revision Rigid anchoring has a real cost. Sometimes the world genuinely changed and the original goal is no longer the right one — the equipment is being decommissioned, so minimizing its downtime is now pointless. An agent hard-anchored to the charter will pursue a goal that no longer serves anyone. The resolution is not to let the agent re-scope silently. It is to make goal revision a *different kind of event* from plan revision: the agent may rewrite plans freely within the charter, but a change to the charter itself is escalated, with the evidence that motivated it, for a human decision. That boundary is the one most worth being explicit about in a design review, because an agent that can quietly edit its own objective has no meaningful success criteria at all. ## Costs to weigh Re-injecting the charter on every cycle spends tokens on every call — small per call, real at volume. Restate-and-justify adds output tokens and latency to each replan. A periodic auditor is an extra model call. None of this is free, and on short runs it is unnecessary overhead; drift is a long-horizon phenomenon. The judgment call is where your task lengths cross the threshold — roughly, when the number of replanning cycles is large enough that the goal has been reformulated more times than a human reviewed it. ## What a strong answer sounds like Name the two mechanisms (attenuation and local-objective substitution), give the structural fix (immutable charter separate from mutable state, re-injected verbatim), add at least one independent check that does not share the drifting context, and be honest that anchoring trades against legitimate re-scoping — which is why goal changes escalate rather than happen in-band.
- Why can't the executing agent be trusted to audit its own drift?Because the context that produced the drift is the context it would audit from. Every local repair looked reasonable, and the agent's current understanding of the goal is the drifted one. An auditor given only the original charter and the current plan — with none of the intervening reasoning — is not captured by the same chain and can see the gap.
- When is a goal change legitimate rather than drift, and how should the agent handle it?When the world genuinely changed such that the original objective no longer serves anyone — the equipment is being decommissioned, so minimizing its downtime is moot. The agent should not re-scope in-band. It escalates with the evidence and a proposed revision, because an agent that can silently edit its own objective has no meaningful success criteria left.
- What does the restate-and-justify step give you beyond nudging the model?An audit trail. Each replan records the goal as the agent then understood it, so when a run ends badly you can find the cycle where the restatement first slipped and diagnose the cause. Without it, drift is only visible by comparing the final behaviour against the original instruction, long after the divergence point.
saying these in an interview costs you the question
- Assuming the goal stays in context because it was in the first prompt
- Letting condensation summarize the goal along with the run history
- Trusting the executing agent to notice its own drift
- Allowing the agent to silently rewrite its own objective
- Adding drift controls to short runs where they are pure overhead