How does an LLM agent detect that its plan has gone stale mid-execution?
answer
- reality answers only through tools
- errors are the easy case
- the silent success is the dangerous one
- postconditions written at planning time
- never trust the agent's own narration
basics
~20 sDetection comes from checking each step against an expected outcome instead of assuming success. Three signals dominate: an explicit tool error, tool output that contradicts an assumption the plan was built on, and a postcondition check on the step's result that fails.
solid answer
~50 sExplicit errors are the easy case — a tool returns a failure and the agent knows the step did not happen. The dangerous case is the **silent** deviation: the call succeeds, and the answer invalidates the plan anyway. A plant-maintenance agent plans a pump swap where step 4 assumes a spare is in stock; the inventory lookup returns cleanly with a count of zero. Nothing errored, but the plan is dead. So the practical design is to attach an expected **postcondition** to each step at planning time ("after this step, a spare pump is reserved") and to record the **assumptions** each step depends on. After execution you check the observation against both. For fuzzy steps where no assertion is possible, a cheap model call can judge whether the observation actually satisfies the step's intent. What you never rely on is the agent's own narration that the step went fine.
go deeper
Know that agents must verify each step's result rather than assume it worked, and that a tool call can succeed while still breaking the plan — an inventory lookup returning zero is a clean response and a dead plan.
Be ready to name the three signal classes — hard tool errors, contradicted assumptions, failed postconditions — and to explain why postconditions must be written when the step is planned, not after the observation is in context.
Show the operational side: which checks are programmatic versus judged, what you log on every deviation event, and how you tune sensitivity so the agent neither sails past a dead plan nor replans on every harmless surprise.
Own the economics. Per-step verification adds latency and tokens to every run; decide where that spend buys real reliability, which steps get cheap assertions versus judges, and what the organization accepts as the residual rate of undetected silent failures.
## The problem detection solves An agent's plan is a set of predictions about a world it cannot see directly. Every step encodes assumptions — that a resource exists, that an earlier step produced usable output, that a window of time is still open. Tools are the only channel through which reality answers back. Replanning is worthless if the agent never notices reality disagreed, so *detection is the hard half of replanning*, not the trivial preamble to it. ## Three classes of signal **1. Hard failure signals.** The tool returned an error, a timeout, a non-2xx status, a malformed payload, a schema-validation failure. These are cheap, unambiguous, and mostly free — the execution layer surfaces them without extra reasoning. They are also the minority of real deviations. **2. Contradicted-assumption signals.** The call succeeded and its content invalidates the plan. Inventory returns zero. A document exists but is the wrong revision. A query returns an empty list where the plan assumed rows. These produce no error anywhere in the stack, which is exactly why naive agents sail past them and execute three more steps on a dead branch. Catching them requires that the *assumption itself* be written down somewhere checkable. **3. Postcondition signals.** Each step is authored with a statement of what must be true after it runs — "a spare pump is reserved", "the diagnostic report contains a vibration reading", "the ticket is in state Approved". After execution you evaluate that statement against the observation. This is the most general mechanism because it catches both of the classes above and also catches partial success (the step half-happened). ## How to make postconditions actually checkable The strongest postconditions are **programmatic**: a field equals a value, a count is greater than zero, a file exists, a schema validates, a test passes. They cost nothing per check, are deterministic, and cannot be talked out of a verdict. Author them at planning time, when the model is already reasoning about what each step is for — asking for them later, after the observation is in context, invites the model to rationalize whatever happened as success. Where the step's success is genuinely fuzzy ("summarize the incident history usefully"), the fallback is a **judge**: a small, separate model call given the step's stated intent and its observation, asked only whether the intent was met. Judges are probabilistic and cost tokens and latency on every step, so reserve them for steps where no assertion exists. Their bias is toward leniency — a judge sharing the generating context tends to agree with it, which is why the check should see the intent and the observation, not the agent's reasoning about them. ## Assumptions as first-class plan metadata A useful practice is to keep, alongside the step list, an explicit register of the facts the plan rests on and which step each fact came from: *spare pump in stock (assumed, unverified)*, *maintenance window 02:00–06:00 Saturday (from scheduling tool at T0)*. Then any observation that touches a registered fact triggers a comparison. This turns "did something surprising happen?" — an open-ended question a model answers badly — into a small set of equality checks, and it makes the later repair decision tractable, because you know exactly which assumption broke. ## The two failure directions **Under-detection** is the classic demo-agent failure: the agent continues executing a plan whose foundation collapsed at step 4 and confidently reports success at step 9. The cost is wasted steps plus, worse, real side effects committed on a dead plan. **Over-detection** is a real cost too. If every mild surprise trips a replan, the agent thrashes: it burns a planning call per step, latency balloons, and the plan churns without converging. Tune by making postconditions assert *what the plan actually needs*, not everything the step returned. A retrieval step that wanted three relevant documents and got four should not fire a deviation. ## What not to treat as a signal The agent's own claim that a step went well is not evidence. Self-narration is generated from the same context that produced the plan and inherits its optimism; models routinely assert completion of steps whose observations show otherwise. Ground the check in the observation, in a deterministic assertion over it, or in a separately-prompted judge — never in the executing agent's summary of itself. ## What to log Every detection event should record: the step, the assumption or postcondition that failed, the observation that falsified it, and the resulting decision. That record is what makes the subsequent repair well-scoped instead of a blind rewrite, and it is the raw material for diagnosing why an agent failed a week later.
- Why author the postcondition when the step is planned rather than after the tool returns?Because a model shown the observation first will rationalize it. Asked after the fact whether the step succeeded, it reads the output and constructs a story in which it did. Authored up front, the postcondition is a commitment made before the evidence exists, so it can genuinely be falsified — the same reason you write the assertion before running the test.
- Where does a judge model beat a programmatic assertion for deviation detection?Only where success is not expressible as a predicate — open-ended writing, summarization, research adequacy. There a small judge given the step's intent and its observation is the only available check. Pay the token and latency cost per step, accept that it is probabilistic and leniency-biased, and keep programmatic assertions everywhere they exist.
- How do you keep deviation checks from firing on every harmless surprise?Assert what the plan needs, not everything the step returned. If step 4 requires at least one reserved spare, check that, not the exact inventory count or the response shape. Over-tight postconditions produce constant false deviations, and an agent that replans on every wobble converges slower and costs more than one that never replans.
saying these in an interview costs you the question
- Assuming a tool call that returned 200 means the step succeeded
- Trusting the agent's own statement that a step went fine
- Only checking for thrown errors and never for contradicted assumptions
- Asking the model after the fact whether the step worked
- Firing a full replan on any unexpected detail in tool output