How do you stop an agent from thrashing between replans of the same failed step?
answer
- count attempts per subgoal, not per run
- the global cap fires far too late
- each replan must differ observably
- show the model its own failed attempts
- escalate with evidence, never silently
basics
~20 sCount replan attempts per failing subgoal, not just globally, and cap them at a small number such as three. Require each new plan to differ observably from the last and to state what the previous failure taught. On exhaustion, escalate with the failed plans and the deviation evidence attached.
solid answer
~50 sReplan thrash is the loop where an agent regenerates a near-identical plan for a step that keeps failing for a reason the plan cannot address. The pump is discontinued; the agent tries vendor B, then vendor C, then vendor A again, each time producing a fresh, confident, useless plan. Three controls handle it. First, a **per-subgoal attempt counter** — three replans against the same failing step is usually the ceiling — rather than relying on the global iteration cap, which fires far too late and yields no diagnosis. Second, a **difference requirement**: each replan must change something observable (different tool, relaxed constraint, different assumption) and must be shown the previous attempts and why they failed, or it will paraphrase them. Third, a **structured escalation** at exhaustion: hand a human the goal, the attempts, the deviation evidence and the current world state. The failure to avoid is silent surrender — an agent that gives up and reports success.
go deeper
Know that an agent can regenerate the same broken plan over and over, and that there should be a small limit on how many times it retries a failing goal before asking a human.
Explain why the counter belongs to the failing subgoal rather than the whole run, and why each replan must be shown the previous attempts and required to change something observable, or it will paraphrase them.
Show what a good escalation payload contains and why capping is only safe when the output at the cap is diagnostic. Distinguish thrash from a genuine search that is eliminating options attempt by attempt.
Own the calibration across task classes and the escalation contract. Caps that are too tight flood humans with escalations they learn to ignore; caps that are too loose burn spend. Both are organizational costs, not just agent settings.
## What thrash looks like An agent hits a deviation, replans, executes, hits the same deviation, replans again. Each plan is fluent and plausible. None of them can work, because the blocker is a fact about the world that no rearrangement of steps changes — the part is discontinued, the account lacks permission, the API the plan needs does not exist. Left alone, the agent burns its entire budget cycling, and because every cycle is *locally* reasonable, nothing in the trace looks obviously wrong until you compare attempt 1 with attempt 4 and see they are the same plan in different words. ## Why the global loop cap is not the answer Runs carry coarse backstops — maximum iterations, wall-clock limits, spend ceilings — and they will eventually stop a thrashing agent. They are the wrong instrument for three reasons. They fire long after the information needed to diagnose the problem was available, typically after 20+ wasted steps. They produce a generic exhaustion outcome that says nothing about *what* was impossible. And they cannot distinguish a productively long run from a stuck one; raising the cap to accommodate genuinely long tasks also raises the ceiling on thrash. The replan-specific control is finer: attach a counter to *the subgoal*, not the run. Step 4 has now been replanned three times. That is a specific, diagnosable condition, and three is a reasonable default ceiling — beyond it, empirically, the marginal attempt almost never succeeds where the first three failed. ## Making each attempt actually different A fresh planning call over a context that still describes the original situation produces the original plan again. Diversity has to be forced: - **Show the failures.** Each replan prompt should carry the previous attempts and the concrete reason each one failed. Without this, the model has no way to know it is repeating itself. - **Require an observable change.** The new plan must differ in a way you can check — a different tool, a different resource, a relaxed constraint, a different decomposition of the blocked step. A plan that only rewords the last one should be rejected before it executes. - **Require a stated diagnosis.** Ask the replanner to say what the last failure implies about the world before proposing the next plan. "Vendor B also has no stock, which suggests the part is discontinued rather than out of stock at one supplier" is the reasoning that turns attempt 3 into an escalation instead of attempt 4. A useful similarity check is cheap and mechanical: if the new step list is near-identical to a previous one for the same subgoal, treat that as an exhausted attempt rather than a new one. ## Escalating well When the cap is hit, what the agent hands back matters as much as the fact that it stopped. A bare "I could not complete this task" wastes the run's most valuable output — everything it learned about why the task is blocked. A good escalation payload contains: the original goal verbatim; the plan attempts, each with the step that failed and the observation that falsified it; the inferred blocker ("pump model appears discontinued across all three approved vendors"); the current world state including any side effects already committed; and, where the agent can form one, a concrete question that would unblock it ("is an equivalent replacement model approved for this site?"). That payload turns a failed run into a triaged ticket. It is also what makes capping safe: a low cap is only acceptable if what comes out at the cap is useful. ## The failure modes at both ends **Silent surrender** is the worse one. An agent that hits its limit and reports the task as done — or reports a partial result phrased as a success — corrupts everything downstream that trusts its output. Termination on exhaustion must be explicitly labelled as failure, and the run's success criteria should be evaluated independently of the agent's own claim. **Capping too aggressively** is real too. Some tasks legitimately need a second or third approach — the first vendor genuinely is out of stock and the second genuinely has it. A cap of one turns ordinary recoverable friction into escalations and trains humans to ignore the escalations. Tune per task class: routine, well-understood workflows tolerate a low cap; exploratory work needs more room. ## Distinguishing thrash from progress The cleanest discriminator is whether each attempt *learns something*. Attempt 2 that ruled out one vendor and attempt 3 that ruled out another are making progress through a finite search space, even though they look repetitive. Attempt 2 that re-proposes vendor A after vendor A already failed is thrash. Tracking which options have been eliminated — and feeding that set into the next replan — converts what looks like a loop into a search, and lets the agent conclude the space is exhausted rather than cycling in it.
- How do you tell a productive sequence of replans from thrash?By whether each attempt eliminates something. Ruling out vendor A, then vendor B, is a search through a finite space and is progress even though it looks repetitive. Re-proposing vendor A after it already failed is thrash. Tracking the eliminated set and feeding it into the next replan makes the distinction mechanical rather than a judgment call.
- What belongs in the escalation payload when the replan cap is exhausted?The goal verbatim, each plan attempt with the step that failed and the observation that falsified it, the inferred blocker, the side effects already committed, and a specific question that would unblock the run. That turns a dead run into a triaged ticket. A bare 'could not complete' throws away everything the run learned.
- Why not simply raise the iteration cap and let the agent keep trying?Because a higher ceiling raises the cost of every stuck run without improving the odds — the blocker is usually a fact about the world that no rearrangement of steps changes. Raising the cap also blinds you: the diagnosis was available after attempt two, and twenty more attempts add spend and latency while burying it.
saying these in an interview costs you the question
- Relying on the global iteration cap to catch replan loops
- Reporting success or a partial result after exhausting replan attempts
- Replanning without showing the model its previous failed attempts
- Accepting a reworded copy of the failed plan as a new attempt
- Escalating with no evidence beyond 'the task could not be completed'