When should an agent repair a single plan step instead of replanning fully?
answer
- blast radius, not error severity
- which later steps used that fact?
- step, suffix, or whole plan
- a replan cannot un-send an email
- full replan on a dead premise
basics
~20 sRepair locally when the broken assumption belongs only to the failing step. Rewrite the remaining suffix when downstream steps depended on it. Replan wholly when the goal's feasibility or a global constraint changed. Scope follows the blast radius of the invalidated assumption, not the severity of the error.
solid answer
~50 sThe decision rule is: **which other steps rested on the assumption that just broke?** Take a plant-maintenance agent whose step 4 assumed a spare pump from vendor A. If the assumption that broke is only *this vendor has stock*, swapping in vendor B is a step-local repair — every later step still holds. If the assumption that broke is *the maintenance window is Saturday 02:00*, then scheduling, technician dispatch and notification steps all inherited it, so you rewrite the suffix from that point. If the goal itself is no longer reachable — the line cannot be taken down this quarter at all — that is a full replan, and often an escalation rather than a rewrite. Two constraints bound this. Full replans are expensive in tokens and latency and discard reasoning that was still valid. And already-committed side effects are not undone by replanning: a dispatched technician is a fact the new plan must accommodate.
go deeper
Know that not every failure means starting over — often one step can be swapped while the rest of the plan still stands, and that replanning only changes future steps.
Explain the three scopes — step-local, suffix, whole plan — and tie the choice to which later steps consumed the assumption that broke, not to how alarming the error message was.
Demonstrate the operational judgment: computing the dependent set from an assumption register, holding the executed prefix and its side effects fixed, adding compensating actions, and feeding the replanner condensed world state rather than the raw transcript.
Own the policy. Decide where an agent may re-scope autonomously versus escalate with evidence, what a replan costs per run at your volume, and how much irreversible authority you delegate before a human must confirm a rewritten plan.
## Repair scope is a blast-radius question When a step deviates, the instinct is to grade the *severity* of the error and replan harder for worse errors. That is the wrong axis. The right question is structural: **how far does the invalidated assumption reach?** A catastrophic-sounding failure whose assumption was local needs a one-line fix; a mild-sounding surprise whose assumption underpinned eight downstream steps needs a suffix rewrite. This is why an explicit assumption register pays for itself. If each step records which facts it depends on, the repair scope falls out mechanically: find the invalidated fact, find every step that depends on it, and rewrite exactly that set. ## The three scopes **Step-local repair.** The failing step's own preconditions broke; the rest of the plan still holds. You re-author one step — a different vendor, a different tool, a different argument, an added prerequisite step — and resume. This is by far the cheapest option: one small model call, no re-derivation of a plan that was fine, and the run's history stays coherent. *Plant-maintenance example:* step 4 was "order spare pump from vendor A"; vendor A shows zero stock. Vendor B has it at a higher price. Steps 5–9 (schedule window, dispatch technician, swap, test, close ticket) never referenced the vendor. Swap step 4 and carry on. **Suffix replan.** The broken assumption was consumed by later steps, so everything from the failure point onward must be re-derived, while the already-executed prefix stands. This is the middle option and the most commonly correct one in practice. *Same example, different break:* the outage window moved from Saturday 02:00 to the following Thursday. Step 4 is untouched — the pump still needs ordering — but the technician booking, the parts-delivery deadline, the customer notification and the ticket SLA were all derived from Saturday. Keep the prefix, rewrite the tail against the new window. **Full replan.** The goal's feasibility, the top-level constraints, or the strategy itself changed. The pump is discontinued and no drop-in exists; the site cannot be taken offline this quarter; the budget was withdrawn. Here the earlier decomposition rests on premises that no longer hold, and patching it produces an incoherent plan that reads fluent and is unexecutable. A full replan is also the honest point to ask whether the goal is still the right goal — frequently the correct output is an escalation carrying the evidence, not a new plan. ## What bounds the choice **Cost and latency.** Every replan is at minimum one planning call over a context that has grown with every observation so far. Full replans on long runs are the expensive operation, and they throw away reasoning that was still sound. Prefer the narrowest scope that restores validity — but do not under-scope out of thrift, because a locally-patched plan whose downstream steps rest on a dead assumption fails again three steps later, having spent real side effects on the way. **Irreversibility.** Replanning rewrites *future* steps only. It does not un-order the part, un-send the email, or un-dispatch the technician. The executed prefix is a fact, and the new plan must treat committed side effects as part of the current world state — sometimes adding compensating steps (cancel the booking, return the part) that the original plan never contained. This is the single most common gap in candidate answers: they describe replanning as if the plan were pure computation. **Context growth.** By the time a full replan fires on a long run, the planner's context holds every prior observation. Feeding all of it back produces slow, expensive, and often worse plans. Give the replanner a condensed world state — the goal verbatim, the current facts, the committed side effects, the deviation evidence — rather than the raw transcript. ## A workable decision procedure 1. Identify the assumption the observation invalidated. 2. Compute the dependent set: which steps consumed that assumption? 3. If the set is empty beyond the failing step, repair that step in place. 4. If the set is a downstream suffix, re-derive from the earliest dependent step, holding the executed prefix and its side effects fixed. 5. If the invalidated assumption is a goal-level premise or a hard constraint, do not patch — replan wholly, and consider escalating with the evidence attached instead. ## Where candidates go wrong The two symmetric errors are *always patching locally* (cheap, and it leaves the plan quietly invalid) and *always replanning from scratch* (safe-feeling, and it burns tokens, loses valid reasoning, and on long horizons opens the door to drifting away from the original goal). A senior answer names both, ties the choice to dependency structure rather than to how bad the error looked, and remembers that side effects already committed cannot be replanned away.
- What do you feed a full replanner on a long run — the whole transcript?No. By then the transcript holds every observation, which makes planning slow, costly and often worse. Feed a condensed world state: the goal restated verbatim, the current known facts, the side effects already committed, the executed prefix, and the deviation evidence. The planner needs the state of the world, not the history of how you learned it.
- How do irreversible actions in the executed prefix change the new plan?They become fixed facts the plan must accommodate, and sometimes require compensating steps the original plan never had — cancel the technician booking, return the ordered part, retract the customer notification. A replanner that treats the world as if the prefix never ran will produce a plan that double-books or double-orders.
- Why is defaulting to a full replan on every deviation a bad policy?It discards reasoning that was still valid, costs a planning call over an ever-growing context, and adds latency on every wobble. Worse, on long horizons each fresh rewrite is another opportunity to drift from the original goal. Full replan is the right tool for an invalidated premise, not the default reflex.
saying these in an interview costs you the question
- Choosing repair scope by how severe the error looked
- Always replanning from scratch because it feels safer
- Patching the failing step while downstream steps rest on the dead assumption
- Assuming replanning undoes actions already taken
- Feeding the entire run transcript back into the replanner