How do you detect an agent loop that keeps repeating steps without making progress?
answer
- a cap is not a detector
- same call, same args, same result
- hash a window, not just one step
- normalise volatile fields first
- nudge once, then stop
basics
~20 sHash each step — tool name, arguments and the observation it returned — and count repeats. When the same hash recurs a few times, or a window of recent hashes cycles, the agent is stuck. Trip a no-progress stop instead of waiting for the iteration cap.
solid answer
~60 sWaiting for the iteration cap to catch a stuck agent is expensive and slow, so you add a cheap detector. The core signal is **repeated state**: canonicalise each step into a hash over the tool name, its normalised arguments, and the observation returned, then fire when the same hash appears N times — three is a common threshold. A single-step hash catches an agent hammering the same call; hashing a sliding *window* of steps catches oscillation, like a game-QA agent bouncing between door A and door B forever, where no individual step repeats consecutively. Secondary signals: the scratchpad or working artifact has not changed in several steps, observations carry no new information, or a task-specific progress metric (tests passing, items resolved) is flat. On detection you have three options — inject an explicit note telling the model it has already tried this call with these arguments and asking for a different approach, escalate to a human, or stop and return what exists. Do the nudge at most once or twice; a model that ignores it will keep ignoring it.
code
python · 22 linesimport hashlib
from collections import Counter, deque
VOLATILE = {"request_id", "timestamp", "cursor"}
class NoProgressDetector:
def __init__(self, repeat_limit=3, window=6):
self.counts, self.window, self.limit = Counter(), deque(maxlen=window), repeat_limit
def observe(self, tool, args, observation):
clean = sorted((k, v) for k, v in args.items() if k not in VOLATILE)
key = hashlib.sha256(f"{tool}|{clean}|{observation}".encode()).hexdigest()[:16]
self.counts[key] += 1
self.window.append(key)
repeated = self.counts[key] >= self.limit
cycling = len(self.window) == self.window.maxlen and len(set(self.window)) <= 2
return repeated or cycling
d = NoProgressDetector()
stuck = any(d.observe("open_door", {"door": door, "request_id": i}, "room unchanged")
for i, door in enumerate(["A", "B"] * 4))
print(stuck)go deeper
Know that an agent can get stuck repeating the same tool call, and that harnesses watch for identical repeated calls rather than trusting the model to notice.
Explain the hashing mechanism: canonicalise tool name, arguments and observation, count repeats, fire at a threshold. Mention that a sliding window is what catches alternating states.
Show operational judgement — normalising volatile arguments, choosing a threshold against real traces, nudging once with the concrete repetition named, and reading no-progress metrics per tool to find the ambiguous tool description behind the loop.
Frame it as a cost-per-progress control: what you alert on across a fleet, how detector thresholds trade false stops against wasted spend, and why fixing tool descriptions and error messages removes more loops than tuning detectors does.
## Why a cap is not a detector An iteration cap will eventually stop a stuck agent, but it stops it at maximum cost: forty model calls, forty tool executions, a minute of latency, and an answer that is no better than it was at step six. A no-progress detector aims to notice within a handful of steps that the loop has stopped generating new information, so you can spend the remaining budget on something else — or return early and honestly. ## The primary signal: repeated state The cheapest reliable detector is a hash. After each step, canonicalise a small tuple: the tool name, its arguments with keys sorted and volatile fields (timestamps, request ids) normalised away, and a digest of the observation. Hash it, count it, and trip when a count crosses a threshold. Two shapes matter: - **Repetition.** The same call with the same arguments returning the same result. Threshold of three is a sensible default; two can fire on legitimate retries after a transient error. - **Oscillation.** The agent alternates between two or more states, so no hash repeats *consecutively* but the set of recent hashes cycles. An automated game-QA agent exploring a level shows this beautifully: it opens door A, sees the same room, opens door B, sees the same room, opens door A again. Detect it by hashing a sliding window of the last k steps, or by watching for a small set of hashes that dominate a recent window. Canonicalisation is where implementations quietly fail. If arguments include a fresh UUID or a current timestamp, every step hashes differently and the detector never fires even though the agent is plainly stuck. Normalise before hashing, and prefer a digest of the *semantic* part of the observation over the raw bytes. ## Secondary signals Hashes catch identical work. Progress is a broader idea, and a few complementary signals are worth having: - **Unchanged artifact.** If the agent's job is to build something — a scratchpad file, a todo list, a patch — and it has not changed across several steps, no work is landing regardless of how varied the tool calls look. - **No new information.** Diff each observation against everything already in context. An agent re-reading files it has already read is burning budget to learn nothing. - **Flat task metric.** Where the domain gives you a cheap score (tests passing, lint errors remaining, records reconciled, bugs found), plateau on that number is the most meaningful stall signal you have, because it is defined in terms of the goal rather than in terms of activity. - **Rising cost per unit of progress.** Tokens spent since the last measurable improvement is a good single number to alert on. ## What to do on detection Detection is easy; the response is the judgement call. 1. **Nudge once.** Inject an explicit observation into the context: "You have called `open_door` with `{door: A}` three times and received the same result each time. That approach is not working — choose a different action or report what you have." Being concrete about the repetition works far better than a generic "try something else", because the model often genuinely does not notice the pattern in a long transcript. 2. **Escalate.** If the run is user-facing and the agent has authority over anything consequential, hand it to a human with the loop's trace attached. 3. **Stop.** Return the partial result with a status that says why. This is usually right after one or two failed nudges: a model that oscillated through a nudge is unlikely to escape on the third. Avoid an unbounded nudge budget — each nudge costs a full model call and adds tokens to a context that is already crowded, and repeated nudges are themselves a loop. ## Instrument it Emit the detection as a metric with the repeated tool name attached. In aggregate this is one of the most actionable agent signals you have: a single tool dominating no-progress events almost always means its description is ambiguous, its errors are unhelpful ("an error occurred" gives the model nothing to change), or its result does not actually answer the question the agent is asking. Fixing the tool removes the loop far more reliably than tuning thresholds does. ## Threshold honesty There is no universal N. Retry-heavy environments need higher thresholds; expensive tools justify lower ones. Tune it against real traces rather than by intuition, and accept that both directions cost you — too low and you cut off legitimate retries, too high and you pay for the stall you were trying to avoid.
- Your hash-based detector never fires even though the agent is visibly stuck. What is the likely cause?Canonicalisation. If arguments carry a timestamp, a request id, or a cursor that changes every call, every step hashes differently and the counter never reaches its threshold. Strip volatile fields, sort keys, and hash a semantic digest of the observation rather than raw bytes.
- How do you catch an agent that alternates between two actions instead of repeating one?Consecutive-repeat detection misses oscillation by construction. Hash a sliding window of the last k steps, or track the distribution of hashes in a recent window and fire when a small set dominates it. A two-cycle then shows up as two hashes covering nearly every recent step.
- Why is a flat task metric a better stall signal than a flat step count?Because it is defined in terms of the goal. An agent can issue perfectly varied, never-repeating tool calls and still make no progress. Tests passing, bugs found, or records reconciled staying flat over several steps says the work is not advancing, whatever the activity looks like.
saying these in an interview costs you the question
- Relies on the iteration cap alone to catch a stuck agent
- Hashes raw arguments including timestamps or request ids
- Only checks consecutive repeats, so oscillation slips through
- Nudges the model on every detection with no bound
- Treats varied tool calls as proof of progress