skip to content

A ReAct agent repeats the same two searches until its step cap fires — why?

level: seniorimportance: should knowfreq 48%

answer

  1. uninformative feedback reproduces the decision
  2. nothing in the loop scores progress
  3. A fails so B looks good, and back
  4. ask what each observation added
  5. the repeat is telling you something is unachievable

basics

~20 s

Because each pass re-derives the next action from a context that barely changed. When an observation adds nothing usable, the state that produced the last action is effectively reproduced, so the same action is chosen again, and the loop has no built-in notion of progress to notice it.

solid answer

~50 s

Oscillation is what a memoryless-in-effect decision rule does when feedback carries no new information. Each pass picks the action that looks best given the transcript; if the observation was empty, identical, or an error the agent cannot interpret, the transcript is materially unchanged, so the same reasoning recurs and the same action is emitted. Two actions can alternate for the same reason: each one's failure makes the other look attractive again. Nothing in the bare loop tracks whether the run is closer to the goal, so it spins until an external cap fires. The fixes are to give the loop a progress signal: hash recent action-and-argument pairs and detect repeats, inject an explicit note that this exact call already returned that result, require the next thought to state what new information it has, force a different tool or a broadened query after N repeats, then escalate or return a partial result.

go deeper

for a junior

Recognise the pattern in a trace: the same action with the same arguments appearing repeatedly, with observations that keep returning the same nothing.

for a middle

Explain the mechanism — each pass re-derives the action from a context that an uninformative observation left materially unchanged — and name repeat detection over action-and-argument pairs as the first countermeasure.

for a senior

Diagnose from the trace by asking what each observation actually added, trace the cause upstream to an unsatisfiable query or a broken tool, and choose deliberately between injecting a repeat notice, constraining the next action, and stopping with a partial result.

for a principal

Own the policy across a fleet: what fraction of runs may oscillate before it is a product defect, whether budget is spent on detection or on tool and data quality upstream, and how repeated-action rate is surfaced as an operational metric.

## The mechanism, not the symptom An oscillating agent is not confused in the human sense. It is doing exactly what the loop asks: choose the best next action given the transcript. The pathology comes from the interaction of two facts. **Fact one: the decision is a function of the context.** Pass N's action is generated from the accumulated transcript. If two passes present near-identical evidence about the goal, they will tend to produce near-identical actions — the second call is a rational choice from a state that has not usefully changed. **Fact two: an unhelpful observation does not change the evidence.** An empty result set, a timeout, a permission error, or a result identical to the previous one adds text to the transcript but adds nothing about the goal. The agent is where it was, and the action that looked best then still looks best now. Put together: **the loop converges on repetition precisely when feedback stops being informative.** Alternation between two actions is the same effect one step wider — A returns nothing, so B looks better; B returns nothing, so A looks better again, and the pair cycles until something external stops it. ## Why the agent cannot see it from inside The transcript does contain the earlier attempts, so people reasonably ask why the agent does not simply notice. Three reasons, in practice: - **Nothing asks it to.** The loop's implicit question each pass is "what next?", not "am I making progress?". Progress is not a term in the decision. - **Repetition reads as consistency.** Having just written a rationale for an action, the same rationale remains the most locally plausible continuation. Coherence with its own recent output pulls toward repeating rather than abandoning. - **Attention to a long, repetitive middle is weak.** In a transcript stacked with near-identical passes, the older attempts are exactly the material least likely to steer the next token — and they all say the same thing anyway. ## Diagnosing it in a real trace Read the trace and answer one question: **what did each observation add?** In a genuine oscillation the honest answer for several consecutive passes is "nothing about the goal". That distinguishes it from slow-but-real progress, where each pass narrows something even if the run is long. A second tell is that the thoughts stop referring to specifics from the observations and start restating the goal. The root cause is usually upstream of the loop: a query the corpus cannot satisfy, a filter combination with no matching rows, a tool that returns an error the agent has no route to fix, or a goal that was never achievable with the tools provided. Oscillation is the loop's way of showing you an impossible sub-goal. ## Fixes, cheapest first **Detect the repeat.** Keep a set of recent action-plus-normalised-arguments pairs. On a hit, you know before spending another call. This is mechanical and costs nothing. **Tell the agent what it already did.** On detection, append an explicit note: this exact call was already made and returned this result, so a different approach is required. Surfacing the repeat as a fact usually breaks the cycle where merely having it buried in history does not. **Make progress explicit.** Require each thought to state what new information the last observation provided. An agent forced to write "nothing new" is far more likely to change strategy than one that never has to assess it. **Constrain the next action.** After N repeats, exclude the repeated action for a pass, or force a broadened or differently shaped query. This is a harness-side restriction rather than a request. **Give up well.** If the repeats continue, the sub-goal is probably unachievable with the available tools. Escalate to a human or return what was established with an explicit incomplete marker. A cap that fires with no explanation is a much worse outcome than a run that says "the catalogue has no product matching these filters". ## What separates a senior answer Weak answers say the agent is stuck and propose raising the cap — which funds more of the same loop. Strong answers explain *why* an uninformative observation reproduces the previous decision, distinguish oscillation from legitimate slow progress by asking what each observation added, and treat the repeated action as a signal that some sub-goal is unsatisfiable rather than as a bug in the agent's willpower.

  • How do you tell oscillation apart from a run that is simply long?
    Ask what each observation contributed toward the goal. Slow progress still narrows something every pass — a filter tightened, a candidate eliminated, a fact confirmed. Oscillation produces several consecutive passes whose honest contribution is nothing, and whose thoughts drift from citing observation specifics to restating the goal.
  • Why is raising the step cap the wrong first response?
    Because the cap is not the cause; it is the only thing currently ending the run. If feedback has stopped being informative, more passes reproduce the same decision at higher cost, and the failure surfaces later and more expensively. Fix the information problem, or stop early with a partial result, then set the cap from observed successful-run lengths.
  • What does a repeated identical action usually indicate about the task itself?
    That some sub-goal is not satisfiable with the tools and data provided — no matching rows, a corpus gap, or a permission the agent cannot obtain. The agent has no way to conclude "this cannot be done" unless something asks it to, so it keeps proposing the only action that looked plausible. Treat it as a signal about the environment, not a defect of nerve.

It is a person re-checking an empty mailbox: each trip is individually reasonable, and nothing in the routine ever raises the question of whether the letter is coming at all.

saying these in an interview costs you the question

  • Raising the step cap as the primary fix
  • Assumes the agent must notice repeats itself
  • Blames the model rather than uninformative feedback
  • Confuses oscillation with legitimate slow progress
  • Never considers that the sub-goal may be impossible

context