How do you stop an agentic RAG retrieval loop from running away?
answer
- success condition, not just a limit
- same documents back means stop
- cap rounds, clock and spend
- set caps from p95, not intuition
- exhaustion must not mean answer anyway
basics
~20 sGive the loop an explicit sufficiency test and hard limits: stop when a grader says the gathered evidence answers the question, stop when a reformulated query returns documents already seen, and cap iterations, wall-clock time and spend so the worst case returns a partial answer rather than an unbounded bill.
solid answer
~60 sAn agentic RAG loop exposes retrieval as a tool the model may call repeatedly, grading results and reformulating between calls. Without stop conditions it can search forever, because the model's default behaviour when evidence is thin is to try again. Three kinds of stop belong in the design. A **success stop**: a sufficiency judgement over the accumulated evidence — can the question be answered from what we now hold — evaluated after each round rather than only at the end. A **no-progress stop**: track the document identifiers returned each round and halt when a reformulation returns the same set, since another rewording of a query the corpus cannot serve will not help. A **budget stop**: a maximum number of rounds, a wall-clock deadline and a spend cap, chosen from the p95 rounds observed in evaluation rather than guessed. What matters as much is the exit behaviour. On exhaustion, return the best evidence found with an explicit statement of what is missing. A loop that silently answers from insufficient context has converted a latency problem into a correctness problem.
code
python · 13 linesdef agentic_retrieve(question, retrieve, grade, rewrite, max_rounds=4):
query, seen, best = question, set(), []
for _ in range(max_rounds):
docs = retrieve(query)
ids = frozenset(d["id"] for d in docs)
best = docs
if grade(question, docs) >= 0.7:
return {"docs": docs, "stop": "sufficient"}
if ids in seen:
return {"docs": docs, "stop": "no_progress"}
seen.add(ids)
query = rewrite(question, docs)
return {"docs": best, "stop": "budget_exhausted"}go deeper
Know that an agentic RAG system can call retrieval more than once, and that it needs an explicit limit so it cannot keep searching forever.
Explain the three stop types — sufficiency reached, no new documents, budget exhausted — and why repeated identical result sets mean the corpus, not the query, is the problem.
Demonstrate operating judgement: derive caps from the observed round distribution, define exhaustion behaviour that reports the gap rather than answering anyway, and instrument termination reason alongside quality.
Own the economics and the blast radius. Per-question cost becomes variable, so decide the spend ceiling per query class, how the system degrades under a correlated traffic spike, and what an unanswerable question should cost before it becomes a corpus-coverage work item.
## What the loop is In agentic RAG the retriever is not a fixed pipeline stage; it is a tool. The model decides whether to call it, with what query, reads what comes back, judges it, and may call again with a reformulated query or against a different source. That is powerful — multi-hop questions become expressible, because the second query can depend on the first result — and it is unbounded by construction. The pathology is easy to reproduce. Ask a question the corpus cannot answer. The model retrieves, judges the evidence insufficient (correctly), rewords, retrieves near-identical passages, judges them insufficient again, and repeats until something external stops it. Every round costs at least one generation call plus a retrieval, so cost and latency grow linearly while information gained approaches zero. ## A worked case An on-call SRE assistant answers incident questions from an internal runbook corpus. During an incident someone asks why a service's health checks began failing after a deploy. The runbooks were written against the previous deploy topology, so every retrieval returns passages about a component layout that no longer exists. The grader is right to reject them. The agent rewords — "health check failure", "readiness probe timeout", "post-deploy rollout failure" — and keeps getting the same stale pages, burning minutes during an incident, which is the worst possible time to spend latency on a search that cannot succeed. The correct behaviour is to detect that the corpus is not going to answer, stop, and either escalate to an external source or hand back what was found with a plain statement that the runbooks predate the current topology. That answer, delivered in fifteen seconds, is far more useful during an incident than a perfect answer that never arrives. ## Success conditions The primary stop should be positive, not merely a limit. After each round, evaluate the accumulated evidence against the question: is every part of the question now covered by retrieved material? Decomposing the question into required facts up front makes this checkable — the loop stops when each required fact has supporting evidence, rather than when a model feels done. A pure "do you have enough?" self-judgement works, but it is the same model that decided to keep going, so it tends toward optimism early and pessimism late. ## No-progress detection The cheapest and most reliable safeguard costs nothing in tokens: hash the set of document identifiers returned each round. If round three returns the same set as round two, the reformulation did not reach new material and further rewording is unlikely to. A softer variant tracks the fraction of newly seen documents per round and stops when it falls below a threshold, which also catches slow oscillation between two phrasings. The same logic applies to the queries themselves: keep the reformulations you have already tried, both to detect repeats and to feed them to the model so it does not propose one again. ## Budgets Hard limits are the backstop for everything the smart conditions miss: maximum rounds, wall-clock deadline, and a cost ceiling per question. Set them from data. Run your evaluation set, record how many rounds successful answers actually needed, and place the cap above the p95 — a cap below it silently truncates the questions the loop exists to serve. In interactive settings the wall-clock deadline usually binds first and should be the one derived from the user experience. Budgets also need to hold across concurrency. One runaway question is a rounding error; a retry storm during an incident, when many people ask overlapping questions about the same broken system, is a bill and a rate-limit event. ## Exit behaviour is part of the design What happens at exhaustion decides whether the caps make the system safer or merely cheaper. Three defensible exits: return the best evidence gathered with an explicit statement of the gap; escalate to a fallback source or a human; or refuse. The indefensible exit is to answer anyway from whatever is in context, because that turns a bounded latency cost into an unbounded correctness cost, and it does so precisely on the questions the corpus could not support — the ones where being wrong is most likely. Surface the loop's state as well. An agent that reports "searched four times, found only pre-migration runbooks" gives the on-call engineer something actionable; one that reports a confident wrong answer gives them a second incident. ## What to instrument Rounds per question as a distribution, not a mean; the share of questions terminating on each condition (success, no-progress, cap); cost and latency per answered question; and answer quality split by termination reason. If a large share of questions terminate at the cap, either the cap is too low or the corpus has a coverage hole — and those two look identical on a dashboard that only tracks averages.
- How would you choose the maximum number of retrieval rounds?Empirically. Run the evaluation set with a generous cap and record how many rounds the questions that eventually succeed actually consumed. Set the cap above the p95 of that distribution so you are not truncating the multi-hop questions the loop exists for, then check the cost implication at your traffic volume. Revisit it whenever the corpus or the question mix changes, since both shift the distribution.
- The agent keeps reformulating and getting the same passages back. What is the underlying problem, and does a better rewriter fix it?Usually the corpus simply lacks the answer, and no rewording reaches material that is not indexed. A better rewriter helps only when the failure is vocabulary mismatch, which the first or second reformulation would already have resolved. Treat repeated identical result sets as evidence of a coverage gap, log the question for corpus work, and route to a fallback source or a human instead of paying for more rounds.
- What is the risk of letting the model itself decide it has enough evidence?It is the same model that chose to keep searching, so the judgement is not independent, and it is systematically miscalibrated: optimistic when a plausible-looking passage arrives, and prone to continuing when it has enough but the question feels hard. Decomposing the question into required facts up front and checking coverage of each gives a more mechanical stop, and an external grader gives an independent one.
saying these in an interview costs you the question
- Relying only on a max-iteration cap with no success condition
- Answering from insufficient context when the budget runs out
- Assuming another reformulation will eventually reach missing documents
- Setting caps by intuition instead of the observed round distribution
- Ignoring that each round costs a generation call as well as a retrieval