skip to content

In recursive RAG retrieval, each sub-answer forms the next query — what stops the chain?

level: seniorimportance: should knowfreq 34%

answer

  1. nothing terminates a loop by default
  2. hops are sequential, errors compound
  3. hard cap outside the model's control
  4. no new chunks means no progress
  5. say so when you stop early

basics

~20 s

Nothing stops it by construction, so you impose stops: a hard depth cap, a check that the accumulated evidence answers the original question, and a novelty test that halts when a hop returns only chunks already in context. Budget exhaustion is the backstop.

solid answer

~50 s

Recursive retrieval is for dependent chains — "which carrier serves our Rotterdam lane, and what is that carrier's claims SLA?" cannot retrieve the SLA until hop one names the carrier. Each hop feeds its answer into the next query, so the loop is unbounded by default and every hop compounds error: a wrong carrier name sends hop two confidently into the wrong documents. Practical stopping rules stack. A **depth cap** (usually two or three hops) is the non-negotiable hard stop, because it bounds latency and cost regardless of what the model believes. A **sufficiency check** asks whether the accumulated sub-answers now answer the original question. A **novelty check** halts when the newest hop retrieves only chunks already gathered — the strongest no-progress signal available. Never rely on the model declaring itself finished as the only stop; that judgement is exactly what fails on the queries that loop.

go deeper

for a junior

Know that some questions need a second search that depends on the first search's answer, and that such loops need an explicit limit on how many rounds they run.

for a middle

Explain the stopping rules and why they stack: a hard depth cap, a sufficiency check on the accumulated answers, and a novelty check that halts when a hop returns nothing new.

for a senior

Demonstrate the operational view — compounding error across hops, mechanical stops kept outside the model's control, logging which stop reason fired, and reporting an unresolved sub-question instead of a silently truncated answer.

for a principal

Own the tradeoff between depth and reliability: decide the depth budget from measured question distributions and latency targets, and recognize when a recursive chain should be replaced by a purpose-built lookup or a structured query path.

## When a chain is unavoidable Most decomposition produces independent sub-questions you can retrieve for in parallel. Recursive — or iterative — retrieval is the case where you cannot: the next query is not knowable until the previous sub-answer exists. A logistics assistant asked "what is the claims SLA on our Rotterdam lane?" must first retrieve which carrier holds that lane, then use the carrier's name to retrieve the contract clause. Two searches, strictly ordered. This is powerful and expensive. Each hop costs a retrieval plus at least one model call, and hops are sequential, so latency grows linearly with depth. Worse, quality degrades multiplicatively: if each hop is ninety percent reliable, three hops land around seventy percent, and the failure is silent — a wrong intermediate value produces a perfectly confident retrieval from entirely the wrong part of the corpus. ## Why a stop condition must be imposed Nothing in the loop's structure terminates it. Given any accumulated context, a model asked "what should we look up next?" will nearly always produce a plausible next question. Left to itself the chain wanders: it drifts from the original question, chases tangents that seem informative, and can oscillate between two related queries indefinitely. Termination is a design decision, not an emergent property. ## The stopping rules that work **Depth cap.** A fixed maximum number of hops, typically two or three for retrieval chains. This is the only rule that holds when every other signal is compromised, because it does not consult the model at all. Set it from your latency budget, not from ambition. Most real question sets have a shallow tail: a small fraction of questions truly need three hops, and questions needing five are usually questions you should have answered differently. **Sufficiency check.** After each hop, evaluate whether the accumulated sub-answers now answer the original question. This is a model call, so it costs, and it is imperfect — models err toward continuing when the material is thin and toward stopping when the context is long and comfortable-sounding. Treat it as the ordinary stop, not the safety net. **Novelty check.** Compare the chunk ids returned by the newest hop against everything already gathered. If the hop introduced no new chunks, the chain has stopped learning and further hops will only rephrase what you have. This is cheap, purely mechanical, and the single most reliable no-progress detector available, because it does not depend on any model's self-assessment. A near-variant is to compare the generated sub-query against previous ones and halt on repeats, which catches oscillation between two phrasings of the same lookup. **Budget exhaustion.** A wall-clock deadline or a token ceiling for the whole request, checked before starting each hop. When it trips, stop and answer from what you have, with an explicit note about what could not be resolved. These stack rather than compete. The typical loop runs: hop, novelty check, sufficiency check, depth and budget check, repeat. ## Never let the model be the only stop The attractive design is to let the model decide it is done. The problem is that the queries where the model's judgement is unreliable are precisely the queries that loop — ambiguous questions, thin corpora, and chains where an early hop returned something wrong but fluent. The model's confidence is highest exactly where it should be lowest. Mechanical stops must sit outside the model's control. ## What to do at the stop How you terminate matters as much as when. Three options: answer from the accumulated context and state which part is unresolved; return the partial chain to the user and ask a clarifying question; or fail explicitly. The one option to avoid is silently answering as though the chain had completed — the user has no way to tell a two-hop answer that resolved from a two-hop answer that ran out of budget halfway. Carrying the intermediate sub-answers into the final response as an audit trail is worth the tokens on any workflow where the answer will be acted on. When a recursive answer is wrong, the failing hop is almost always identifiable at a glance from that trail, and it is the only artifact that shows *which* hop went wrong rather than merely that the answer did. ## Instrumentation Log per request: hop count, which stopping rule fired, the generated query at each hop, and the distinct-chunk count per hop. The distribution of stop reasons is the health metric. If depth caps fire often, either your questions are deeper than your cap or the loop is drifting; if novelty checks fire on the first or second hop routinely, the chain is not earning its cost and those queries should be routed to plain retrieval instead.

  • Why is a novelty check on retrieved chunks more trustworthy than asking the model whether it has enough information?
    The novelty check is mechanical — it compares chunk ids against what is already gathered and cannot be talked into an optimistic answer. Model self-assessment fails hardest on exactly the queries that loop: ambiguous questions and chains where an early hop returned something fluent but wrong. Use the model's judgement as the normal stop and mechanical checks as the ones that cannot be overridden.
  • How would you keep an error in the first hop from poisoning the rest of the chain?
    Carry the intermediate value with its evidence rather than as a bare string, and require the next query to be grounded in a retrieved chunk, not in a value the model asserted. Where the intermediate is an entity that can be validated — a carrier code, an account id — check it against a source of truth before hop two. And keep the chain short: two hops fail far less often than four.
  • What should the system return when the depth cap or budget stops the chain early?
    Answer from what was gathered, and state plainly which part of the question remains unresolved — ideally with the sub-question that was still open. Silently answering as though the chain completed is the worst outcome, because the user cannot distinguish a resolved answer from a truncated one. On workflows where the answer will be acted on, return the intermediate sub-answers as an audit trail too.

saying these in an interview costs you the question

  • Letting the model decide alone when to stop retrieving
  • Assuming the loop will terminate naturally
  • Running deep chains without a wall-clock or token budget
  • Treating each hop as independently reliable
  • Answering from a truncated chain without saying so

context