How many reflection rounds should an agent run before it stops revising?
answer
- Cap is a backstop, not the rule
- Stop when the checker passes
- No progress means stop
- Round one does most of the work
- Keep the best artifact, not the last
basics
~20 sMost of the gain arrives in the first one or two rounds, so cap revision at two or three and stop earlier on signal: the verifier passes, or a round fails to reduce the failure count. Every extra round costs a full generation and its latency.
solid answer
~50 sTreat the round cap as a safety net, not the primary stopping rule. The primary rule should be signal-driven: stop when the verifier passes, and stop when a round produces **no measurable improvement** — the same tests still fail, the critic returns the same findings, or the diff churns text without addressing anything. Empirically the curve is steep then flat: round one fixes the obvious defects, round two catches a few more, and rounds three and beyond mostly rewrite. Without a verifier the curve is flatter still, and often noisy in both directions, which argues for a cap of one. Set the cap per task class rather than globally — a patch loop against a test suite justifies more rounds than a summary. Log rounds used and outcome per round, so the cap is a measured number rather than a guess, and account for the fact that each round adds full generation latency to the user-visible response.
go deeper
Know that revision loops need a limit, that a couple of rounds captures most of the benefit, and that each round costs another full model call.
Explain the progress-based rules alongside the cap: stop when the checker passes, when the failure count does not drop, or when the critique repeats itself.
Show the production instincts — per-task-class caps chosen from measured data, retaining the best-scoring artifact rather than the last, honest reporting when the cap is exhausted, and the latency cost of each round.
Own the quality-versus-time-to-answer tradeoff as a product decision, and the cost-per-successful-task arithmetic that decides whether a revision tier belongs in the interactive path at all.
## Why a cap is not enough on its own A fixed cap answers "when must we stop?" but not "when should we?". Running three rounds on a draft that was correct after one wastes two full generations and their latency; running three on a draft that is not converging wastes the same and ships a churned artifact. Good loops carry both a cap and a progress rule, and the progress rule fires far more often. ## Signal-driven stopping rules **Verifier satisfied.** The strongest rule: tests pass, the schema validates, the compiler is clean. Stop immediately — further rounds can only regress a passing artifact, and revising something that already passes is a real source of regressions. **No improvement across a round.** Compare the verifier state before and after. If the failing-test count did not drop, or the same schema violation persists, the revision did not act on the feedback and another round is unlikely to. Some loops allow one "no progress" round before stopping, on the grounds that a repair can be partial; two in a row is decisive. **Critique convergence.** If the critic returns substantially the same findings it returned last round, the generator is not able to address them and the loop is spinning. This also catches the case where the critique is unactionable — vague findings produce cosmetic revisions that never clear them. **Regression detection.** Keep the best artifact seen so far, scored by the verifier. If a round makes things worse, return the retained best rather than the latest. Loops that always return the final artifact silently ship regressions. ## Why the returns fall off so fast The first round consumes the highest-value information: the largest, most localized defects with the clearest error messages. What remains after that is either subtler, harder to localize, or not actually a defect. Meanwhile each round adds the previous draft and critique to the context, which lengthens it and increases the chance of the model latching onto its earlier commitments. So the marginal value of information falls while the marginal cost and the drift risk rise. That is the general shape whether or not a verifier is present; a verifier simply makes the first round much more valuable and makes the falloff detectable. ## Setting the number Do not pick one number for the whole system. Reasonable starting points: - **Ungrounded critique** (summaries, advisory text): cap at one, and be prepared to remove the round entirely if measurement does not justify it. - **Structural validation** (schema, format): cap at two — a schema violation almost always clears on the first repair, and a second failure usually means the schema or the instruction is wrong, not the output. - **Executable verification** (tests, type checks): three to five is defensible, because each round has a real oracle and a genuine chance of progress. Cost per round is high, so pair it with the no-progress rule. Then measure. Instrument rounds-used and per-round outcome across your task suite, and plot the success rate against the cap. If almost no task succeeds on round four, the cap is three. ## Budgets, latency, and user experience A revision round is a full generation and, with a verifier, a tool execution as well. Under an interactive latency budget, two rounds may already be unacceptable, which pushes the loop offline or behind a streamed progress indication. The honest tradeoff is quality against time-to-answer, and it should be an explicit product decision rather than a default buried in the loop. When the cap is reached without success, the right behaviour is usually to return the best artifact seen along with an explicit note that verification did not pass — not to present the last revision as if it were finished. False success claims after an exhausted loop are worse than an honest failure, because downstream systems and users act on them. ## What to log At minimum, per task: rounds used, verifier state after each round, whether the final artifact was the last or a retained earlier one, and total tokens spent on revision. Those four fields are enough to tune the cap, catch loops that never converge, and prove to a sceptic that the revision tier earns its cost.
- What should the agent return when the cap is reached and verification still fails?The best artifact seen so far, scored by the verifier, together with an explicit statement that it did not pass and which checks failed. Presenting the final revision as complete is a false success claim, and downstream systems act on it. Returning the retained best also avoids shipping a regression introduced by the last round.
- Why keep the best artifact rather than the latest one?Because revision is not monotonic — a round can make things worse, especially once the earlier drafts and critiques crowd the context. Scoring each candidate with the verifier and retaining the best turns the loop into a search with a floor, so the outcome is never worse than the strongest attempt already produced.
- How would you justify the specific cap you chose to a reviewer?With data from the task suite: rounds used per task, verifier state after each round, and success rate plotted against the cap. If the success curve flattens after round two, the cap is two or three. Choosing the number by measurement also gives the cost-per-successful-task figure that justifies the revision tier at all.
saying these in an interview costs you the question
- Runs a fixed number of rounds with no progress check
- Returns the last revision rather than the best-scoring one
- Assumes more rounds monotonically improve quality
- Uses one global cap for every task class
- Reports success when the cap was exhausted without passing