How do you tell that a self-refine loop has stopped improving output and is just changing it?
answer
- more rounds is not more better
- keep the best, not the last
- score with something the critic doesn't control
- diffs shrink to synonym swapping
- critique repeats, then reverses itself
basics
~20 sScore every round against a signal outside the critic rather than trusting the critic's own satisfaction, keep the best-scoring version instead of the last one, and stop when the diffs become paraphrase-level churn or the critique starts repeating and contradicting itself.
solid answer
~40 sThe failure to avoid is treating the final round as the best round. Critique loops do not converge monotonically: a marketing email rewritten five times often peaks at round two and then degrades, because once the substantive problems are gone the critic still has to say something and starts inventing stylistic objections. So instrument the loop. Score each candidate against something the critic does not control — a constraint check, a held-out rubric, a judge that never saw the critique thread — and retain the argmax, not the last output. Watch the critique stream too: repeated objections, an objection that reverses a change from two rounds ago, or edits that only reword are all signals that the loop has exhausted its useful range. Iteration caps are a backstop for cost, not a quality criterion.
go deeper
Know that running more critique rounds does not automatically make output better, and that the last version is not necessarily the best one produced.
Explain the mechanism: a critic prompted to criticize will always find something, and each rewrite regenerates the whole artifact, so late rounds trade real content for hedged fluency.
Demonstrate instrumentation — an external scorer the critic does not control, argmax retention across candidates, and churn heuristics like shrinking diffs, repeated objections and oscillating edits as stop signals.
Own the budget policy: how many rounds each task class gets, whether refinement is worth running at all there, and what it means operationally when a class routinely hits the cap instead of converging.
## Refinement is not monotone The intuitive model of a critique loop — each pass makes the output a bit better, so more passes are better — is wrong, and interviewers ask about it precisely because the wrong model is so intuitive. Two effects break it. First, **the critic must produce output**. Ask a model to critique an already-good artifact and it will not say "this is fine"; it will find something, because the prompt asked for criticism. Once the substantive defects are gone the objections become stylistic, subjective, or invented, and the rewrite that honours them makes the artifact worse. Second, **each rewrite is a full regeneration risk**. A rewrite is not a surgical patch; the model reproduces the whole artifact and can drop a good line, blunt a strong opening or reintroduce a problem that an earlier round fixed. The longer the loop runs, the more chances there are to lose something that was working. The marketing-email case is the canonical one: round one fixes a genuine structural problem, round two tightens it into the best version, and rounds three through five sand off the specifics until the copy is fluent, hedged and forgettable. If the loop returns "the last one", it returns round five. ## Keep the best, not the last This single change removes most of the damage. Treat the loop as a search that emits candidates rather than a pipeline that transforms one. Score each candidate, keep the argmax, and let extra rounds be cheap exploration instead of a risk of regression. It also makes the loop safe to run a round longer than necessary, because a bad round costs tokens but not quality. The scoring signal has to come from somewhere the critic does not control. Options in rough order of trustworthiness: a mechanical check of the properties that matter (length, required elements, forbidden claims, factual fields matching a source); a rubric grader that sees the candidate alone with no critique history; a human on a sample. A critic that has been arguing for its own edits across five rounds is the worst possible judge of whether those edits helped, because it is scoring its own advocacy. ## Signals that the loop is finished Several cheap heuristics detect exhaustion without needing a scorer at all: - **Diff size and kind.** When successive versions differ only by synonym substitution and clause reordering, the substantive work is over. - **Repeating critique.** The critic raising the same objection it raised two rounds ago means the actor is not able to satisfy it, or the objection is not real. - **Oscillation.** Round three removes a sentence, round five puts it back. That is a sign there is no shared objective function, just successive opinions. - **Softening verdicts.** Critique that has moved from concrete defects to "could be more engaging" has run out of findings. - **Score plateau or decline.** With an external scorer, two consecutive rounds without improvement is a reasonable stop. ## Where caps fit A hard iteration cap belongs in the loop, but it is a budget control, not a quality criterion — it bounds the worst case for cost and latency. Stopping on evidence of diminishing returns usually fires well before the cap on easy tasks, and the cap catches the pathological ones. Reporting which of the two ended the loop is useful telemetry: a task class that routinely hits the cap is a task class where refinement is not working and something structural needs to change. ## Refinement can also mask a wrong approach The most expensive version of this failure is polishing an artifact that should have been thrown away. Five rounds of critique on an itinerary built from a misread requirement produce a beautifully refined wrong answer. That is why the strongest loops check the coarse thing first — does this satisfy the hard constraints, does it answer the question asked — before spending rounds on quality, and treat a hard-constraint failure as a restart signal rather than another refinement pass. ## Interview framing The answer that lands has three moves: refinement is not monotone and I can say why; therefore I keep the best-scoring candidate rather than the last; and I judge with something the critic does not control, with cheap churn heuristics as a secondary stop. A candidate who says only "cap it at three iterations" has controlled cost and left quality entirely to chance.
- Why is keeping the best-scoring candidate a bigger win than tuning the stop condition?Because it makes an extra round harmless. If you always return the last version, one bad round destroys the run and the stop condition has to be exactly right. If you retain the argmax, the loop becomes a search: extra rounds cost tokens but can only find something better or be discarded. The stop condition then only has to control spend, not quality.
- The critique keeps raising the same objection round after round. What does that tell you?Either the actor cannot satisfy it — missing information, a constraint conflict, a capability limit — or the objection is not real and the critic is manufacturing findings. Both mean more rounds will not help. The useful response is to stop refining and escalate: surface the objection to a human, fetch the missing input, or restart with a different approach rather than iterating.
- How would you catch a loop that is polishing a fundamentally wrong answer?Check the coarse properties before spending rounds on quality: does it satisfy the hard constraints, does it answer the question that was asked, is it grounded in the right source. Treat a failure there as a restart signal, not a refinement one. Otherwise you spend the whole budget producing a well-written wrong answer, which is harder to catch downstream than a rough one.
saying these in an interview costs you the question
- Assumes each critique round strictly improves the output
- Returns the last version instead of the best-scoring one
- Uses the critic's own satisfaction as the stop criterion
- Treats a fixed iteration cap as a quality guarantee
- Keeps refining an answer that violates a hard constraint