Why does agent self-correction improve results with a test suite but often not without one?
answer
- Where does the critique come from?
- Same model, same blind spots
- Tests and schemas are outside evidence
- Ungrounded reflection is theatre
- No oracle, no reliable gain
basics
~20 sSelf-correction needs a signal the model did not produce. Failing tests, compiler errors and schema violations are outside evidence of a defect. Pure self-judgment re-samples the same model that wrote the output, so it rarely catches what it already missed.
solid answer
~50 sReflection is only as good as the critique signal it runs on. When an agent can execute something — run the test suite, invoke a type checker, validate a payload against its JSON Schema — the failure output is **evidence generated outside the model**, and feeding it back gives the next attempt information the first attempt did not have. That is why revise loops work well on code and structured output. With no verifier, the critique comes from the same model, with the same knowledge and the same blind spots, so it mostly restates the draft's assumptions in a confident voice. The empirical pattern is that self-correction without external grounding produces small or zero gains and sometimes net harm, while verifier-driven loops give large, repeatable gains. So the design question is never "should the agent reflect?" but "what can it check against?" — and if the answer is nothing, build a cheap checker before you buy a critic pass.
code
python · 8 linesdef repair_with_verifier(generate, run_tests, task, max_rounds=3):
patch = generate(task)
for _ in range(max_rounds):
result = run_tests(patch)
if result.passed:
return patch
patch = generate(task, previous=patch, failure=result.output)
return Nonego deeper
Know that self-correction means the agent reviews its output and tries again, and that it works best when something concrete, like a test run, can say the output is wrong.
Be able to name the verifier tiers — schema validation, compiler, tests — explain why feedback from outside the model carries information the model lacked, and say why ungrounded self-review often changes little.
Show you would measure the loop: ablate the verifier, compare one-shot against one and two rounds, and account for the extra latency and token cost per round in production.
Own the decision of whether to invest in building a verifier for a domain that lacks one, since that investment, not the critic prompt, is what determines whether reflection pays across the whole agent portfolio.
## The shape of a reflection loop A reflection (or self-correction) loop wraps the ordinary generate step in a cycle: produce a candidate, obtain a critique of that candidate, then produce a revised candidate that takes the critique into account, and repeat until some stopping rule fires. Named versions of this idea in the literature include **Self-Refine** (the model critiques and rewrites its own output), **Reflexion** (the agent writes a short verbal lesson after a failed attempt and carries it into the next attempt), and **CRITIC** (the critique is grounded in external tools rather than the model's opinion). What separates a loop that pays for itself from one that just burns tokens is where the critique comes from. ## Two kinds of critique signal **External verifier signal** is produced by something other than the model: - a unit test suite that fails with an assertion message and a stack trace - a compiler or type checker that names a symbol, a line and a type mismatch - a JSON Schema validator that reports which required field is absent or which enum value is unrecognized - a linter, a query planner, a simulator, a physical unit checker, a database that rejects a constraint violation **Self-judgment signal** is produced by the model itself: a second prompt asking "is this correct?", a self-assigned confidence score, a rubric the model grades its own answer against. The first kind carries information the generator did not have when it wrote the draft. The second kind, in the general case, does not. The model that failed to notice an off-by-one in generation is drawing on the same weights and the same context when asked whether the code is right, so its probability of noticing is not much better the second time. Worse, the critique itself is generated text and can be wrong, which means a confident but incorrect critique can degrade a correct draft. ## Why this asymmetry is so large in practice Code is the canonical case where reflection works, and it works because the environment is a cheap, precise, non-negotiable oracle. A patch either makes the failing test pass or it does not, and the failure output is specific enough to localize the defect. A loop of the form "apply patch, run pytest and the type checker, feed the failure text back, patch again" turns the model from a one-shot guesser into a search process over a verifiable space. Contrast a summarization or advisory task with no oracle. Asking the model to critique its own summary produces plausible-sounding notes ("the summary could mention the timeline more explicitly") that are unfalsifiable. Revisions churn the text, cost a full extra generation per round, add latency, and move quality around rather than up. This is reflection as theatre: the loop looks rigorous in the trace and delivers nothing measurable. ## Build the cheapest verifier you can Before reaching for a model-based critic, look for a deterministic check: 1. **Structural checks** — does the output parse? Does it validate against the schema? Are required fields present and enums in range? This is nearly free and catches a large share of real failures. 2. **Executable checks** — run it. Tests, type checks, a dry-run of the query, a compile. 3. **Consistency checks** — do the numbers in the output sum to the total? Does every cited identifier exist in the retrieved source? 4. **Model-based critique** — only for what the first three cannot express, and knowing it is the weakest tier. Running them in that order also controls cost: a schema violation should never consume a model call to discover. ## What self-judgment can still do Self-critique is not worthless. It helps most where the failure is a *format or instruction* miss rather than a *fact*: did the answer address every part of the request, is it in the required register, does it violate a stated constraint. Checking a draft against an explicit, externally supplied checklist is far more effective than an open-ended "is this good?", because the checklist supplies the missing information. Giving the critic material the generator did not see — the retrieved source documents, the original ticket, a style guide — also converts self-judgment into something closer to verification. On frontier models as of mid-2026, extended internal reasoning already absorbs much of what a naive extra self-critique pass used to buy, which further narrows the case for ungrounded reflection rounds. ## How to know which regime you are in Measure it. Run the task suite with reflection off, with one round on, and with the verifier feedback ablated. If reflection with a verifier beats one-shot and reflection without the verifier does not, you have your answer, and you know exactly which component is doing the work. Teams that skip this measurement routinely ship a critique pass that doubles cost for noise.
- You have no automated verifier for the task. What would you do before adding a self-critique pass?Look for a cheap partial oracle first: schema or format validation, arithmetic and totals consistency, checking that every cited identifier appears in the retrieved sources. If none exists, consider whether a checklist-driven critique with externally supplied criteria is enough, and measure the change on a held-out task set before shipping it. If the measurement shows no gain, drop the pass rather than paying for it.
- Can a model-based critic ever act like an external verifier?Partially, when it is given information the generator did not have. A critic that sees the retrieved source documents, the original requirements or a written rubric is checking the draft against evidence rather than against its own prior. That is materially stronger than an open-ended self-review, though still weaker than an executable check, since the critic's own output can be wrong.
- Why is JSON Schema validation worth running before any model critique?It is deterministic, costs microseconds rather than a model call, and catches a common and unambiguous failure class — missing required fields, wrong types, out-of-range enum values. Spending a full generation to discover that a field is absent is pure waste, and the validator's message is precise enough to make the repair attempt nearly always succeed.
A proofreader who never leaves the room and only has the draft in front of them will keep re-reading their own sentences as correct; handing them the printer's error report is what actually finds the defect.
saying these in an interview costs you the question
- Claims self-critique reliably catches the model's own factual errors
- Adds a critic pass without any measurement of whether it helps
- Treats a model's self-reported confidence as a verification signal
- Runs an expensive model critique before cheap schema validation
- Assumes more reflection rounds always mean better output