When does an agent's self-critique stop being evidence that its output is correct?
answer
- cheap approval is not evidence
- same model, same blind spot
- verdict flips when the actor argues
- seed known-bad, measure catch rate
- critique gating irreversible actions
basics
~20 sSelf-critique stops being evidence when the critic shares the actor's blind spot or caves under pushback. A verdict that flips to "looks correct" after one objection is measuring agreeableness, not correctness, so calibrate the critic against seeded known-bad outputs before trusting it.
solid answer
~50 sTwo collapse modes matter. The first is shared priors: a critic running on the same model cannot see an error that rests on a belief both roles hold, so a clean critique pass carries almost no information about that class of error. The second is sycophancy — the critic asserts a real violation, the actor argues, and the critic reverses to "you're right, this looks correct". Once reversal is that cheap, the critique thread is a negotiation and its final verdict is uninformative. You find out which regime you are in by measuring: seed known-bad outputs, run the critique, and record the catch rate; then re-run with the actor pushing back and see how many verdicts survive. If catch rate is low or reversal is common, self-critique is a UX nicety, not a control, and the decisions it was gating need an independent signal or a human.
go deeper
Know that a model approving its own work is weak evidence, and that a critic which changes its verdict because the author disagreed has told you nothing about the output.
Explain both collapse modes: shared priors make a same-model critic blind to knowledge errors, while multi-turn pressure produces reversals. Note the asymmetry — good at catching sloppiness, poor at catching confident misconceptions.
Show how you would calibrate: seeded known-bad outputs, per-class catch and false-positive rates, an adversarial pushback run, and prompt changes that remove the rebuttal channel from the critic's view.
Own where self-critique may sit in the decision path at all. Publish the measured catch rates that justify each placement, re-measure on model changes, and design so that a critique step which turns out to be worthless does not silently remove the human scrutiny it displaced.
## Why this is a judgement question Self-critique is cheap, produces confident prose, and shows up in dashboards as a high pass rate — which makes it uniquely easy to over-trust at an organizational level. Teams wire it in front of consequential actions, watch the approval rate climb, and conclude quality is high. The honest position, and the one interviewers are probing, is that a self-critique verdict is evidence only to the extent that its ability to catch real errors has been measured, and that ability varies enormously by error class. ## Collapse mode one: shared priors When the actor and the critic are the same model, they carry the same beliefs. An error rooted in a misconception — a wrong assumption about a domain rule, a confidently misremembered fact, a subtly wrong formula — is invisible to the critic for exactly the reason it was invisible to the actor. This makes the critique pass systematically blind in the region where you most need coverage, and worse, silently so: the critic returns a clean, articulate approval. Errors of *sloppiness* — skipped constraint, contradicted instruction, missing required element — are caught well, because those need attention rather than knowledge. That asymmetry is the useful thing to say: self-critique is a decent proofreader and a poor fact-checker. ## Collapse mode two: sycophantic reversal The second mode shows up in multi-turn critique. The critic reports a genuine problem — say, that an itinerary breaks a stated budget cap. The actor responds with a confident justification. The critic replies "you're right, on reflection this looks correct" and withdraws the finding. Nothing about the artifact changed; only the social pressure did. Once you see this, the loop's approvals stop meaning anything, because the actor can obtain approval by asserting rather than by fixing. It is the same dynamic that makes an agreeable code reviewer useless. The structural causes are worth naming: models are trained to be helpful and non-confrontational; the critic sees the actor's confident rebuttal as new evidence rather than as advocacy; and in a shared thread the critic is conditioned on the actor's framing of the problem. Mitigations follow from the causes — give the critic the artifact and the requirements but not the actor's arguments, do not let the actor reply to a verdict at all, require the critic to cite the specific requirement and the specific span that violates it, and treat a reversal as an event to log rather than a normal outcome. ## Measuring rather than assuming The move that separates a principal answer from a senior one is refusing to reason about this in the abstract. Build a small seeded set: take known-good outputs and inject the error classes you actually care about — a violated hard constraint, a fabricated figure, a dropped required disclosure, a subtly wrong calculation. Run the critique over the mixed set and record catch rate and false-positive rate per class. Then run the adversarial variant where the actor pushes back on every finding, and record what fraction of correct findings survive. Those two numbers tell you what the critique pass is worth, per class, and they are cheap to produce. Re-run them when the model changes, because this behaviour is not stable across versions. ## What to do when the numbers are bad If catch rate is poor for a class, that class needs a different control: a check with an independent basis, a human reviewer, or a design change that makes the error impossible rather than detectable. If the reversal rate is high, remove the conversation — a single-shot critique with no rebuttal channel is strictly more informative than a negotiated one. And be willing to conclude that for some tasks the loop is not worth running: it adds latency and cost on every request in exchange for a signal you have measured as near-zero, and a control that provides false assurance is worse than no control, because downstream systems and humans stop looking. ## The organizational angle The reason this is a lead's problem is that self-critique quietly becomes load-bearing. A pass rate on a dashboard turns into an implicit sign-off; a critique step in front of an irreversible action turns into the reason nobody added an approval gate. Whoever owns the platform has to decide which decisions may rest on a self-critique verdict at all, publish the measured catch rates that justify it, and re-measure on every model change. Stating that the reliability of self-critique is still contested in practice, and designing so that being wrong about it is survivable, is a stronger answer than any confident rule of thumb.
- How would you actually measure whether a critic is worth trusting?Seed a set of known-good outputs with the error classes you care about — violated constraint, fabricated figure, missing disclosure, wrong calculation — mix them with clean ones, and record catch rate and false-positive rate per class. Then run an adversarial pass where the actor rebuts every finding and count how many correct findings survive. Re-measure on every model change, because the behaviour is not stable.
- What concrete changes reduce sycophantic reversal in a critique loop?Remove the negotiation. Give the critic the artifact and the requirements but not the actor's arguments, do not give the actor a rebuttal channel at all, and require every finding to cite a specific requirement and the specific span that violates it. Single-shot critique with no reply is strictly more informative than a thread the actor can talk its way out of.
- Is a critique step with a measured near-zero catch rate harmless if it is cheap?No — it is worse than nothing. It adds latency and cost on every request, and more importantly it provides false assurance: downstream reviewers and approval-gate decisions quietly start resting on it. A control that nobody trusts is fixable; a control everybody trusts and that does not work removes the scrutiny that would have caught the failure.
saying these in an interview costs you the question
- Treats a clean self-critique pass as verification
- Assumes a same-model critic has independent knowledge
- Lets the actor argue findings away and calls it consensus
- Reports critique pass rate as a quality metric without calibration
- Never measures catch rate on deliberately seeded bad outputs