skip to content

Why does a prompt-rewrite loop critique the prompt before editing it?

level: middleimportance: should knowfreq 35%

answer

  1. diagnosis first, edit second
  2. a written reason, not just a new prompt
  3. the critique is the reviewable artifact
  4. generic critiques mean no signal left
  5. a hypothesis, confirmed only by re-scoring

basics

~20 s

Splitting diagnosis from editing forces the model to state, in writing, why the prompt failed. That written reason constrains the next edit to the observed defect, makes the change reviewable by a human, and gives a stopping signal when critiques go generic.

solid answer

~60 s

A one-shot "here are the failures, rewrite it" call gives you a new prompt and no reason. The critique-then-edit shape — sometimes described as a textual gradient, by analogy with a gradient telling you which direction to move — splits the call in two: first ask for a written diagnosis of *why* the current instruction produced those outputs, then ask for an edit that applies only that diagnosis. Three things improve. The edit becomes targeted rather than a wholesale rewrite that discards working sentences. The reason is inspectable, so a human reviewing the loop can reject a bad diagnosis before it becomes a bad prompt. And the critiques themselves become a stop signal: once every round produces "the instruction should be clearer and the model should double-check", the loop has run out of signal and further rounds will only add words. The critique is still an unverified model output, so ground it in actual failing cases and confirm the edit by re-scoring, not by how convincing the critique reads.

go deeper

for a junior

Know that the loop asks two questions, not one: first why the prompt failed, then how to change it. The written reason is what makes the second step targeted.

for a middle

Explain the textual-gradient framing, why the critique must be conditioned on real failing outputs, and why generic critiques such as 'be more careful' signal the loop should stop.

for a senior

Demonstrate that you treat the critique as an unverified hypothesis confirmed only by re-scoring on unseen data, cap the rounds, keep the champion rather than the latest, and read the critique stream for evidence the real defect is not the prompt at all.

for a principal

Own whether the critique trail becomes a durable engineering asset: retaining diagnoses as the rationale behind every production prompt clause, deciding who signs off on model-written edits, and using recurring diagnoses to redirect effort from prompt wording to the system defect underneath.

## The loop shape The naive self-refinement loop is one call: current prompt plus failures in, new prompt out. The critique-then-edit variant splits that into two calls with different jobs. 1. **Critique.** "Here is the instruction, here are cases where it produced the wrong output. Explain why the instruction allowed those outputs." The output is prose — a diagnosis, not a prompt. 2. **Edit.** "Here is the instruction and this diagnosis. Rewrite the instruction to address exactly this, changing nothing else. Output only the new instruction." The analogy that gave the pattern its name is optimization: the critique plays the role of a gradient — a written statement of which direction to move — and the edit plays the role of the step. The analogy is loose (there is no metric being differentiated and no guarantee of descent), but it captures the useful part: separating *where to go* from *taking the step*. ## Why the extra call earns its cost **Targeted edits.** Asked to rewrite directly, models tend to produce a fresh instruction rather than a modified one — which throws away phrasing that was working, and makes each round's diff unreadable. Conditioning the edit on a specific diagnosis keeps the change local and the diff small enough to review. **Inspectability.** The diagnosis is the artifact a human can actually audit. "The instruction says 'classify the incident' but never lists the allowed labels, so the model answers in free text" is a claim you can agree or disagree with in seconds. The rewritten prompt alone gives you nothing to check except by re-reading it in full. **Reusable signal.** Diagnoses across rounds aggregate into a picture of how the prompt is failing — ambiguity in one clause, missing output format, an unstated tie-breaking rule. That picture often tells you the prompt was never the problem: if the diagnosis keeps saying "the model lacks the policy text", the fix is retrieval, not phrasing. **A stopping signal.** This is the underrated one. As long as critiques name specific defects, the loop has signal. When they degrade into content-free exhortation — "the instruction should emphasise care", "the model should double-check its answer" — the critic has nothing left to diagnose and every further round will append inert clauses. That degeneration is easier to spot in the critique stream than in the prompt itself, where each addition looks locally reasonable. ## Grounding the critique A critique of a prompt read in isolation is speculation: the model will find things to say about any instruction. The critique must be conditioned on real failing outputs, ideally with the expected output alongside, so the diagnosis explains an observed discrepancy rather than an imagined one. Include a passing case or two as a control; a critique that condemns behaviour that is actually correct is a signal that the critic is confabulating. ## Separating critic from editor Using one model for both roles is common and works, but the critic inherits the editor's blind spots. Some setups use a stronger model as critic, or run the critique at a lower temperature than the edit. Either way, the critique is an unverified model output. It is a hypothesis about why the prompt failed, and the only thing that confirms it is that the edited prompt scores better on data the critic never saw. ## Failure modes - **Critique that flatters.** Ask a model whether a prompt is good and it will often say yes. Ask *why these specific outputs were wrong* and it has something concrete to work with. - **Diagnosis that names the output, not the cause.** "The model gave the wrong label" restates the failure; "the instruction never says what to do when two labels both apply" names a cause you can edit. - **Edit that exceeds its brief.** If the editor rewrites everything anyway, add an explicit constraint ("change at most two sentences") or diff the result and reject oversized edits. - **Trusting the critique as evidence.** A persuasive diagnosis followed by a prompt that scores worse is common. Score, then decide. - **Running the loop with no cap.** Even with good critiques, each round costs two calls plus an evaluation; set a round cap and keep the best-scoring version rather than the last one. ## Practical shape A workable configuration is three to five rounds, each round: score the current prompt, critique against the worst failures, edit under a size constraint, re-score on a held-out split, keep the champion. Stop early when the critique goes generic or when the held-out score stops moving beyond noise. Keep every round's critique — when someone later asks why the production prompt contains a strange clause, the diagnosis that produced it is the answer.

  • How do you tell a useful critique from a useless one before spending the edit call?
    A useful critique names a property of the instruction that permitted the wrong output — a missing rule, an ambiguous clause, an unspecified format. A useless one restates the failure ("the answer was wrong") or offers exhortation ("be more careful"). You can filter cheaply on that distinction: if the critique does not quote or refer to a specific part of the instruction, discard the round.
  • Should the critic and the editor be the same model?
    They can be, and usually are for cost reasons, but the critic then shares the editor's blind spots — it will not flag ambiguity it does not itself perceive. Using a stronger or simply different model as critic surfaces defects the editor misses. Either way the scoring model must be the production executor, since that is whose behaviour you are trying to change.
  • Does a persuasive critique justify keeping the edited prompt?
    No. The critique is an unverified hypothesis about why the prompt failed; models produce fluent diagnoses for prompts that were fine. Only a score on data the critic never saw settles it. Keep the best-scoring prompt across rounds rather than the most recent one, so a convincing-but-harmful round cannot become the champion.

It is the difference between a code review that says "rewrite this file" and one that says "this branch never handles the empty case" — the second constrains the change and can be argued with.

saying these in an interview costs you the question

  • Treats a convincing critique as proof the edit is an improvement
  • Asks the model to critique the prompt without showing any failing outputs
  • Lets the editor rewrite the whole prompt instead of applying the diagnosis
  • Keeps the last round's prompt rather than the best-scoring one
  • Runs rounds indefinitely as long as the critiques keep coming

context