Why split an agent's generator and critic into two separate prompts?
answer
- Same conversation means inherited framing
- Authorship makes the model agreeable
- Give criteria, not "is this good?"
- Withhold the generator's scratchpad
- Two prompts, still one agent
basics
~20 sSeparation stops the critic from inheriting the generator's framing. A fresh prompt with its own system message, the task, the draft, and explicit failure criteria — but not the generator's reasoning — reviews the artifact on its merits instead of defending the path that produced it.
solid answer
~50 sWhen one prompt both writes and reviews, the review is conditioned on everything that led to the draft: the chain of reasoning, the assumptions taken as settled, the fact that the text is the model's own. All three bias it toward approval. Splitting the roles means the critic gets a **different system message** ("you are reviewing a candidate solution against these criteria"), the task statement, the artifact itself — and deliberately *not* the generator's scratchpad. It should be given concrete failure criteria to check rather than an open-ended "is this good?", and ideally material the generator never saw, such as the source documents or the original requirements. The cost is a second full call per round plus the tokens to restate the task. The benefit is a critique that actually finds things. Note this is still one agent with two prompts, not two deployed agents debating.
code
python · 14 linesCRITIC_SYSTEM = (
"You review a candidate solution against the stated criteria. "
"Return a JSON list of findings, each with location, issue and severity, "
"then a top-level verdict of pass or fail."
)
def critique(call_model, task, artifact, criteria):
return call_model(
system=CRITIC_SYSTEM,
messages=[{
"role": "user",
"content": f"TASK\n{task}\n\nCRITERIA\n{criteria}\n\nCANDIDATE\n{artifact}",
}],
)go deeper
Know that the critic works better as its own prompt with its own instructions, and that showing it the draft alone — not the whole earlier conversation — is the usual setup.
Explain the three biases separation removes: inherited framing, self-preference toward its own text, and commitment to the reasoning trace already in context.
Discuss what to supply the critic — restated task, explicit failure criteria, evidence the generator lacked, a structured findings contract — and the latency and token cost of a second call per round.
Own the decision of where the critique tier sits at all: on every request, on high-stakes outputs only, or as sampled monitoring, weighed against building the mechanical check that would replace it.
## The problem with critiquing in place The simplest reflection design appends "now review your answer and improve it" to the same conversation. It is cheap and it under-performs, for three compounding reasons. **Framing inheritance.** The draft was produced by a particular reading of the task. That reading — which requirement mattered, which ambiguity was resolved which way, which approach was chosen — sits in the context as settled fact. A review conditioned on it can find local slips but almost never questions the frame itself, which is where the expensive errors live. **Self-preference.** Models rate text more favourably when it is presented as their own. Continuing the same conversation makes authorship maximally salient. **Commitment to the reasoning trace.** If the generator's step-by-step reasoning is in context, the critic is effectively asked to contradict a chain of statements it just produced. The likelihood landscape favours consistency. ## What separation actually means A separated critic is a distinct model call with: - **Its own system message** casting it as a reviewer, with a standard to hold the work to. - **The task statement, restated** — not inherited from the conversation, so the critic reads the original requirement rather than the generator's paraphrase of it. - **The artifact alone**, with the generator's scratchpad and intermediate reasoning omitted. - **Explicit failure criteria**: the specific defects to look for, expressed as checks. "Verify every figure in the summary appears in the source table" beats "check for accuracy" by a wide margin. - Optionally **evidence the generator did not have**: the retrieved documents, the style guide, prior incident notes, the acceptance criteria. - A **structured output contract**: a list of findings, each with a location and a severity, plus an explicit pass or fail. Free-form prose critique is hard to act on and easy to ignore. That last point matters for the repair step too. The generator then receives findings it must address, not a paragraph of general advice, which makes the next revision targeted rather than a rewrite. ## Should the critic see the reasoning? Usually not, with one exception. If the goal is to catch *process* errors — a wrong intermediate calculation, an unjustified leap — the trace is the thing being reviewed and must be supplied. If the goal is to catch *artifact* errors, hiding the trace is better, because the trace is precisely what makes the artifact look justified. Decide which failure class you are hunting and supply context accordingly; supplying everything by default is the common mistake. ## What separation does not buy you A separate prompt does not make the critic knowledgeable about things the model does not know. The same weights are still doing the judging, so a factual error rooted in a knowledge gap remains invisible. Separation reduces bias; it does not create an oracle. Where a mechanical check exists — tests, a schema, a lookup — run it before or alongside the critic, and let the critic handle only what the check cannot express. It also does not require a different model, although using a different one is a reasonable variation. Using a *cheaper* model as the critic is common for cost reasons and works acceptably when the criteria are concrete checks; it works poorly when the critique requires the same reasoning depth as the generation. ## Cost and when to skip it Every round is now at least two full calls plus the tokens to restate the task and the artifact, and on a long artifact those tokens dominate. Under a latency budget, a separated critic can be the difference between a two-second and a six-second response. Reasonable ways to spend less: run the critic only when a cheap check is inconclusive; run it on a sample of traffic to monitor quality rather than on every request; or gate it to high-stakes outputs only. ## Keep the scope straight This is a single agent with two prompts in one loop. It is not a debate between deployed agents, not a voting panel, and not an orchestrator delegating to worker agents — those are different architectures with different failure modes and much higher coordination cost. The generator-critic split is the cheapest useful step up from critiquing in place, and for most systems it is where the returns stop.
- Should the critic ever see the generator's reasoning trace?Only when process errors are the target — a wrong intermediate step or an unjustified leap can only be caught by reading the trace. When artifact errors are the target, withhold it: the trace is exactly what makes a flawed output look justified, and including it biases the critic toward consistency with reasoning it just produced.
- Does using a cheaper model for the critic work?It works when the criteria are concrete, checkable items — required fields present, figures matching the source, format rules honoured — because verification against a checklist is easier than generation. It works badly when the critique itself needs the generator's reasoning depth, in which case the small model produces confident, shallow findings that waste repair rounds.
- How do you keep critic findings actionable in the repair step?Require structured output: each finding carries a location, the issue, and a severity, with an explicit pass or fail verdict at the top. The repair prompt then addresses an enumerated list rather than a paragraph of advice, which keeps revisions targeted and lets you drop low-severity findings when the budget is tight.
It is the difference between proofreading your own draft with your notes open beside you and handing a clean copy to a colleague who has only the brief and a checklist.
saying these in an interview costs you the question
- Appends "now review your answer" to the same conversation and calls it a critic
- Passes the full generator scratchpad to the critic by default
- Asks the critic "is this good?" with no concrete criteria
- Confuses a generator-critic split with a multi-agent debate setup
- Expects separation to fix factual gaps the model simply does not know