skip to content

Why give a critic agent a fresh context instead of the author agent's thread?

level: middleimportance: should knowfreq 45%

answer

  1. a second opinion, not the same one twice
  2. the author's window is long and anchored
  3. context rot at the end of a session
  4. review the artifact, not the story
  5. re-supplying the spec is the cost

basics

~20 s

A critic started in a fresh context judges the artifact rather than the story that produced it. Continuing the author's thread anchors the critic on the author's assumptions and inherits a long, degraded window, so it tends to ratify work instead of finding defects.

solid answer

~50 s

In a proposer/critic pair, the critic is a separate agent with its own window, and that isolation is the point. If you append a "now review your work" turn to the author's thread, the reviewer's context is saturated with the author's reasoning — every justification for every choice is already in the window, so the model is anchored and tends to agree with itself. A fresh critic sees only the artifact and the requirements, which is closer to what a real reviewer gets. It also dodges context rot: the author's thread is long by the time the work is done, and quality degrades with window length, whereas the critic starts short and focused. The cost is real — you must re-supply the spec, the constraints and any evidence, and you lose the rationale behind deliberate choices, which produces objections to decisions that were already considered.

go deeper

for a junior

Know that the critic is a separate agent with its own context, and that it is shown the finished artifact plus the requirements rather than the author's conversation. Saying why — a fresh reader spots what the author cannot — is enough at this level.

for a middle

Explain the two mechanisms: anchoring on one's own reasoning, and degradation as the author's window fills. Be able to list what you would put in the critic's context and what you would deliberately leave out, and name the cost of re-supplying it.

for a senior

Demonstrate you have tuned this loop — severity-ranked bounded findings, verifier tools so criticism is grounded, a revision cap with a defined exit, and a short intent summary to cut false objections. Be ready to describe how you measured whether the critic was catching real defects.

for a principal

Own the discipline question: additional agents add judgement, not writes. Be prepared to argue where in a pipeline critique gates belong, what a critic that never fires or always fires tells you about the system, and how you would justify the extra latency in a review path.

## The pattern A proposer/critic pair — sometimes called an actor-critic or generator-verifier pair — splits production from judgement across two agents. The proposer produces an artifact: a patch, a plan, a policy decision, a draft. The critic receives that artifact and returns findings. A controller then decides whether to accept, ask for a revision, or escalate. The distinguishing design choice, and the thing interviewers probe, is that the critic runs in its own context window rather than as another turn in the proposer's conversation. ## Why isolation changes the output Three effects stack up. **Anchoring.** A model conditioned on a long chain of its own reasoning has, in its window, an argument for every choice it made. Asked to review, it is being asked to contradict the most recent and most heavily represented content in its own context. Models are poor at that; the usual result is a review that restates the design and declares it sound. **Context rot.** Quality degrades as the window fills — retrieval of details from the middle of a long context weakens, and instructions given early lose force. By the time an agent has finished a substantial task, its window holds tool results, dead ends and superseded intermediate states. A reviewer inheriting that window is working in the worst part of the quality curve. A fresh critic starts with a short, clean window containing only what matters. **Framing.** The author's thread encodes the problem as the author decomposed it. A defect that comes from a wrong decomposition is invisible from inside that framing, because every subsequent step is locally consistent with it. A critic given the requirements and the artifact, but not the decomposition, can notice that the artifact does not satisfy the requirements at all. Code review is the pattern's most-reported production use: a reviewer agent handed the diff and the requirements, with no access to the authoring session, is reported in 2026 practice to surface a couple of genuine bugs per pull request, a majority of them substantive rather than stylistic. That is a materially different result from asking the author agent to double-check itself. ## What the critic should actually receive Build the critic's context deliberately rather than by inheritance: - The artifact under review, complete — the full diff, the full plan, the full decision. - The requirements, acceptance criteria or policy the artifact must satisfy. - The rubric: what counts as a finding, and how severity is ranked. - Any evidence the critic needs to check claims — tests it can run, documents it can retrieve, the source data. Deliberately withhold the proposer's justification narrative. If a choice cannot be defended from the artifact and the requirements alone, that is itself a finding. ## What you lose Isolation is not free and a good answer says so. You pay to re-supply context that the author already had, which is real tokens and, for large artifacts, real latency. You lose the rationale, so the critic will sometimes object to a constraint the author considered and consciously rejected — noise the controller has to filter. And you lose the ability to score the process: a critic that never saw the trajectory cannot tell you the author took twelve steps where three would do. The practical compromise is to pass a short, structured summary of intent — the goal and any explicit constraints or exclusions — while withholding the reasoning transcript. That kills the worst noise without restoring the anchoring. ## Failure modes of the pair **Nitpick storms.** An unconstrained critic will always find something, because "find problems" is satisfiable at any depth. Require severity ranking and cap the number of returned findings, so triage is done by the critic rather than downstream. **Ungrounded criticism.** A critic with no tools is arguing from priors and will invent defects as readily as the author invented code. Give it the same verifiers the author had — run the tests, check the schema, query the source. **Ping-pong.** Author revises, critic objects again, forever. Cap the cycles, and on exhaustion return the artifact with the open findings attached rather than blocking. **Single-writer discipline.** Keep the critic advisory. It contributes intelligence, not edits. Two agents writing the same artifact is the design that reliably corrupts state; the 2026 consensus is that additional agents should add judgement while writes stay single-threaded. ## The one-sentence version A critic in a clean context is a second opinion; a critic in the author's context is the same opinion asked twice.

  • What would you still pass from the author to the critic, if anything?
    A short structured statement of intent: the goal, the acceptance criteria, and any constraints the author was deliberately working under. That removes the most common noise — objections to choices that were made knowingly — without importing the reasoning transcript that causes anchoring. Anything longer starts re-creating the author's framing inside the critic's window.
  • How do you stop the critic and author ping-ponging forever?
    Cap the revision cycles, typically at two or three, and make the exit behaviour explicit: on exhaustion, return the artifact with the remaining findings attached rather than blocking or looping. Requiring the critic to rank findings by severity and return a bounded list also helps, because it stops each new round from surfacing a fresh layer of cosmetic objections.
  • Should the critic be allowed to edit the artifact directly?
    No — keep it advisory. Current practice is that additional agents contribute judgement while writes stay single-threaded, because two agents mutating the same artifact produces conflicting or lost edits that are extremely hard to debug. The critic returns findings; the author, or a dedicated applier, makes the change and the critic re-reviews the result.

saying these in an interview costs you the question

  • Says a self-review turn in the same thread is equivalent to a separate critic
  • Passes the author's full transcript to the critic to "give it context"
  • Lets the critic edit the artifact directly alongside the author
  • Runs an unbounded critique loop with no revision cap
  • Gives the critic no tools, so it argues from priors instead of evidence

context