Why does Reflexion split the agent into actor, evaluator and self-reflection roles?
answer
- produce, judge, diagnose — three jobs
- the writer is a poor grader of itself
- a score is not an instruction
- sparse verdict into actionable text
- separate passes, not one combined prompt
basics
~20 sThree different jobs with different failure modes: the actor produces the attempt, the evaluator judges it and returns a verdict or score, and the self-reflection step turns that thin verdict into concrete written advice the actor can act on next time.
solid answer
~50 sThe split exists because judging and producing pull in opposite directions. An actor that grades its own work in the same pass is anchored on the reasoning that produced the error, so it rationalizes rather than checks. Giving the evaluator its own pass — its own prompt, its own view of the output, sometimes a deterministic check — restores some independence. The third role exists because an evaluator's output is usually thin: a score, a failed constraint, a pass/fail. That tells the actor it lost but not what to change. The self-reflection step reads the trajectory *and* the verdict and writes the actionable version. Consider a travel itinerary generated against hard constraints — budget cap, no red-eye flights, museum closed Mondays. A separate critic can check those mechanically and report "violates the Monday constraint"; the reflection step converts that into "reorder day three so the museum falls on Tuesday".
go deeper
Know the three names and what each does: actor makes the attempt, evaluator judges it, reflection writes the advice for the next try. Being able to list them cleanly is enough at this level.
Explain why the passes are separate — a model grading its own output in the same pass is anchored on the reasoning that produced it — and why a bare score cannot be acted on without a diagnosis step.
Show where you would put a deterministic evaluator instead of a model, and be honest that same-model separation buys framing independence but not freedom from shared blind spots. Mention the per-attempt cost of tripling passes.
Own the policy question: which task classes get the full three-role loop at all, given it multiplies latency and spend on exactly the requests that are already failing. Tie that to how expensive an undetected wrong answer is in each class.
## Three roles, three jobs Reflexion decomposes a self-improving loop into roles that are often collapsed in practice, and the decomposition is the interesting part. **Actor.** Produces the attempt — the reasoning, the actions, the observations, the final answer. It is optimizing for completing the task, and it holds all the assumptions that produced the attempt, correct and incorrect alike. **Evaluator.** Judges the produced trajectory and returns a signal. That signal can be programmatic (a constraint check, a comparison against an expected result) or a model asked to grade against a rubric. Its output is typically sparse: pass/fail, a score, or a named violated constraint. **Self-reflection.** Reads the trajectory together with the evaluator's verdict and writes a short natural-language explanation of the likely cause plus a concrete change. Its output is what is carried into the next attempt. ## Why not let the actor judge itself The practical argument is anchoring. An actor asked, in the same pass, "is this right?" is conditioned on the chain of reasoning that produced the answer. Everything in that chain looks locally justified — that is why it was written — so the self-check tends to confirm. Giving the evaluation its own pass changes the framing: the evaluator sees an artifact to be attacked, not a plan it just committed to. It can also be given a different, stricter prompt, a rubric the actor never saw, or a narrower brief ("check only the hard constraints") that resists the actor's persuasion. The independence is partial, not absolute. If both roles are the same model, they share priors and blind spots: a misconception about how a library behaves will be invisible to both. Role separation buys framing independence, not epistemic independence, and a strong candidate says so rather than overselling it. ## Why the reflection step is separate from evaluation This is the part people usually miss. Evaluation and diagnosis are different tasks. "Score 0.3" or "violates the Monday constraint" is a judgement, and it is exactly the signal an actor cannot act on — it says where it landed, not which decision to change. The reflection role's whole purpose is signal conversion: from sparse verdict to specific instruction, grounded in the trajectory that produced it. Keeping it separate from the evaluator also keeps the evaluator honest; an evaluator asked to both grade and suggest fixes drifts toward grading its own suggested fix favourably. ## The itinerary example An agent plans a five-day trip under hard constraints the requester supplied: total budget, no departures before 06:00, and one museum that closes on Mondays. The actor writes a plausible itinerary. A separate critic pass, holding only the constraint list, reports two violations — the budget is exceeded by 12% and day three visits the museum on a Monday. Neither statement tells the actor what to do; the budget one in particular has many possible remedies. The reflection step reads both the itinerary and the violations and writes: "swap days three and four so the museum falls on Tuesday; the 12% overage is concentrated in the two internal flights, replace one with rail". That is what the next attempt consumes. ## Practical shapes In real systems these roles are rarely three separate services. They are commonly three prompts against the same model, sometimes with a cheaper model as evaluator, sometimes with the evaluator replaced entirely by a deterministic check where one exists. What matters for the interview is that you can articulate *why* they are distinct passes rather than one instruction saying "solve this, then check your work, then fix it". You should also be able to say what the split costs: three passes instead of one, roughly triple the latency and token spend on every failing attempt, which is why teams gate reflection on the tasks where a wrong answer is expensive. ## What weak answers sound like "You add a critic so the model checks itself" misses both the reason the pass separation matters and the existence of the third role. "Separating them makes the critic objective" overstates it — same model, same blind spots. The calibrated answer names the anchoring argument, the signal-conversion argument, and the residual shared-prior limitation.
- If the actor and evaluator are the same model, what independence have you actually gained?Framing independence, not epistemic independence. The evaluator pass is not anchored on the reasoning chain that produced the answer and can be given a stricter rubric the actor never saw, which catches sloppiness and constraint violations. It cannot catch a shared misconception — if the model believes an API behaves a certain way, both roles believe it. That is the argument for a deterministic check where one is available.
- What does this split cost you at runtime?Roughly three model passes per failing attempt instead of one, so latency and tokens scale with attempts times roles. On high-volume, low-stakes tasks that is rarely worth it. Teams typically gate the full loop on tasks where a wrong answer is expensive or hard to detect downstream, and let cheap tasks fail fast to a human or a default.
- Can the evaluator be non-model code?Often yes, and that is usually the stronger design when the task has checkable properties — constraints, invariants, schemas. A mechanical evaluator does not flatter, does not drift, and costs nothing per call. The self-reflection role still earns its keep, because a mechanical verdict is even sparser than a model's and needs converting into a specific change.
saying these in an interview costs you the question
- Claims a same-model critic is objective or independent
- Collapses evaluation and diagnosis into one role
- Says the actor can just check its own work in the same pass
- Ignores the extra latency and token cost of three passes
- Assumes the evaluator must be another LLM