skip to content

Reflexion & Self-Critique

Letting the agent write a short verbal post-mortem after a failed attempt and carry that note into the retry, instead of blindly repeating the same action. It is the usual follow-up to "how do you stop an agent from looping?".

on this pageshow

questions

4

What does a Reflexion-style agent carry into its next attempt after a failed one?

level: middleimportance: must knowfreq 62%

answer

  1. feedback in words, not weights
  2. the retry can read the note
  3. failure summarized into the next prompt
  4. episodic lesson, scoped to the task
  5. verbal reinforcement, no gradient step

basics

~20 s

A short self-written note in plain language about why the attempt failed and what to do differently. That note is appended to the next attempt's context, so the retry starts from an explicit lesson instead of repeating the same action.

solid answer

~50 s

After a failure the agent is asked to write a brief post-mortem — what it tried, what the failure signal was, what it will change — and that text is injected into the retry's context. The retry is therefore no longer an independent sample: the model conditions on its own diagnosis of the last attempt. This is usually called **verbal reinforcement**, because the learning signal is text in the context window rather than a gradient step; no weights change and nothing persists once the context is dropped. A coding agent that fails hidden tests and writes "I assumed the input list was sorted" often passes on attempt two, because the note names the faulty assumption that the raw stack trace did not. The notes are episodic: scoped to this task, kept for the last few attempts, discarded when the task ends.

code

python · 20 lines
python
def build_prompt(task: str, lessons: list[str]) -> str:
    notes = "\n".join(f"- {n}" for n in lessons) or "- none yet"
    return f"TASK: {task}\nLESSONS FROM PREVIOUS ATTEMPTS:\n{notes}"


def reflect(error: str) -> str:
    return f"Attempt failed with: {error}. Validate that assumption before coding."


lessons: list[str] = []
prompt = ""
for _ in range(3):
    prompt = build_prompt("sort-and-dedupe the input list", lessons)
    passed = False  # stand-in for running the real attempt
    error = "test_unsorted_input failed"
    if passed:
        break
    lessons.append(reflect(error))

print(prompt)

go deeper

for a junior

Be able to say plainly that the agent writes a short note about why the attempt failed, and that the note is put into the next attempt's prompt so the retry differs from the first.

for a middle

Explain the mechanics: trajectory, failure signal, reflection text, injection into the retry. Stress that the feedback is text in the context window, not a weight update, and that it disappears with the context.

for a senior

Show the production judgment: bound the note buffer, constrain the reflection prompt to a localized cause plus a concrete next action, and recognize that a thin failure signal yields confident but wrong diagnoses that steer the retry away from a near-miss.

for a principal

Own the tradeoff between episodic notes and durable learning — every session rediscovering the same lesson is a real cost, but promoting notes into a shared store buys invalidation, scoping and contamination problems. Decide deliberately which failures deserve to outlive their episode.

## The problem it solves An agent that simply retries after a failure is drawing a fresh sample from almost the same distribution: same prompt, same tools, same context. Temperature buys variation, not insight, which is why unguarded agents call the same broken action three, five, ten times. Reflexion's contribution is to make the retry *conditioned on an explanation of the previous failure*, written by the model itself, in natural language. ## The loop 1. The agent attempts the task and produces a trajectory — its reasoning, its actions and their observations, and a final result. 2. Something judges that trajectory and returns a signal. It may be as thin as pass/fail. 3. A reflection step reads the trajectory plus the signal and emits a few sentences: the likely cause of the failure and a concrete change for next time. 4. That reflection text is appended to a small buffer and injected into the next attempt's prompt, typically under a heading like "lessons from previous attempts". 5. Repeat until the task passes or an attempt budget is exhausted. ## Why "verbal" reinforcement The framing borrows the vocabulary of reinforcement learning — attempt, reward signal, policy improvement — but nothing is optimized. The policy is the prompt, and the update is a paragraph of English appended to it. This has three practical consequences. It is fast: improvement lands in the next request rather than in the next training run. It is inspectable: the "weights" are sentences you can read, audit and delete, which matters enormously when debugging why an agent suddenly changed behaviour. And it is fragile: the moment that context is compacted, truncated or the session ends, the lesson is gone. Nothing has been learned by the model in any durable sense. ## What a useful note contains A good reflection localizes the fault and names an action. "The tests failed" is worthless — the agent already saw that. "I assumed the input list was sorted, so my binary search returned wrong indices for unsorted input; next time sort or validate the input first" is useful because it converts an outcome into a corrected assumption. In practice you get better notes by constraining the reflection prompt: ask for at most three sentences, forbid restating the error verbatim, and require a specific next action rather than a resolution to "be more careful". ## Where the notes live They belong to the episode. A common shape is a bounded buffer of the last two or three reflections for the current task, injected verbatim into each retry. Bounding matters for two reasons: context is finite and reflections are cheap to generate, and stale advice from a superseded approach actively misleads later attempts — an agent told "the API needs pagination" after it has already switched to a bulk endpoint will spend an attempt honouring an obsolete lesson. Promoting a note into a durable, cross-session store is a different design decision with different hazards, and interviewers expect you to distinguish the two: episodic notes are cheap and self-limiting, persistent memory needs deduplication, invalidation and scoping. ## When it fails The technique inherits the model's ability to diagnose itself. If the failure signal does not localize the fault — "incorrect answer", no more — the reflection is a plausible-sounding guess, and a confident wrong diagnosis is worse than none because the retry now optimizes against a phantom cause. Reflexion also cannot rescue a task the model fundamentally cannot do; it converts sparse feedback into direction, it does not add capability. And notes accumulate: past a handful of attempts, most of the context is the agent talking to itself about its own failures, which crowds out the actual task. ## How to talk about it The strong answer names the mechanism (text feedback carried across attempts), the boundary (no weight update, no durable learning), and the dependency (the value of the note tracks the richness of the failure signal). The weak answer describes it as "the model learns from its mistakes", which is the sentence an interviewer is probing behind.

  • Where do those notes live, and why not keep every one of them forever?
    They are episodic — held in a small bounded buffer for the current task and dropped when it ends. Keeping all of them costs context and, worse, keeps advice that a later change of approach has invalidated. An agent that abandoned an endpoint will still honour a lesson about that endpoint's pagination. Promoting a note into durable cross-session memory is a separate decision that needs deduplication and invalidation.
  • What if the only failure signal is "wrong answer", with no detail?
    Then the reflection is a hypothesis, not a diagnosis, and its quality drops sharply. The model will still produce a confident-sounding cause, and the retry will optimize against it — sometimes moving away from a nearly correct attempt. Where you can, make the signal localizing: which check failed, on which input, with what expected value. Rich signals are most of what makes this technique work.
  • Does any of this improve the underlying model?
    No. There is no weight update; the improvement lives entirely in the context window and evaporates when the session ends or the context is compacted. Two users hitting the same failure both pay for the same rediscovery. You can later distil successful trajectories into training data, but that is a separate offline process, not part of the loop.

saying these in an interview costs you the question

  • Says reflection updates or fine-tunes the model's weights
  • Describes it as resending the same prompt and resampling
  • Keeps every past reflection in context indefinitely
  • Assumes the self-diagnosis is correct because the model wrote it
  • Claims it lets an agent solve tasks the model cannot do

context

open as a page

Why does Reflexion split the agent into actor, evaluator and self-reflection roles?

level: middleimportance: should knowfreq 44%

basics

~20 s

Three different jobs with different failure modes: the actor produces the attempt, the evaluator judges it and returns a verdict or score, and the self-reflection step turns that thin verdict into concrete written advice the actor can act on next time.

open as a page

How do you tell that a self-refine loop has stopped improving output and is just changing it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Score every round against a signal outside the critic rather than trusting the critic's own satisfaction, keep the best-scoring version instead of the last one, and stop when the diffs become paraphrase-level churn or the critique starts repeating and contradicting itself.

open as a page

When does an agent's self-critique stop being evidence that its output is correct?

level: principalimportance: should knowfreq 33%

basics

~20 s

Self-critique stops being evidence when the critic shares the actor's blind spot or caves under pushback. A verdict that flips to "looks correct" after one objection is measuring agreeableness, not correctness, so calibrate the critic against seeded known-bad outputs before trusting it.

open as a page