skip to content

How does failure-mode analysis of wrong outputs drive the next prompt revision?

level: seniorimportance: should knowfreq 46%

answer

  1. the score says nothing about what to fix
  2. read failures, label them, count them
  3. group before you edit
  4. mode determines the class of fix
  5. re-count per mode after the revision

basics

~20 s

Read the wrong outputs, group them by hand into a few named failure modes, and count each. The biggest group sets the next revision's target, and each mode implies a different edit — or tells you the fix is not a prompt edit at all.

solid answer

~50 s

An aggregate score tells you a prompt is 78% right; it tells you nothing about what to change. So pull forty wrong outputs, read them, and label each with a short cause — "confuses escrow with taxes", "ignores the format rule on long inputs", "invents an account status not present in the input". Iterate the labels until they stabilise into five or six named modes, then count them. Now the next revision has a target instead of a hunch, and the mode names the intervention: category confusion wants a decision rule and a contrasting example pair; instruction decay on long inputs wants restructuring or repositioning; missing domain facts want retrieval or fine-tuning, not more prose. After the revision, re-count *per mode* on the same items — a change that halves one mode while doubling another is not an improvement, and only the per-mode counts reveal that.

go deeper

for a junior

Know the habit itself: when a prompt underperforms, read the wrong outputs and group them by cause before editing anything, rather than rewriting from intuition.

for a middle

Explain how the taxonomy is built and why counts matter — sample rather than cherry-pick, let labels emerge, consolidate to a handful of named modes, and target the largest one first.

for a senior

Show that the mode dictates the intervention, including recognising mislabelled gold data and missing-knowledge modes that no prompt edit can fix, and validate revisions per mode so a fix that trades one error for another is caught.

for a principal

Own the loop as a practice: a shared failure taxonomy that becomes the slice vocabulary in eval reports, prioritisation weighted by business consequence rather than raw counts, and a clear read on when prompt iteration has hit diminishing returns and a different lever is required.

## Why the score is not a work item "The prompt is 78% accurate" is a status report, not a plan. Two prompts at 78% can be failing for entirely unrelated reasons, and the edits that would help them have nothing in common. Error analysis — reading failures and organising them — is the step that converts a number into a decision about what to write next. It is unglamorous, it is manual, and it is the single highest-yield habit in prompt iteration, because almost everything else people do at this stage is guessing. ## Building the taxonomy The procedure is deliberately low-tech: 1. **Sample failures, don't cherry-pick them.** Take all the wrong items if there are few; take a random sample if there are many. Sampling only the ones that annoy you biases the taxonomy toward memorable cases. 2. **Read the input, the expected output and the actual output together.** The cause is often visible only in the input — an unusual phrasing, a second topic buried in paragraph four. 3. **Write a short free-text cause for each.** Do not start from a category list; let the labels emerge. 4. **Consolidate.** Merge synonymous labels into named modes. Thirty to fifty failures typically yield five or six modes, and you stop when new samples stop producing new labels. 5. **Count.** The distribution is the point. A mode holding 24 of 40 failures is where the next hour goes; a mode with one instance is noise until proven otherwise. Keep the taxonomy written down with an example per mode. It becomes the vocabulary the team uses in reviews, and the slice definition your regression suite reports against. ## Modes map to different interventions The reason the taxonomy pays off is that the mode determines the class of fix: - **Category confusion** — two labels whose definitions overlap in the prompt's wording. Fix with an explicit discriminating rule ("if both an escrow shortage and a payment change are mentioned, classify as escrow") plus a contrasting example pair showing the boundary in both directions. - **Instruction not followed on hard inputs** — the rule is present but loses out to a long or unusual input. Fix structurally: move the rule closer to the task, make it an explicit constraint rather than a passing remark, or simplify competing instructions. - **Format drift** — output shape degrades on some inputs. This is a formatting problem with its own toolkit rather than a knowledge problem. - **Under-specification** — the model is not wrong so much as the task never said what to do; the gold label encodes a convention the prompt never states. Fix by writing the convention down, and check the labels are self-consistent. - **Missing knowledge** — the input does not contain the fact and neither does the model. No amount of instruction fixes this; it wants retrieval, a tool, or a different data path. - **Bad gold labels** — a real and common mode. If 6 of 40 "failures" are actually correct outputs against wrong labels, your true accuracy is higher and your eval set needs adjudication before any prompt change. That last one is worth calling out in an interview: the first pass of error analysis frequently improves the *eval set*, not the prompt, and doing it in the other order wastes revisions chasing phantom failures. ## Prioritising by cost, not only by count Count ranks by frequency; production ranks by damage. A mode with four instances that sends a legal notice to an automated response queue outranks a mode with twenty instances that mislabels two interchangeable low-stakes categories. Weight each mode by its business consequence before choosing the next revision, and be explicit that you did — that reasoning is what an interviewer is listening for. ## Closing the loop A revision is validated per-mode, not in aggregate. After the edit, re-score the same items and re-count each mode. Three outcomes matter: - The targeted mode shrinks and the others hold — the fix worked. - The targeted mode shrinks and another grows — the new rule is over-firing, usually because it was written from three examples rather than as a general policy. Narrow its trigger. - Nothing moves — the edit did not reach the behaviour, and adding more words rarely helps. Escalate to a structural change or a different technique. Because a prompt edit can trade one mode for another silently, per-mode counts belong in the regression report alongside the headline number. And every mode you close is worth a permanent case in the suite, so that the next person to reorganise the prompt learns immediately if they have reopened it. ## Diminishing returns After several rounds the modes get small, heterogeneous and hard to name — a long tail of one-off oddities. That is the signal that prompt iteration has extracted what it can and the remaining error needs a different lever: better inputs, a different decomposition of the task, more capable models, or accepting the residual and designing the product around it.

  • How many failures should you label before you trust the taxonomy?
    Label until new samples stop producing new categories — typically thirty to fifty for a single task. What matters is saturation, not a fixed number: if the fortieth failure still introduces a mode you have not seen, keep going. Then rely on counts rather than impressions, because memorable failures are systematically over-weighted when you work from recollection.
  • How do you avoid fixing one failure mode and quietly creating another?
    Report per-mode counts, not just accuracy, before and after each revision on the same items. A new rule written from three examples often over-fires on cases it was never meant to touch, and the aggregate can stay flat while the error distribution shifts underneath. Narrow the rule's trigger when a neighbouring mode grows, and keep a permanent eval case for every mode you have closed.
  • What does it look like when error analysis says the problem is not the prompt?
    Failures cluster on inputs whose answer depends on a fact that appears neither in the input nor plausibly in the model's knowledge — an internal policy, a customer-specific term, a recent change. Restating the instruction cannot help. That mode points at retrieval, a tool call, or fine-tuning, and recognising it early saves several revisions of increasingly elaborate prose.

saying these in an interview costs you the question

  • Rewrites the whole prompt after skimming a few bad outputs
  • Cherry-picks memorable failures instead of sampling them
  • Judges the revision on aggregate accuracy alone
  • Assumes every wrong output is the model's fault, never the label's
  • Keeps adding instructions when the gap is missing knowledge

context