An AI-powered feature returned a wrong answer. Which cheap experiments separate the possible causes, and in what order do you run them?
answer
- cheapest and most decisive first
- deterministic checks before probabilistic ones
- the first rungs need no model call
- one variable per experiment
- compare rates, never single outcomes
basics
~20 sRun the cheap deterministic checks first: replay the recorded call, compare the shown answer with the returned one, and read the instructions and material exactly as sent. Only then vary one layer at a time and re-measure the failure rate.
solid answer
~50 sOrder experiments by cost and by how much they can rule out. 1. **Replay the recorded call unchanged**, twenty or so times, to learn whether the failure is systematic or occasional. 2. **Compare what came back with what was shown** — a mismatch ends the investigation in ordinary product code. 3. **Read the instruction text and material exactly as sent.** Missing, stale or truncated material is visible without calling the generative model at all. 4. **Hand the needed fact directly to the step.** If the answer becomes correct, the failure is in how material is assembled; if not, wording or ceiling remains. 5. **Make one minimal wording edit** and re-measure on the same fixed inputs. 6. **Restore any setting that differs from the tested one.** Each experiment must be able to change your mind, must change exactly one thing, and must be judged on a rate rather than a single attempt.
code
pseudocode · 19 linesREPEATS = 20
baseline = failureRate(repeat(recordedCall, REPEATS))
ladder = [
{ name: "rendering", probe: compare(recordedCall.returnedText, complaint.shownAnswer) },
{ name: "material", probe: assertPresent(neededFact, recordedCall.suppliedMaterial) },
{ name: "material", probe: rerun(recordedCall with neededFact injected) },
{ name: "instructions", probe: rerun(recordedCall with oneWordingEdit) },
{ name: "configuration", probe: rerun(recordedCall with testedSettings) },
{ name: "task shape", probe: rerun(recordedCall split into twoSmallerAsks) }
]
for rung in ladder: // cheapest first, one change each
rate = failureRate(repeat(rung.probe, REPEATS))
record(rung.name, baseline, rate)
if rate materially below baseline:
return owner = rung.name // this layer owns the ticket
return owner = "nothing available moved the rate"go deeper
Be ready to say what you would look at first: the exact request as it was sent, and whether the answer the user saw matches what came back. Those two checks cost nothing and often finish the investigation on their own.
Explain the mechanics of a clean experiment on a feature whose output varies: fix the input, change exactly one layer, repeat enough times to compare failure rates, and record the outcome of each attempt as evidence.
Demonstrate ordering under real time pressure. Rank experiments by cost and by how many candidates they eliminate, keep the deterministic ones ahead of the probabilistic ones, and stop as soon as a layer's change moves the rate repeatably.
Own the discipline across the team. Argue for an investigation order that is written down rather than improvised, for evidence retained per attempt, and against the culture where wording is edited until the complaint stops and nobody can say what fixed it.
## Order by cost, then by discriminating power Investigating a wrong answer from an AI-powered product feature is a search over five candidate layers: the ordinary product code around the generative step, the instruction text sent with the request, the material supplied alongside it, the configuration in force, and the capability ceiling of the generative model itself. A good order is not the order of suspicion; it is the order that eliminates the most candidates for the least effort, and it puts every deterministic check ahead of every probabilistic one. Two properties decide the ranking. **Cost** is time, money and risk: reading a recorded request costs nothing and touches no users, while shipping a reworded instruction to production costs a release and puts the change in front of people before it is understood. **Discriminating power** is how many candidates an outcome removes: a check that can only ever confirm what you already believe is not an experiment. The best early moves are cheap *and* decisive, and the worst early move is the tempting one, which is to rewrite the wording and see if the complaint stops. ## The ladder | Experiment | Cost | If it comes out clean | If it comes out dirty | | --- | --- | --- | --- | | Replay the recorded call unchanged, many times | Very low | The failure is occasional; you are chasing a rate, not a defect | The failure is systematic and one layer is likely responsible | | Compare the returned text with the answer the user saw | Very low | The deterministic shell is exonerated | The investigation ends here, in ordinary product code | | Read the instruction text and material exactly as assembled | Very low | The request was well formed | Missing, stale or truncated material is the finding, with no model call needed | | Place the needed fact directly into the supplied material | Low | Wording or ceiling remains in play | The assembly of material owns the failure | | Make one minimal wording edit and re-measure | Medium | Wording is exonerated for these inputs | Instruction text owns it, and the edit is now testable | | Restore each setting to its tested value | Medium | Configuration is exonerated | An untested setting reached users, which is its own finding | | Decompose the task into two smaller asks | Medium | The task shape, not the model, was the obstacle | The ceiling explanation survives another attempt | The first three rungs need no call to the generative model at all, which is why they come first: they are free, they are repeatable, and on a mature feature they resolve a large share of complaints on their own. ## What makes an experiment honest - **Fix the input.** Use the literal text the user sent, not a paraphrase. A paraphrase is a second change you did not intend to make. - **Change one thing.** Two simultaneous changes that fix the failure leave you unable to say which one mattered, and you will carry both forever. - **Judge on a rate, not an outcome.** Because the generative step can return different text for identical input, an experiment needs enough repeats that a change in the failure rate is distinguishable from noise. Ten to thirty repeats on a small fixed input set is usually enough to tell a large change from nothing. - **State in advance what would change your mind.** Writing down the expected result before running the experiment is what stops a marginal improvement from being read as proof. - **Keep the results.** Each rung's outcome is the evidence that the eventual attribution rests on, and a later reader will want to know what was already ruled out. ## Knowing when to stop Stop as soon as one layer's change moves the failure rate materially and repeatably: that layer owns the ticket, and the experiment that found it is also the test that protects the fix. Stop for a different reason when the ladder is exhausted with the rate unmoved, because that is the only honest route to concluding that no change available to the team fixes the behaviour. The failure mode of an undisciplined investigation is not that it reaches the wrong answer; it is that it reaches an unfalsifiable one. Wording gets edited three times, material gets enriched, a setting gets nudged, the complaint stops, and nobody can say what fixed it. Six weeks later the same complaint returns with none of the changes safe to revert. An ordered ladder costs an hour more up front and leaves behind a fix whose owner, mechanism and regression test are all known.
- Why place the first three experiments before any call to the generative model?They are free, repeatable and fully deterministic, and each can end the investigation outright. Comparing the returned text with the shown answer convicts the product code; reading the material as assembled convicts the assembly. Spending a probabilistic experiment before those is paying for a noisy result when a certain one was available.
- How many repeats does an experiment on this feature need?Enough that a change in the failure rate is larger than the variation between two identical batches. In practice, run the unchanged call and the changed call at the same repeat count on the same fixed inputs, and only believe a difference that is large and reproduces on a second batch. A single passing attempt is not a result.
saying these in an interview costs you the question
- Starts by rewriting the wording and shipping it
- Changes several layers at once and calls the result a fix
- Accepts one passing attempt as evidence a change worked
- Runs the expensive experiment before the free deterministic ones
- Never writes down what result would change their mind