skip to content

How do you turn a production failure trace into a permanent regression test case?

level: middleimportance: should knowfreq 48%

answer

  1. capture, redact, assert, verify
  2. freeze the tool responses too
  3. assert the property, not the wording
  4. it must fail before the fix
  5. tag by failure mode, then prune

basics

~20 s

Freeze the trace: capture the inputs and the tool responses it saw, redact personal data, then write an assertion about the property that actually failed rather than the exact wording. Confirm the case fails before the fix and passes after, and tag it by failure mode so the suite stays prunable.

solid answer

~50 s

The loop has four steps. Capture: pull the full trace — user turns, retrieved context, tool calls and their responses — so the case can be replayed without live dependencies. Sanitise: strip or pseudonymise personal data before it lands in a repository, keeping the shape that mattered. Assert: decide what actually went wrong and encode that as a check, not an exact match on free text. If the assistant promised a refund it had no authority to promise, the assertion is that no unauthorised commitment appears, not that the reply equals a golden string. Verify: run it against the pre-fix build and confirm it fails, then against the fix and confirm it passes — an assertion that never fails is not protecting anything. After a bad week you might promote around forty traces this way, tagged by failure mode so duplicates can be collapsed and the suite does not bloat past its time budget.

code

yaml · 13 lines
yaml
id: unauthorized-refund-promise-001
source_trace: 8f2c41ae
failure_mode: unauthorized-commitment
input:
  user_turns:
    - "my order never arrived, I want my money back"
frozen_tools:
  lookup_order:
    status: delivered
    delivery_scan: left at front door
assertions:
  - not_contains_any: ["refund approved", "we will refund you"]
  - contains: "dispute"

go deeper

for a junior

Know the shape of the loop: a real failure gets captured, cleaned of personal data, and turned into a test that runs from then on. Say that the test must check what went wrong, not the exact wording of the reply.

for a middle

Explain each step and why it exists — freezing tool responses for determinism, property assertions instead of exact match, and confirming the case fails on the broken build before you trust it.

for a senior

Show the operational view: how promoted cases get tagged and deduplicated, how the suite's runtime and cost stay bounded, how intermittent failures are handled with repeated runs, and how the suite is wired into the gate for prompt, retrieval and model changes.

for a principal

Own it as an institutional asset: argue why a suite grown from your own incidents predicts real risk better than public benchmarks, who is accountable for promoting traces during an incident, and how coverage of failure modes is reported to people outside the team.

## Why the loop exists Every production incident in an LLM system is an expensive discovery: a case nobody imagined, found by real users. The waste is letting that discovery evaporate once the fix ships. The trace-to-eval loop is the discipline of converting each live failure into a permanent, automatically-checked case, so that a prompt tweak, model upgrade or retrieval change six months later cannot silently reintroduce it. It is the mechanism that makes a quality programme cumulative rather than reactive. ## Step 1: capture enough to replay A replayable case needs everything the system saw, not just what the user typed: - the full sequence of user turns, in order - the system instructions and prompt version in force - whatever was retrieved or injected as context - **every tool call and the exact response it received** - the model and configuration identifiers The tool responses are the part teams most often skip and most often regret. If the case calls a live service at replay time, it will drift: the order that was undelivered in July is closed by September, and the test becomes flaky or quietly stops exercising the failure. Freeze the responses with the case so it is hermetic. ## Step 2: sanitise before it becomes an artefact A production trace contains real people. Once promoted into a repository it will be copied, read in CI logs and possibly sent to a third-party model. Redact or pseudonymise names, addresses, account and payment identifiers, and free text that could identify someone — while preserving the structural features the failure depended on. If the bug only reproduces when an account number has an unusual format, keep the format and change the digits. Record provenance (source trace id, date, incident) so the case can be traced back internally even after the content is scrubbed. ## Step 3: assert the property, not the transcript This is where the loop most often goes wrong. Free-text output is not stable across model versions, so an exact-match assertion against a golden reply will fail on the next upgrade for reasons unrelated to the bug, and the team will delete it. Instead, name the defect and check for it directly. Useful assertion shapes: - **Must-not** — the reply does not contain an unauthorised commitment, does not disclose another customer's data, does not invent a policy. - **Must** — the reply cites at least one retrieved document, or offers the escalation path, or refuses. - **Structural** — the response validates against the required schema; the tool was called with the correct arguments. - **Programmatic verification** — where the task has a checkable end state, assert that state rather than the prose. The test is: would this assertion have caught the original failure, and would it stay quiet on a correct answer worded differently? If the second half fails, the assertion is too tight. ## Step 4: verify the case has teeth Run the new case against the build that produced the failure. It must fail. This one step separates a real regression case from decoration: an assertion that passes on the broken build is testing something other than the bug. Then run it against the fix; it must pass. Only then does it join the suite. ## Step 5: tag, batch and prune Cases arrive in bursts. After a bad week you might promote around forty traces at once. Without structure that becomes an undifferentiated pile: - **Tag by failure mode** — unauthorised commitment, retrieval miss, tool-argument error, loop, premature termination. Tags let you report which failure families are covered and collapse near-duplicates. - **Deduplicate.** Twenty traces of the same failure mode add runtime and no information; keep the two or three that differ meaningfully in trigger. - **Budget the suite.** Cases cost inference time and money on every run. Track total runtime and cost, and prune cases whose failure mode is now covered by a cheaper deterministic check. - **Handle nondeterminism.** A single pass can pass by luck. For cases where the failure was intermittent, run repeats and require consistent success rather than one green result. ## What closes the loop The loop is only closed when the case runs automatically on every candidate change and its failure blocks the change. A suite that is run manually before big releases is a checklist, not a regression barrier. Wire it into the same gate that guards the prompt, the retrieval configuration and the model version — those are the three things most likely to resurrect an old failure. ## The payoff Six months in, the suite is a written record of every way this system has actually broken in front of users — far more predictive of real risk than any generic benchmark, because every case is drawn from your traffic, your tools and your policies. That is the argument for spending an hour promoting a trace on the day of the incident, when the context is still fresh and someone still remembers exactly what went wrong.

  • Why freeze the tool responses instead of calling the real services during replay?
    Because live services change underneath the test. The order that was undelivered when the failure happened gets resolved, the knowledge base article gets edited, and the case stops reproducing the condition that triggered the bug — usually without anyone noticing, since it still passes. Frozen responses make the case hermetic and deterministic, so a failure means the system changed, not the world.
  • What is wrong with asserting that the reply exactly matches a known-good answer?
    Free-text output legitimately varies across model versions, temperatures and prompt edits, so exact-match cases fail constantly for reasons unrelated to the defect. Teams then mute or delete them and the coverage is lost. Assert the property that failed — no unauthorised commitment, a citation present, the right tool called — so the case tolerates rewording while still catching the bug.
  • How do you stop a growing regression suite from becoming too slow and expensive to run?
    Tag every case with its failure mode and treat the tags as the coverage unit rather than raw case count. Collapse near-duplicates that share a mode and trigger, retire cases whose mode is now caught by a cheap deterministic check, and track suite runtime and cost as first-class numbers with a budget. Coverage of distinct failure modes matters; case volume does not.

saying these in an interview costs you the question

  • Asserting exact string equality against a golden reply
  • Replaying against live tools instead of frozen recorded responses
  • Promoting traces without confirming the case fails pre-fix
  • Copying raw production traces with customer data into the repository
  • Adding every failing trace without deduplicating by failure mode

context