A team quotes a monthly trend of failures from a scheduled promptfoo red-team job, and this month it dropped by half. The report does not say whether the suite was regenerated or whether responses were replayed from cache. What do you check to explain the drop, and what should each run record so the question is answerable next time?
answer
- denominator first: case count and identity
- duration + provider usage = was it replayed
- grader and judge model changed?
- target identity last
- record suite hash, cache mode, counts not just rate
basics
~20 sCheck whether the graded population changed: case count and case identity, then whether responses were served from cache, then whether the grader or its judge model changed. Only after those hold constant does a target change explain the drop. Each run should record suite identity, case count, target identity and cache mode alongside the number.
solid answer
~50 sWork outward from the number to everything that could have moved it, cheapest check first. 1. **The denominator.** Compare case counts and case identities against last month. A regenerated suite is the most common cause of a large jump in either direction, and it is visible immediately. 2. **Whether calls happened.** Compare run duration and provider usage. A near-instant run with negligible usage was replayed from promptfoo's cache and describes an older deployment. 3. **The grader.** The object deciding pass or fail is usually itself a model. If its judge model or its assertion text changed, the same responses get scored differently. 4. **The target.** Only now is "the application got safer" a supportable explanation — and it still needs the deployment identity to have changed in a way you can point to. The fix is provenance: every run emits suite hash, case count, target and deployment identity, grader configuration and judge identity, cache mode, and run duration, stored with the result rather than in a chat message.
go deeper
Knows to ask whether the same tests ran, rather than accepting the drop at face value.
Checks case count and cache state, and can explain why either alone accounts for a large move in the number.
Runs the checks in cost order including the grader and its judge model, then specifies the run-record fields — suite hash, target and deployment identity, cache mode, duration, counts — that make the next investigation unnecessary.
Makes provenance a property of the reporting system rather than of a diligent engineer, and rules that a trend line is broken at every suite change.
### Why this happens at all A red-team number acquires authority by being repeated. Six months on, the chart is quoted in a review by someone who never saw the job configuration, and nothing in the chart records that the suite behind month three is not the suite behind month six. The defect is not statistical noise; it is missing metadata about the instrument. Every layer of a promptfoo red-team run can move independently of the target, and each one moves the number. ### The diagnostic order, cheapest first **1. The denominator — did the graded population change?** Compare the case count and the case identities against last month's run. If the job runs `promptfoo redteam run`, generation happens on every invocation and the cases are model-written afresh each time, so a large move in either direction is expected from the suite alone. This is a one-line comparison and it ends most investigations. **2. Did calls actually happen?** Compare run duration and recorded provider usage. promptfoo reports token counts per result; a run served from its response cache shows near-zero usage and finishes in a fraction of the usual wall clock, and it describes whichever deployment answered originally, not today's. A partial hit is the nastier version: some cases replayed, some live, producing a number no single system ever generated. **3. The grader.** The object deciding pass or fail is usually itself a model — a model-graded assertion with a rubric and a judge model. If the rubric text was edited, or the judge model was upgraded underneath you (a provider default moving is enough), identical responses get scored differently. Diff the assertions and pin the judge model explicitly so this stops being a variable. **4. The target.** Only now is "the application got safer" supportable, and it still needs a deployment identity you can point at: a model version, an image tag, a prompt revision. "Same endpoint name" is not an identity. ### What it costs to answer versus to prevent Answering after the fact is expensive: an afternoon of one engineer's time, plus possibly a full cold rerun of the suite (hundreds to low thousands of metered calls, tens of minutes) to reproduce last month — and if the old suite was never persisted, the question is simply unanswerable and the point has to be dropped. Preventing it is a few dozen bytes written alongside each result. The asymmetry is the whole argument. ### What a defensible run record contains ``` suite: hash of the generated case file, case count target: provider, endpoint, model/deployment version identifier grader: assertion set version, judge model and its version runtime: cache mode, duration, request + token counts, promptfoo version result: absolute pass/fail counts, not only a rate ``` Stored with the result — in the eval output, in the artefact store, in the row that backs the chart — not in a chat message or someone's memory. ### Where the number misleads Two readings do most of the damage. First, **a rate hides its denominator**: failures falling from 60 to 30 is a halving whether the case count held at 600 or fell to 300, and only absolute counts make the difference visible. Report counts and the denominator alongside any percentage. Second, **a continuous line implies a continuous instrument**. Drawing one line across a regeneration invites the reader to interpret an instrument change as a behaviour change, and they will, because that is what a line means. Break the series at every suite change and annotate the break. A third, quieter one: even when everything checks out, the improvement is an improvement *on those cases*. Expect and pre-empt the follow-up question — does the gain survive a freshly generated exploratory suite? A drop that appears only on the pinned cases is a tuning artefact. ### What you check before crediting the application Confirm case count and suite hash match; confirm the run was cold and the provider recorded the expected traffic; diff the grader configuration and confirm the judge model version; then obtain the target's deployment identity from whoever owns it and confirm it actually changed. If all four hold, say what changed and by how much, in counts, and mark the deployment on the chart so the next reader does not have to repeat this.
- Case counts match exactly and the cache was cold. What is still left to rule out before crediting the application?The grader: the assertion text or the model judging it may have changed. Also confirm the target's deployment identity actually differs from last month, rather than assuming it does.
- How do you present a month where the suite was deliberately regenerated?Break the series. Mark the regeneration on the chart, report the old and new suite's counts for that month if you can run both, and never interpolate across the boundary.
- Why record absolute counts rather than only a failure rate?A rate hides the denominator. If the suite grew or shrank, the rate can move with no change in behaviour, and the record gives no way to notice.
Reporting a rate without its denominator is like saying the pass mark went up without saying the exam changed. Regenerating the suite mid-trend swaps the exam, and the line drawn through both papers is the part that lies.
saying these in an interview costs you the question
- Announcing the improvement before checking whether the suite changed
- Reporting only a rate, so a changed case count is invisible
- Drawing one continuous trend line across a suite regeneration
- Treating a chat message about the configuration as the run record
- Ignoring that the judging model may have changed underneath the grader