A long-running promptfoo red-team matrix has one provider replaced by a different one. What happens to the pass-rate history, and how do you stop a team that watches that trend from drawing the wrong conclusion?
answer
- pass rate belongs to a pairing
- two series, not one line
- paired run on a frozen case set
- thresholds fitted to the old system
- composition shifts under a flat aggregate
basics
~20 sTreat it as a new series, not a continuation. The old points describe a provider you no longer call, so mark the break, keep the retired provider's last full run as the baseline, and re-run a frozen case set against the new provider before anyone reads a trend across the swap.
solid answer
~50 sA pass rate is a property of a pairing, not of the product. Swap the provider and every earlier point was measured on something you no longer ship, so plotting them as one line invites two wrong readings: a drop looks like a regression the team caused, and a rise looks like the swap improved safety when it may only have shifted which behaviours fail. The handling has three parts. **Mark the break** in whatever renders the trend, so the discontinuity is visible without reading a changelog. **Back-fill a comparison**: re-run a frozen case set against both the retired and the new provider once, on the same prompt version, so you have one honest paired measurement rather than an inferred one. **Re-baseline**: declare the first post-swap run the new zero, and reset any threshold that was tuned on the old series. And say what changed qualitatively, not only by how much — which behaviour families moved.
go deeper
Should recognise the results are not directly comparable after the provider changes.
Says to re-baseline and to re-run the suite against the new provider before trusting the trend.
Runs both providers on a frozen prompt and case set, diffs per-case outcomes, and resets any gate threshold fitted to the old series.
Adds the governance: who may change a standing suite's axis, what they owe when they do, and how the break is annotated so historical numbers stay readable months later.
### Why the line breaks Everything a promptfoo red-team suite reports is conditional on the **pairing under test** — this provider, this prompt version, this case set, this grader. A pass rate is a property of that pairing, not of your product. Replace an entry in the `providers:` list and every earlier point on the chart was measured against a system you no longer call. Plotting the old and new points as one line is not a small presentational sin; it silently invites two specific wrong readings. A drop looks like a regression the team introduced and triggers a hunt through merges that contain nothing. A rise looks like the swap improved safety, when it may only have changed *which* behaviours fail. There is a mechanical casualty too. Any gate threshold derived from the old series — "block the release below 92%" — was fitted to a system that no longer exists. Carried across the swap it will either wave regressions through or block everything, and either way somebody will eventually turn the gate off rather than debug it. ### What to actually run One **paired run** is worth more than a quarter of unpaired history. Freeze the prompt version and the exact case set (check the generated cases in rather than regenerating them), run both the retiring and the incoming provider, and write per-case outcomes for both with `promptfoo eval --output`. That gives you the delta as a *set of behaviours that moved* rather than a single number, and behaviours are the only form a team can act on. Expect the interesting result to be a **composition swap**: the new provider holding cases the old one failed while failing others the old one held, with a nearly unchanged aggregate. An aggregate-only comparison reports "no change" and hides the entire finding. ### What it costs One extra column: providers x prompts x cases for the retiring provider, once, plus grader calls. On a typical matrix that is minutes and single-digit-to-tens of dollars — trivially less than the cost of a quarter of decisions made on a broken trend. The larger real cost is engineer time: someone has to read the per-case diff and characterise it in a sentence, and that sentence is the deliverable. Note also that the swap changes the **cost per run** itself. A pricier or slower incoming provider silently changes what cadence your budget buys, so the swap is a schedule decision as well as a safety one. ### Where the number misleads The trap that catches experienced teams: **a rising pass rate can be an over-refusal regression**. A model that refuses more readily scores better on an adversarial suite by construction — the suite asks whether the target did something it should not, and a target that does nothing passes everything. Reported alone, "88% to 97% after the swap" reads as a safety win and may be a product disaster. The control is a benign, in-scope set of requests the assistant is *supposed* to answer, run beside the adversarial suite; if the pass rate rose and the benign set's answer rate fell, you bought your score with utility. Two further misreads: an unchanged aggregate read as "neutral" without a per-case diff, and any comparison of the new provider against the *old history* rather than against a paired run — the history was collected under other prompt versions and, if cases were generated, other cases. ### What you check, and who owns it Annotate the break in whatever renders the trend, so the discontinuity is visible without reading a changelog; keep the old series visible but visually separated rather than deleting it, because it is the baseline you need to characterise the swap. Restate the gate threshold as provisional until enough post-swap runs exist to fit one. Write one sentence of interpretation beside the first post-swap number naming which behaviour families moved. If cases are generated per run, say so — then even the post-swap points are not strictly comparable to each other. The organisational half is the part that actually holds: decide in advance who may change an axis of a standing suite and what they owe when they do — the paired run, the annotation, the threshold reset, the over-refusal control. Without that rule, provider swaps happen for cost or availability reasons, the eval re-baselines itself in silence, and six months later nobody can say which of the numbers on the wall were measured against what.
- The aggregate pass rate is identical before and after the swap. Is that reassuring?Not on its own. Identical aggregates routinely hide an exchange of failures — new cases failing, old ones fixed. Diff per-case outcomes and compare behaviour families before calling it neutral.
- Can you avoid the paired run by comparing the new provider against the old history?Only weakly. The history was collected under other prompt versions and, if cases are generated, other cases. The paired run costs one column and removes all of that doubt.
Swapping the provider is like changing the scale under a patient halfway through a weight chart: the line continues across the page, but the two halves were never measured by the same instrument.
saying these in an interview costs you the question
- Plots pre- and post-swap runs as one continuous trend.
- Keeps a gate threshold that was tuned on the retired provider.
- Reads an unchanged aggregate as 'the swap was neutral' without diffing per-case outcomes.
- Deletes the old series instead of marking the break, destroying the history.
- Lets the prompt version change in the same run as the provider swap.