skip to content

Your promptfoo red-team suite costs too much to run on every merge, so provider-by-prompt cells have to go. How do you decide which pairings survive, and what can the trimmed run no longer tell you?

level: seniorimportance: must knowfreq 40%

answer

  1. tier by cadence, do not delete
  2. shipped pairing every merge
  3. past failures pinned permanently
  4. rotation loses attribution and dating
  5. interaction cells die first

basics

~20 s

Keep the pairing you actually ship on every run and rotate the rest, so each provider and each prompt appears somewhere without crossing them all. Then state the loss out loud: a rotated matrix says a failure happened, not which provider-and-prompt pairing owns it. Re-cross the full matrix before a release.

solid answer

~50 s

Split the suite by cadence rather than deleting cells outright. On every merge, run the **shipped pairing** — the provider and prompt version actually in production — against the full case set. That is the run whose result maps to a real deployment, and it must never be the part you trim. On a slower cadence (nightly or weekly), rotate the remaining pairings so every provider and every prompt variant is exercised regularly even though they are never all crossed at once. Before a release or a model swap, run the full cross. The cost of the trim is attribution. In a rotated matrix a given pairing appears rarely, so when a failure shows up you cannot immediately say whether the provider or the prompt owns it, and you cannot date the regression to a specific merge. Write that limitation next to the pass rate, because the number will otherwise be read as if the whole matrix ran.

go deeper

for a junior

Should at least say to keep the configuration you actually ship and to run the rest less often, not never.

for a middle

Proposes cadence tiers, protects a regression set of previously-failing cases, and names cost per cell as the trim criterion.

for a senior

Names what the trim costs — attribution, dating, and interaction pairings — and insists the pass rate is reported with its coverage.

for a principal

Owns the policy: which tier gates a release, who may move a cell between tiers, and how coverage is communicated so 'green' is never over-read.

### The real decision Trimming a promptfoo matrix is not optional at merge cadence — a few hundred cells with model-graded assertions is minutes of wall-clock and real money on every push. Trimming is fine. Trimming **silently** is the defect, because the pass rate that comes out the other side keeps the same name, the same units and the same dashboard tile while quietly standing for something smaller. ### Mechanism: tier by cadence, do not delete A cell deleted from `promptfooconfig.yaml` is gone and nobody notices its absence six months later. A cell moved to a slower tier still runs, just less often, and the tiering is visible in the repo. Concretely, keep one config and select subsets at run time — `promptfoo eval --filter-providers <regex>` restricts which providers run, `--filter-pattern` restricts which cases run by description, and `promptfoo eval --filter-failing <previous-results.json>` re-runs only the cells that failed last time, which is the cheapest high-value run available to you. Three rules keep the tiering honest: **Protect the shipped pairing.** Whatever else moves to a slower tier, the provider and prompt version actually in production keep the full case set on every merge. That is the only cell whose result maps onto a real deployment; it must never be the part you trim because "it always passes". **Pin the regression set.** Any case that has ever failed stays in the per-merge tier permanently, regardless of which pairing it belongs to. These are the failures you already know how to produce; losing them to a rotation is how a fixed bug comes back unobserved. **Cut the widest axis first, and cut cheap information before scarce information.** Providers you do not serve, prompt variants nobody would deploy, and near-duplicate cases inside one behaviour family are cheap cuts. Sampling the case set uniformly across all families is the expensive cut, because it thins the rare tail behaviours a red-team suite exists to find in exact proportion to the common ones it would have found anyway. ### What the trim costs you, in order of how often it bites - **Attribution.** With pairings rotated rather than crossed, a failure cannot be assigned to the model or to the wording without a follow-up run. - **Dating.** A weekly-tier failure could have been introduced by any merge that week, so the bisect becomes manual work rather than a glance at the run history. - **Interaction.** The pairing that only fails *in combination* — this prompt on that provider — is the first thing a budget trim removes, because it looks redundant precisely when both halves pass everywhere else. That cell is also the most interesting one in the matrix. ### Where the number misleads The dangerous artefact is a **denominator swap that is invisible in the units**. The full matrix reported 96%; the trimmed matrix reports 96%; they are percentages of different sets of cells, and nobody comparing the two tiles can see that. Worse, trimming tends to remove cells that were failing (unshipped providers often fail more, being untuned for your prompt), so the trimmed rate drifts *upward* for a reason that has nothing to do with the system getting safer. The other misread is the word "green". "Suite green" is read by everyone downstream as "the matrix passed". After a trim it means "the shipped pairing passed the full case set this merge; the other three pairings last ran on Tuesday", and only the second sentence bounds the claim. ### What you check Before shipping the trimmed config: confirm the shipped pairing still runs the full case set; confirm every previously-failing case is in the per-merge tier; count the cells in each tier and record them. After each run: print the tier name, the cell count and the last-run date of every slower tier immediately beside the pass rate, so no reader can consume the number without its coverage. Then re-cross the full matrix before a release or a model swap — the full grid is a release gate, not a per-merge one.

  • Why is uniformly sampling the case set a worse cut than dropping an unshipped provider?
    Uniform sampling thins every behaviour family including the rare ones, which is where red-team value concentrates. Dropping a provider nobody serves costs a comparison you were not going to act on.
  • How do you stop a trim from being forgotten?
    Keep the trimmed cells in the config as a slower tier rather than removing them, and print the tier and last-run date beside the pass rate so every reader sees the coverage the number actually has.

saying these in an interview costs you the question

  • Deletes cells from the config rather than moving them to a slower tier.
  • Trims the shipped pairing because it 'always passes'.
  • Samples cases uniformly across behaviour families and calls the result equivalent coverage.
  • Reports the trimmed suite's pass rate with no statement of what ran.
  • Claims a trimmed matrix still tells you whether the model or the prompt is at fault.

context