You own a pinned, committed promptfoo red-team suite that several teams run against different applications. How do you decide how often it is regenerated, and what do you do so numbers reported before and after a regeneration remain meaningful?
answer
- suite = the ruler, version it
- triggers: capability, catalogue, overfit evidence, long-stop
- bridge run: old and new, same day
- shared core vs per-app cases
- break the line, never interpolate
basics
~20 sRefresh on events, not on a calendar alone: a new application capability, a catalogue update, or evidence that scores are being tuned against the frozen cases. Treat each refresh as a reviewed commit, run old and new suites together once to bridge the series, and break the trend line at the boundary rather than interpolating across it.
solid answer
~50 sTwo failure modes bracket the decision. Refresh too rarely and the suite becomes a target: prompts get tuned until those specific cases pass, coverage stops reflecting what the catalogue now generates, and the trend flatters everyone. Refresh too often and no two months are comparable, so nobody can attribute a change to anything. I set refresh triggers rather than a bare cadence: a material change in what the applications can do, a meaningful catalogue or tool change, a suspicion of overfitting raised by a fresh exploratory run outperforming the pinned one, and a long-stop interval so drift cannot accumulate forever. The continuity mechanism is a **bridge run**: on the refresh, run both the outgoing and incoming suites against the same targets on the same day. That gives one dated point in both series, so the two series can be read together honestly without pretending they are one line. Sharing across teams needs a shared core plus per-application cases; only the shared core is compared across teams.
go deeper
Understands that regenerating the suite changes what is being measured and that this should not happen silently.
Argues both sides — staleness versus comparability — and proposes a reviewed, versioned refresh rather than an automatic one.
Defines event triggers, requires a bridge run so the series can be read across the change, and separates shared from per-application cases.
Owns the instrument as a versioned artefact across teams: adoption policy, what may and may not be compared, refusal of blended cross-team scores, and acceptance that some history ends rather than being fudged.
### Frame it as versioning a measuring instrument The pinned suite is not test code that should track the application; it is the ruler. Any change to a ruler invalidates comparison with measurements taken before it, unless you calibrate across the change. Everything below follows from that one framing, and the framing is what separates a policy from a preference. ### Refresh triggers, in the order they usually fire - **Capability change.** The applications gained a tool, a retrieval source, a new user-facing surface. Cases generated before it cannot reach it, so the suite's coverage claim is now false regardless of its age. - **Catalogue or tooling change.** promptfoo's plugin and strategy catalogue gains categories and transformations the pinned file never sampled. A suite generated a year ago tests a smaller space than the same configuration would test today. - **Overfitting evidence.** A freshly generated exploratory run against the same target finds failures the pinned suite no longer catches. This is the strongest trigger because it is direct evidence rather than a proxy, and it is the only one that fires on the actual failure you care about. - **Long-stop age.** A maximum interval, so a quiet quarter does not silently freeze the instrument for a year. A bare calendar cadence without triggers refreshes when nothing changed and fails to refresh when everything did. ### Bridging, and what it costs On refresh day, run the outgoing and the incoming suite against the same targets. The old series ends with a dated point that has a partner in the new series, so a reader can see the level shift the instrument itself caused and subtract it from any application change. Without that, the first point of the new series is uninterpretable — and it is precisely the point someone will quote. The bill is concrete: a bridge run is roughly double a normal run's calls, once, at each refresh, plus the generation calls for the new suite. With several teams and a quarterly refresh that is a handful of extra full runs per quarter — for a several-hundred-case suite, a four-figure call count each — plus a reviewer's hour per refresh. That is the honest price of a comparable trend line, and it is small enough that "too expensive" is never the real objection; the real objection is that nobody scheduled it. ### Sharing across teams One suite shared verbatim across different applications is the wrong shape. Cases that assume a capability an application does not have are unreachable there, and its failure rate is depressed for a reason that has nothing to do with safety. Split the artefact: a **shared core** covering cross-cutting behaviour, versioned centrally, and **per-application cases** each team generates against its own description. Only the shared core is ever compared between teams, and only between teams on the same core version. Let teams adopt a new core on a schedule they can absorb, with a bridge run required at adoption. ### Where the number misleads Three readings, all common. **The first point after a refresh** looks like a step change in safety and is usually a step change in the ruler. **A blended cross-team score** averages rates whose denominators differ in reachability, so it moves when the mix of teams changes and nobody notices. **A rising pinned score with no exploratory check behind it** is the overfit signal presented as good news — the more stable the suite, the more it is worth someone's while to tune to it, so the trend is most flattering exactly when it is least trustworthy. ### What you check Before a refresh: run an exploratory generation and record how far it diverges from the pinned suite, so you know whether you are refreshing for coverage or for drift. At the refresh: confirm the bridge run happened on the same day against the same deployment, and record both suites' hashes and case counts. After: confirm the chart shows a break, not a line, and that any cross-team comparison names the core version both sides ran. ### What the policy refuses, and what it accepts Refused: a single blended cross-team score, a continuous line drawn across a regeneration, an automated regeneration committed without review, and a refresh triggered by an unwelcome number — that last one makes the instrument a function of the desired answer. Accepted: that refreshing costs a bridge run, and that some historical comparisons simply end. Ending a series honestly is far cheaper than defending a number nobody can reconstruct.
- What concrete evidence tells you a pinned promptfoo suite is being overfit rather than the application genuinely improving?A freshly generated exploratory run against the same target still finds failures at the old level while the pinned suite's failures fall. The gap between the two is the overfitting signal.
- Why not simply keep every historical suite and rerun them all each month?Cost and noise. Every extra suite is metered calls, and a stack of stale suites mostly re-measures behaviours already covered. A bridge run at each refresh buys the continuity without the ongoing bill.
- How should a shared core suite be adopted by teams on different schedules?Version it, let each team adopt explicitly, and require a bridge run at adoption. Cross-team comparison is only valid between teams on the same core version.
saying these in an interview costs you the question
- A pure calendar cadence with no event triggers
- Automated regeneration committed without review, then charted as one continuous line
- One verbatim suite shared across applications with different capabilities, compared as a single score
- Refreshing whenever a score looks bad
- No mechanism at all for detecting that the pinned suite is being tuned against