Several teams run promptfoo red-team suites on a schedule, each with its own written description of the application under test. How do you keep those descriptions from quietly rotting, and who should own them?
answer
- drift always flatters the score
- app team authors, security reviews
- co-locate with the code
- stamp scores with description version
- rising pass rate = investigate
basics
~20 sTreat each description as a reviewed, version-controlled artifact owned with the application. When an app gains a tool, role or data source and the text does not, promptfoo's generator stops attacking the new surface and its graders stop failing it, so the pass rate rises while real risk grows. Diff it every release.
solid answer
~60 sThe description is the specification the run is derived from, so drift in it is silent and always flatters you: new capability, no new cases, no new rules, an unchanged or improved score. **Ownership.** It belongs with the application team, who alone know when a tool or data source is added, with security as reviewer rather than author. A central security team writing every description scales badly and produces descriptions that are already stale. **Controls that actually work.** - Keep it in the app's repository beside the code that gives it capabilities, so a PR that adds a tool touches the file that describes it. - Make its review part of the change that adds a capability, not a quarterly audit. - Publish scores with the description version they were graded against; a pass rate with no attached description is not a reportable number. - Watch for the tell: a pass rate that improves across a release where nothing about safety changed usually means the description fell behind the app. Accept the residual: even a current description only bounds the run to what someone thought to declare.
go deeper
Says the description should be kept up to date and stored with the project rather than pasted ad hoc.
Explains that a stale description means new capabilities are neither attacked nor graded, and puts the file in version control next to the app.
Ties the review to capability-changing pull requests, stamps reported scores with the description version, and treats a description edit as breaking the trend line.
Sets the ownership split, makes an unstamped pass rate unreportable across teams, treats a rising score as a drift alarm, and states the residual limit that no description makes the number a safety proof.
### Why drift in this particular file is asymmetric `redteam.purpose` is the only input that decides both what gets attacked and what counts as a failure. That double duty is what makes staleness dangerous rather than untidy: any gap between the text and the deployed system moves the score in exactly one direction. A capability the description omits is never generated against, so it cannot lower the pass rate; and if some case reaches it anyway, there is no declared rule for the grader to fail it on, so it cannot lower the pass rate that way either. Drift always flatters. Nothing in promptfoo detects the mismatch — the tool has no view of your application beyond that paragraph. ### Where drift comes from A tool is added to the assistant. A second user tier appears. A retrieval corpus grows to include a new document class. A refusal is relaxed by a product decision nobody routed through security. The underlying model is swapped and the app now attempts things the old one declined. Every one of these changes the attack surface, and none of them touches a text file that no pull-request template mentions. ### Ownership Author: the application team. Reviewer: security. The rationale is currency versus judgment. The app team is the only party that knows on the day that a tool was added; security is the party that knows which declared refusals actually matter and which are decorative. Making security the author creates a queue and guarantees the descriptions are stale by construction; leaving security out entirely produces descriptions that quietly flatter the app, because the author is also the person the score reflects on. ### Mechanisms, in rough order of value 1. **Co-locate** the description with the code that grants capabilities, so the diff that adds a tool sits beside the diff that should describe it. 2. **Put a line in the change template** for anything that adds a tool, a role, a data source, or relaxes a refusal. 3. **Version it and stamp every reported score** with the description version it was graded against. An unstamped pass rate is not a reportable number. 4. **Read generated cases periodically**, not only scores. If no case in this quarter's scheduled run mentions a capability shipped last month, the description missed it — and that check catches drift that no diff review will. 5. **Treat a description edit as a suite change.** It invalidates the trend line; annotate the series rather than drawing a continuous line through it. ### What this costs Per application: a few minutes at each capability-changing pull request, plus a reviewer pass. The standing cost is subtler and worth budgeting for — every legitimate edit resets the comparable series, so a fleet with monthly runs and an actively developing app will rarely have more than a few comparable points in a row. That is the honest cost of an instrument whose rules move with the system, but it means you must stop asking teams for quarter-over-quarter deltas that cannot exist. The tempting saving — one shared description reused across applications to cut the maintenance — destroys the only property that makes any of the numbers useful, which is that the cases were written for that specific system. ### Where the numbers mislead Two readings, and both are common. First, a pass rate that improves across a release where no safety work shipped is far more likely to be drift than progress; good news from this instrument deserves more scrutiny than bad news, not less. Second, any fleet-wide average or leaderboard built on these pass rates mostly measures who wrote the most detailed purpose. The team with the most thorough description has the most declared rules to fail and scores worst, which creates a standing incentive to under-describe. Say that out loud when someone proposes an org-wide red-team scorecard, because the metric will be gamed by writing less, and writing less is exactly the failure the whole leaf is about. ### What you would check Diff the description against the release notes each cycle, and treat any capability in the notes that is absent from the text as an open finding. Sample the scheduled run's generated cases for mention of recently shipped surface. Refuse an unstamped score at the reporting boundary. Spot-check that each declared refusal still draws cases after an edit. And keep the residual limit stated in every report: this governance stops the number from becoming a lie, but it never turns it into a proof of safety.
- A team's scheduled pass rate improves month over month with no safety work shipped. What is your first hypothesis?The application grew capabilities the description does not mention, so the new surface is neither attacked nor gradeable. Diff the description against the release notes before celebrating.
- Why not auto-generate the description from the application's code and tool definitions?It can seed the capability half, but refusals, off-mission boundaries and role policy are human decisions no generator knows. Generated text also stops being read, which is how drift hides.
A stale description is an exam written from last year's syllabus: the course added three modules, the paper never did, and the class average keeps climbing while the students know less of the material.
saying these in an interview costs you the question
- Centralising authorship in security and expecting freshness.
- One shared description reused across several applications.
- Reviewing descriptions on an annual audit cadence rather than at capability changes.
- Reporting a pass rate without the description version it was graded against.
- Treating an improved score as good news without asking whether the description fell behind.