A promptfoo eval reads its system prompt from a file the product team edits in place, and each run records only the resulting pass rate. Why does the week-over-week trend become uninterpretable, and what do you change?
answer
- a path is not an identity
- record prompt version + provider + case set
- one axis moves at a time
- keep per-case outcomes, not just the rate
- prompt edit = re-baseline
basics
~20 sThe pass rate stops being a trend and becomes a coincidence: a drop could be the edited prompt, a change in the model behind the provider, or a different generated case set. Pin each prompt as a versioned artefact, record which version every run used, and move one axis at a time.
solid answer
~50 sA trend line only means something if exactly one thing changed between points. Here three things can move without anyone touching the config: the prompt text on disk, the model behind a hosted provider name, and — in red-team mode — the generated case set. When the run record stores only a number, a drop is unattributable and the usual outcome is that someone blames the model and nobody checks. The fix is to make each axis identifiable in the record. Treat the prompt as a versioned artefact with a content hash or a tag, and store that identifier with the run alongside the provider identity and the case-set identity. Then a drop can be read as 'prompt version changed, others fixed' or 'nothing we control changed', which is the far more interesting finding. Operationally: freeze the prompt for the eval, let the product team ship a new version explicitly, and re-baseline when they do.
go deeper
Should say the prompt must be pinned or versioned so runs are comparable.
Names all the things that can drift under a fixed config — prompt text, the model behind the provider, the generated case set — and stores an identifier for each with the run.
Adds per-case outcome diffs, deliberate re-baselining, and the overfitting risk of tuning a prompt against a fixed suite.
Sets the record-keeping standard across suites so any historical number can be reconstructed and read with its provenance.
### Why the trend line dissolves The entire value of running the same promptfoo matrix week after week is comparison, and comparison requires that each axis be **identified**, not merely present. A config entry such as `prompts: [file://prompts/system.txt]` names a *location*, and a location whose contents are edited by another team is not an identity. Two runs a week apart can share every line of the config and still have tested two different systems. Four things drift under a config that never changed: - **Prompt text**, edited in place by whoever owns product copy, usually for tone or length and usually without telling anyone that a safety-relevant clause moved. - **Provider behaviour**, because a hosted endpoint's underlying weights, system-side filtering or default sampling can change while the provider string you configured stays byte-identical. - **The case set**, when red-team cases are generated per run from `redteam.plugins` and `redteam.numTests` rather than checked in, so this week's suite is not literally last week's suite. - **Graded verdicts**, when an assertion is decided by a model (`llm-rubric` and friends) rather than by a deterministic rule; the judge is itself a drifting system. All four produce the identical symptom — the pass rate moved — and a run record that stores only the aggregate can distinguish none of them. The usual organisational outcome is that somebody blames the model, nobody checks, and the eval quietly stops informing anything. ### The mechanism of the fix Make each axis carry an identifier into the run record, and store per-case outcomes rather than only the aggregate. ```yaml prompts: - id: file://prompts/system.txt label: system-v7 # you set this; promptfoo does not fingerprint the file for you ``` Note the trap in that snippet. promptfoo's `label` is what the comparison view and the summary key results by, and it is a string you typed. Edit the file and leave the label alone and two different texts are recorded as one prompt version — the rendered text does land in the raw output JSON, but nothing compares it across runs on your behalf. So the pin has to be real: version the prompt as a content-addressed artefact (a commit sha or content hash in the label), and treat a text change as a new label. Alongside it, record the provider identity and, if cases are generated, a case-set identifier; then write results with `promptfoo eval --output results-<date>.json` so the per-case outcomes survive. ### What it costs Almost nothing: a label convention, a copied file, and a results artefact per run measured in kilobytes. The cost sits entirely on the other side — an unpinned suite means every historical number you have is uninterpretable, and there is no way to recover the attribution after the fact, because the text that produced last month's number no longer exists anywhere. ### Where the number misleads Two readings go wrong. First, the **confident misattribution**: the prompt was edited on Tuesday and the pass rate fell on Wednesday, so the model gets blamed, or the reverse. With one record per axis, the run reads as "prompt version changed, everything else fixed" — or, far more interesting, "nothing we control changed and it moved anyway", which is evidence the provider moved under you. Second, and more insidious: a **healthy number that is healthy because the prompt was tuned against this exact suite**. Pass rates that climb steadily over months while the case set is fixed usually mean somebody optimised the wording against the cases the suite happens to contain. Pinning versions makes this visible, because the rate rises exactly at the prompt edits, and you can then ask the only question that matters: does the improvement hold on cases the tuner never saw? ### What you check Diff per-case outcomes between the two result files rather than comparing aggregates — two runs with an identical rate can differ entirely in *which* cases passed, and an offsetting regression hides perfectly behind a flat number. Re-run any case that flipped several times (`promptfoo eval --repeat n --no-cache`) to separate stochastic variance from a real change. Move one axis at a time: when the product team ships new prompt text, label that run a deliberate re-baseline rather than a point on the old line, and if two axes moved in one week, re-run last week's prompt against this week's everything else — one extra column buys the attribution back.
- The prompt was pinned and the provider string never changed, yet the pass rate fell. What now?Diff per-case outcomes to see which cases flipped, re-run those cases a few times to separate stochastic variance from a real change, and treat a stable flip as evidence the model behind the provider moved.
- Why keep per-case outcomes rather than only the aggregate?Because two runs with an identical pass rate can differ entirely in which cases passed. The diff is the diagnostic; the aggregate hides an offsetting regression.
A file path is like a room number on a hotel door: it tells you where to knock, not who is inside this week. Recording that you tested room 402 says nothing about which occupant you tested.
saying these in an interview costs you the question
- Assumes a stable provider name implies stable model behaviour.
- Stores only the aggregate pass rate per run.
- Lets the prompt file change and the model change in the same week, then attributes the drop to one of them.
- Reports a rising pass rate as improved safety when the prompt was tuned against that exact suite.