You have budget to widen a promptfoo eval by exactly one axis: a second provider or a second prompt variant. What does each one buy you, and how do you choose?
answer
- provider axis = model control
- prompt axis = wording control
- buy the attribution you can act on
- an added axis is permanent
- shrink cases to afford both
basics
~20 sA second provider tells you whether a weakness lives in the model or in your prompt. A second prompt variant tells you whether your wording is what is holding. If you ship on exactly one provider, add the prompt variant; if you are still choosing a model, add the provider. Either one doubles the run.
solid answer
~50 sBoth axes cost the same — they double the cells — but they answer different questions. Widening the **provider** axis is a control on the model. If the same prompt fails on one provider and holds on another, the weakness is model-side and your mitigation belongs in the prompt or in a guard. That is worth paying for while you are still selecting a model, or when you serve traffic across more than one. Widening the **prompt** axis is a control on your own text. Two variants of the same system prompt against one provider tell you how much of your safety margin comes from wording — which is the fragile part, because product teams edit that text. The deciding question is what you can act on. If you ship on one provider and cannot change it this quarter, a second provider produces a number you will not act on; the prompt variant produces one you can. Reverse it during model selection.
go deeper
Should notice that either addition doubles the number of calls and that they are not interchangeable.
Explains that the provider axis attributes a failure to the model and the prompt axis attributes it to wording, then picks by what they ship.
Adds that an axis is permanent cost, and proposes shrinking the case set so both attributions stay affordable at the cadence they need.
Decides which axes the org standardises on across suites, and who is allowed to add one given the cadence it costs.
### Two lists that look symmetric and are not In `promptfooconfig.yaml` the `providers:` list and the `prompts:` list are structurally identical: both are arrays, both get crossed with everything else, and adding an entry to either one costs the same — one extra column of cells, on every run, forever. What differs is the **conclusion each one licenses**. A red-team suite exists to attribute failure, and these two axes attribute it to different things. ### The mechanism: a small factorial design With one provider and one prompt you have a single measurement and no attribution at all. A failure could belong to the model, to your system prompt's wording, or to the case itself, and the result cannot separate them. - Widening **`providers:`** holds your prompt fixed and varies the model. If the same prompt fails against one provider and holds against another, the weakness is model-side, and your mitigation belongs in the prompt, in a guardrail, or in the choice of vendor. - Widening **`prompts:`** holds the model fixed and varies your own text. Two variants of the same system prompt tell you how much of your safety margin is carried by wording — which is the fragile part, because product teams edit that text for readability, tone and length without ever consulting the eval. You cannot buy both attributions for one axis' price. So you buy the one whose answer changes a decision you are actually able to make. ### What it costs | choice | cells added | recurring? | what it buys | |---|---|---|---| | second provider | prompts x cases | yes, every run | model-vs-prompt attribution | | second prompt variant | providers x cases | yes, every run | wording-vs-model attribution | The recurring column is the one people forget. An axis added for one investigation is permanent until somebody deliberately removes it, and it halves the cadence you can afford at the same spend. There is engineer time in it too: every extra column is more cells to triage, and triage — not API spend — is usually the scarcer budget. ### Where the number misleads The sharp failure here is the **pooled pass rate**. Once a second provider is in the matrix, promptfoo's headline figure averages cells from a system you serve with cells from a system nobody calls. That single number now describes no deployed configuration. It can flatter you (a refusal-happy comparison model lifts the average) or damn you (a weak comparison model drags it down), and in both directions the direction of the error is invisible from the number alone. Report per-cell or per-provider rates, and let the pooled figure exist only as a rough activity indicator. The second misread is the **strawman variant**. A second prompt written to fail — terse to the point of absurdity, or missing the safety section entirely — produces a large, satisfying gap that means nothing, because no team would ship that text. A prompt variant is only informative if it is something somebody would actually deploy: the previous release's wording, or the shorter version a product team keeps asking for. ### What to check before adding the axis Write down, for each candidate, the finding you expect and the action it would trigger. "Provider B fails a behaviour family that provider A holds" triggers something only if switching or dual-serving is genuinely on the table this quarter; if you are locked to one vendor, that column produces a number you will read and not act on. "The terser prompt variant fails a family the verbose one holds" triggers something almost always, because prompt text is the fastest thing you control and the most likely to change under you. Then check whether the axis is in the serving path at all, and check how the results are aggregated before anyone sees them. If the honest answer is that you need both attributions, the move is usually not to double twice but to shrink a different axis: run the four-cell provider-by-prompt cross against a smaller curated case set, and keep the wide case set for the single pairing you actually ship. That keeps both attributions available at roughly the old price, and keeps the suite something you can repeat.
- When is adding a provider you do not ship still worth it?When it is a genuine fallback or a migration candidate, or when you need a reference point to tell 'our prompt is weak' from 'this whole model family is weak' on a specific behaviour.
- How do you keep a second prompt variant honest?Make it a variant somebody would actually ship — the previous release's text, or the terse version a product team keeps asking for — not a strawman written to fail.
saying these in an interview costs you the question
- Treats the two axes as interchangeable because they cost the same.
- Adds a provider nobody ships or plans to ship, then never removes it.
- Uses a strawman second prompt that no team would deploy, and reports the gap as a finding.
- Never asks what action the extra column would trigger.