skip to content

Providers and Prompts

The provider and prompt lists multiply, so a matrix you can afford to repeat weekly is worth more than an exhaustive one you run once. Interviewers ask because this is where eval budgets quietly die.

on this pageshow

explore

questions

5

A promptfoo eval config lists 3 providers, 4 prompt variants and 25 test cases. How many target-model calls does one run make, and what should that number change about your plan?

level: juniorimportance: must knowfreq 45%

answer

  1. matrix multiplies, never adds
  2. providers x prompts x cases
  3. cost per run x cadence
  4. grader calls on top
  5. affordable to repeat beats exhaustive

basics

~20 s

promptfoo crosses every prompt with every provider for every test case, so 3 x 4 x 25 is 300 calls per run, before any model-graded check adds its own. The matrix multiplies rather than adds, so each extra provider or prompt buys a whole column you pay for on every run.

solid answer

~50 s

promptfoo builds a matrix: each prompt is sent to each provider for each test case, so the base cost is `providers x prompts x cases` = 300 calls. That is a floor, not a total. Any model-graded check spends a judge call per case on top; a multi-turn attack strategy spends several calls inside a single cell; retries repeat cells. What the number should change is cadence, not just whether you press run. A suite is only worth building if you can afford to repeat it, because a red-team number means little as a one-off and a lot as a trend. So price one run, multiply by the cadence you actually want (per merge, nightly, weekly), and trim the matrix until that product is affordable — rather than building the exhaustive matrix and discovering after one run that nobody will ever pay for the second.

go deeper

for a junior

Should get the multiplication right and say that the matrix multiplies rather than adds.

for a middle

Adds that the product is a floor — graded checks and multi-turn strategies add calls — and prices the run against a cadence.

for a senior

Frames the whole decision as cost-per-run times cadence, and argues for a repeatable smaller matrix over a one-off exhaustive one.

for a principal

Sets the standing rule for the org: what the eval budget buys per cadence, who approves widening an axis, and how trims are recorded so results stay interpretable.

### What the config is actually describing A promptfoo evaluation is defined by a YAML file (conventionally `promptfooconfig.yaml`) with three independent top-level lists: `providers:` — the model endpoints under test, each entry being one target you pay per call; `prompts:` — the prompt texts or templates, usually loaded as `file://prompts/system.txt`; and `tests:` — the test cases, each supplying variables and the assertions that decide pass or fail. When you run `promptfoo eval`, the tool does not walk these lists in parallel. It **crosses** them: every prompt is rendered for every case and sent to every provider. The unit of work — one prompt, one case, one provider — is a **cell**, and the run is the full grid of cells. ### The arithmetic, and why it is only a floor ``` providers 3 x prompts 4 x cases 25 = 300 target-model calls + 1 grader call per model-graded case -> up to ~600 calls x weekly cadence -> ~2,400-4,800 calls / month ``` Four things routinely push the real number above the product: - **Model-graded assertions.** An assertion of type `llm-rubric`, `answer-relevance`, `factuality` or similar is decided by a second model call. That grader is often a *larger* model than the target, so the cheaper axis of the run can end up being the expensive one on the invoice. - **Multi-turn strategies.** In promptfoo's red-team mode, entries under `redteam.strategies` such as the conversational ones spend several calls inside a single cell, because the cell is a conversation rather than a single request. - **Generated case counts.** In red-team mode the cases are produced by promptfoo from `redteam.plugins` and `redteam.numTests` rather than typed by hand, and strategies then expand that set further. The width of your matrix is therefore an *output* of the generation step. You do not know the case axis until you look. - **Repetition.** `promptfoo eval --repeat 3` multiplies everything by three, which is often the right call on a stochastic system and is always a 3x on the bill. ### What it costs in practice At commodity hosted-model prices and a modest prompt, 300 single-shot calls is small money — single-digit dollars. Add a large grader per case and a multi-turn strategy and the same config is tens of dollars a run; at per-merge cadence on a busy repository that is a real monthly line item. Wall-clock matters as much: promptfoo runs a bounded number of requests concurrently (raise it with `promptfoo eval -j <n>`), so a few hundred calls at a second or two each is minutes, not seconds, and a merge gate that takes ten minutes gets disabled by whoever is waiting on it. ### Where the number misleads The most common misread is the **cheap second run**. promptfoo caches provider responses by default, so re-running an unchanged config costs almost nothing and returns almost instantly. That feels like a confirmation and is not one: cached cells replay yesterday's sampled responses rather than re-measuring a stochastic system. If you want a genuine second measurement you need `promptfoo eval --no-cache` (or `PROMPTFOO_CACHE_ENABLED=false`) and the full price again. The second misread is treating the product as the total and then being surprised by an invoice two or three times the estimate — graders and multi-turn cells are where that gap lives. ### What to check before you press run Print the cell count the config will actually produce and confirm it against your arithmetic; in red-team mode, generate the cases first and count them rather than assuming `numTests`. Establish whether a cell is one call or a conversation. Identify which assertions are model-graded and on which model. Then multiply the whole thing by the cadence you actually want — per merge, nightly, weekly — because a red-team number is worth little as a snapshot and a lot as a trend. If that product is unaffordable, cut an axis deliberately and write down what you cut; an unrecorded trim is what turns "the suite passed" into a claim nobody can bound.

  • Does adding a fifth test case cost the same as adding a fifth prompt variant?
    No. With 3 providers and 4 prompts, one more case adds 12 calls; one more prompt adds 3 x 25 = 75. The axis with the widest product behind it is always the expensive one to grow.
  • Where does the real call count exceed providers x prompts x cases?
    Model-graded checks add a judge call per case, multi-turn strategies spend several calls inside one cell, and retried or repeated cells are paid for again. Treat the product as a floor.

The three lists are the sides of a box, not items on a shopping list: lengthening any one side re-prices the whole volume. And you pay that volume again every time the suite runs, not once.

saying these in an interview costs you the question

  • Adds the axes instead of multiplying them (answers 32).
  • Assumes one cell is always exactly one call, ignoring graders and multi-turn strategies.
  • Treats an exhaustive one-off run as more valuable than a repeatable smaller one.
  • Plans to widen the matrix without ever pricing it against how often it will run.

context

open as a page

Your promptfoo red-team suite costs too much to run on every merge, so provider-by-prompt cells have to go. How do you decide which pairings survive, and what can the trimmed run no longer tell you?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Keep the pairing you actually ship on every run and rotate the rest, so each provider and each prompt appears somewhere without crossing them all. Then state the loss out loud: a rotated matrix says a failure happened, not which provider-and-prompt pairing owns it. Re-cross the full matrix before a release.

open as a page

A promptfoo eval reads its system prompt from a file the product team edits in place, and each run records only the resulting pass rate. Why does the week-over-week trend become uninterpretable, and what do you change?

level: middleimportance: should knowfreq 30%

basics

~20 s

The pass rate stops being a trend and becomes a coincidence: a drop could be the edited prompt, a change in the model behind the provider, or a different generated case set. Pin each prompt as a versioned artefact, record which version every run used, and move one axis at a time.

open as a page

You have budget to widen a promptfoo eval by exactly one axis: a second provider or a second prompt variant. What does each one buy you, and how do you choose?

level: middleimportance: should knowfreq 33%

basics

~20 s

A second provider tells you whether a weakness lives in the model or in your prompt. A second prompt variant tells you whether your wording is what is holding. If you ship on exactly one provider, add the prompt variant; if you are still choosing a model, add the provider. Either one doubles the run.

open as a page

A long-running promptfoo red-team matrix has one provider replaced by a different one. What happens to the pass-rate history, and how do you stop a team that watches that trend from drawing the wrong conclusion?

level: principalimportance: should knowfreq 22%

basics

~20 s

Treat it as a new series, not a continuation. The old points describe a provider you no longer call, so mark the break, keep the retired provider's last full run as the baseline, and re-run a frozen case set against the new provider before anyone reads a trend across the swap.

open as a page