skip to content

A promptfoo eval config lists 3 providers, 4 prompt variants and 25 test cases. How many target-model calls does one run make, and what should that number change about your plan?

level: juniorimportance: must knowfreq 45%

answer

  1. matrix multiplies, never adds
  2. providers x prompts x cases
  3. cost per run x cadence
  4. grader calls on top
  5. affordable to repeat beats exhaustive

basics

~20 s

promptfoo crosses every prompt with every provider for every test case, so 3 x 4 x 25 is 300 calls per run, before any model-graded check adds its own. The matrix multiplies rather than adds, so each extra provider or prompt buys a whole column you pay for on every run.

solid answer

~50 s

promptfoo builds a matrix: each prompt is sent to each provider for each test case, so the base cost is `providers x prompts x cases` = 300 calls. That is a floor, not a total. Any model-graded check spends a judge call per case on top; a multi-turn attack strategy spends several calls inside a single cell; retries repeat cells. What the number should change is cadence, not just whether you press run. A suite is only worth building if you can afford to repeat it, because a red-team number means little as a one-off and a lot as a trend. So price one run, multiply by the cadence you actually want (per merge, nightly, weekly), and trim the matrix until that product is affordable — rather than building the exhaustive matrix and discovering after one run that nobody will ever pay for the second.

go deeper

for a junior

Should get the multiplication right and say that the matrix multiplies rather than adds.

for a middle

Adds that the product is a floor — graded checks and multi-turn strategies add calls — and prices the run against a cadence.

for a senior

Frames the whole decision as cost-per-run times cadence, and argues for a repeatable smaller matrix over a one-off exhaustive one.

for a principal

Sets the standing rule for the org: what the eval budget buys per cadence, who approves widening an axis, and how trims are recorded so results stay interpretable.

### What the config is actually describing A promptfoo evaluation is defined by a YAML file (conventionally `promptfooconfig.yaml`) with three independent top-level lists: `providers:` — the model endpoints under test, each entry being one target you pay per call; `prompts:` — the prompt texts or templates, usually loaded as `file://prompts/system.txt`; and `tests:` — the test cases, each supplying variables and the assertions that decide pass or fail. When you run `promptfoo eval`, the tool does not walk these lists in parallel. It **crosses** them: every prompt is rendered for every case and sent to every provider. The unit of work — one prompt, one case, one provider — is a **cell**, and the run is the full grid of cells. ### The arithmetic, and why it is only a floor ``` providers 3 x prompts 4 x cases 25 = 300 target-model calls + 1 grader call per model-graded case -> up to ~600 calls x weekly cadence -> ~2,400-4,800 calls / month ``` Four things routinely push the real number above the product: - **Model-graded assertions.** An assertion of type `llm-rubric`, `answer-relevance`, `factuality` or similar is decided by a second model call. That grader is often a *larger* model than the target, so the cheaper axis of the run can end up being the expensive one on the invoice. - **Multi-turn strategies.** In promptfoo's red-team mode, entries under `redteam.strategies` such as the conversational ones spend several calls inside a single cell, because the cell is a conversation rather than a single request. - **Generated case counts.** In red-team mode the cases are produced by promptfoo from `redteam.plugins` and `redteam.numTests` rather than typed by hand, and strategies then expand that set further. The width of your matrix is therefore an *output* of the generation step. You do not know the case axis until you look. - **Repetition.** `promptfoo eval --repeat 3` multiplies everything by three, which is often the right call on a stochastic system and is always a 3x on the bill. ### What it costs in practice At commodity hosted-model prices and a modest prompt, 300 single-shot calls is small money — single-digit dollars. Add a large grader per case and a multi-turn strategy and the same config is tens of dollars a run; at per-merge cadence on a busy repository that is a real monthly line item. Wall-clock matters as much: promptfoo runs a bounded number of requests concurrently (raise it with `promptfoo eval -j <n>`), so a few hundred calls at a second or two each is minutes, not seconds, and a merge gate that takes ten minutes gets disabled by whoever is waiting on it. ### Where the number misleads The most common misread is the **cheap second run**. promptfoo caches provider responses by default, so re-running an unchanged config costs almost nothing and returns almost instantly. That feels like a confirmation and is not one: cached cells replay yesterday's sampled responses rather than re-measuring a stochastic system. If you want a genuine second measurement you need `promptfoo eval --no-cache` (or `PROMPTFOO_CACHE_ENABLED=false`) and the full price again. The second misread is treating the product as the total and then being surprised by an invoice two or three times the estimate — graders and multi-turn cells are where that gap lives. ### What to check before you press run Print the cell count the config will actually produce and confirm it against your arithmetic; in red-team mode, generate the cases first and count them rather than assuming `numTests`. Establish whether a cell is one call or a conversation. Identify which assertions are model-graded and on which model. Then multiply the whole thing by the cadence you actually want — per merge, nightly, weekly — because a red-team number is worth little as a snapshot and a lot as a trend. If that product is unaffordable, cut an axis deliberately and write down what you cut; an unrecorded trim is what turns "the suite passed" into a claim nobody can bound.

  • Does adding a fifth test case cost the same as adding a fifth prompt variant?
    No. With 3 providers and 4 prompts, one more case adds 12 calls; one more prompt adds 3 x 25 = 75. The axis with the widest product behind it is always the expensive one to grow.
  • Where does the real call count exceed providers x prompts x cases?
    Model-graded checks add a judge call per case, multi-turn strategies spend several calls inside one cell, and retried or repeated cells are paid for again. Treat the product as a floor.

The three lists are the sides of a box, not items on a shopping list: lengthening any one side re-prices the whole volume. And you pay that volume again every time the suite runs, not once.

saying these in an interview costs you the question

  • Adds the axes instead of multiplying them (answers 32).
  • Assumes one cell is always exactly one call, ignoring graders and multi-turn strategies.
  • Treats an exhaustive one-off run as more valuable than a repeatable smaller one.
  • Plans to widen the matrix without ever pricing it against how often it will run.

context