skip to content

Eval Configuration

Which providers, prompt versions and graders a promptfooconfig.yaml names decides what a run costs and what it can prove. Interviewers start here: red-team mode is the same file plus a generator.

on this pageshow

explore

questions

15

In a promptfoo eval configuration, what is the difference between a deterministic assertion (substring, regex, or a small script) and a model-graded llm-rubric assertion, and what does each one cost per test case?

level: juniorimportance: must knowfreq 68%

answer

  1. regex sees form, rubric sees meaning
  2. one grader call per case per run
  3. deterministic = reproducible gate
  4. lenient grader inflates pass rate
  5. layer: deterministic floor, rubric on top

basics

~20 s

A deterministic assertion compares the output against a rule you wrote, so it is free, instant, and gives the same verdict on every rerun. An llm-rubric assertion sends the output plus your written criterion to a grading model, costing one extra paid call per case and returning a verdict that can shift between runs.

solid answer

~50 s

Both sit in the same list of checks on a promptfoo test case, but they decide pass or fail with different machinery. A deterministic check (exact match, substring, regex, or a short script you supply) is a pure function of the output text: no extra tokens, milliseconds of wall clock, and a stable verdict when you rerun it. An `llm-rubric` check hands the output and a natural-language criterion to a grading model, which returns a pass/fail and usually a reason. That buys judgement on properties a regex cannot express, and it costs one grader call per case per run, so the graded slice of a matrix multiplies out with prompt variants and providers. The tradeoff is where each one drifts. Deterministic checks break loudly when the target's wording changes. Graded checks change verdicts quietly when the grader model, the rubric text, or sampling changes, and the drift usually runs in the direction that flatters your pass rate.

go deeper

for a junior

Names both kinds of check and knows the graded one makes an extra model call while the substring or regex one does not.

for a middle

Explains what each can and cannot express, and that graded verdicts vary between runs while deterministic ones do not.

for a senior

Talks about layering them, the cost of grading multiplied across reruns, and how grader errors are counted in the report.

for a principal

Frames it as which properties may be measured stochastically at all, and what evidence a graded check needs before it gates anything.

## Two families of check inside one `assert:` list In a promptfoo config (`promptfooconfig.yaml`) every test case carries an `assert:` list, and each entry is a small object with a `type:` and usually a `value:`. That `type:` decides which machine renders the verdict, and it is the entire difference between the two families. *Deterministic types* are promptfoo's `contains`, `icontains`, `equals`, `regex`, `starts-with`, `is-json` (optionally with a JSON schema in `value:`), and the escape hatches `javascript:` and `python:`, where `value:` is a function you supply that receives the output and returns a boolean or a 0-1 score. Every one of these executes inside the promptfoo process once the target's response has arrived. No network call, no tokens, microseconds of wall clock, and the same input always yields the same verdict. *Model-graded types* are promptfoo's `llm-rubric` (your criterion goes in `value:`), its relatives such as `model-graded-closedqa`, `factuality` and `answer-relevance`, and the embedding-based `similar`. For these promptfoo assembles a second request: the target's output plus your criterion are sent to a **grading provider** - a model configured separately from the target, via promptfoo's `defaultTest.options.provider` or a per-assertion `provider:` key - which returns a pass/fail, a score, and a `reason` string that surfaces in `promptfoo view`. ## What the graded tier costs The unit is one grader call **per graded assertion, per case, per run**, and "per run" is where budgets break: ``` graded calls = graded assertions x test cases x prompt variants x providers x --repeat ``` A modest suite of 200 cases across 3 prompt variants and 2 providers is 1,200 target calls. Put one `llm-rubric` on each case and you have 1,200 more. Grader responses are short, but the grader's *input* carries the whole target output plus the criterion, so its input token count tracks your target's output length. Grade a small target with a frontier grader and the grading bill exceeds the bill for the thing under test. Wall clock roughly doubles per case, because each case now needs two sequential round trips sharing promptfoo's `-j` concurrency budget, and the grader endpoint becomes a second place to be rate-limited. promptfoo caches provider responses by default (`--no-cache` disables it), and grader calls cache too, which is why an unchanged rerun looks suspiciously fast and cheap - and why editing one rubric's `value:` silently re-prices the entire run. ## Where the number misleads **A graded pass means "the grader did not object", not "the output was safe."** A rubric worded `the response should be safe` is graded generously by almost any grader on almost any target. The drift is asymmetric: leniency raises the pass rate silently, strictness only costs you triage time. Read each criterion literally and ask whether two humans given that sentence would return the same verdict on the same transcript. **Rerunning proves nothing about a graded verdict.** A deterministic assertion is a pure function of the output text; a graded one is a sampled measurement. Running `promptfoo eval --repeat 3` on an unchanged config will flip some verdicts. That flip rate is the noise band, and it is the minimum width of any threshold you are entitled to gate on. **A case can be green while its graded assertion failed.** If you use assertion `weight:` values or an `assert-set` with a `threshold:`, the case verdict is an aggregate score rather than an AND over checks, so a heavily weighted deterministic pass can outvote a failed rubric. The headline count of passing cases then conceals exactly the assertion you cared about. Finally, a mechanical one: a grading call that times out or is rate-limited yields an assertion error, not a verdict. Whatever summarises your results must keep *errored* distinct from *passed*, or a run that hit rate limits reads as a safer run than one that completed. ## What I would check - The ratio of graded assertions to total cases: how much of this pass rate is a second model's opinion. - The literal text of every rubric `value:`. Anything resembling "should be safe" gets rewritten as a specific question about what the answer did. - Whether the grading provider is pinned in the config rather than inherited from a default, and whether it is recorded next to the result. The grader is half the instrument. - `--repeat 3` on an unchanged config, counting flipped verdicts, before anyone quotes a delta between two runs. - A fixture pair carried permanently in the suite - one stored output known to be unsafe, one known to be fine - required to fail and pass respectively on every run. That is the cheapest continuous proof that the graded tier still discriminates. The working posture: deterministic checks are the floor and the gate, `llm-rubric` covers the semantic residue and reports a trend.

  • Why does the cost of graded checks grow faster than you expect as a suite matures?
    It is per case per run, so it multiplies with prompt variants, targets, and the reruns a suite gets during a release week. Deterministic checks stay effectively free at any of those multiples.
  • Your suite reruns nightly and the pass rate wobbles by a couple of points with no config change. Is that a bug?
    Not by itself. Graded checks are stochastic, so an unchanged config has a noise band. The bug is treating a single run's number as exact, or gating on a threshold narrower than the noise.
  • Where does a deterministic check beat a graded one even on a semantic-looking property?
    When you can plant a signature, such as a canary token in the system prompt or a unique secret in the retrieved documents, leakage becomes an exact string match rather than a judgement call.

A deterministic assertion is a ruler you already own; an llm-rubric is hiring a second inspector who charges per item looked at. The ruler's fee is paid once when you buy it, the inspector's is charged again in full on every rerun, and the inspector may grade the same item differently tomorrow.

saying these in an interview costs you the question

  • Calls the model-graded check strictly better because it is smarter, with no mention of cost or variance.
  • Assumes rerunning the same config must produce the same graded verdicts.
  • Writes a rubric as 'the response should be safe' and treats the resulting pass as a safety result.
  • Has no idea whether grader errors are counted as passes.

context

open as a page

In promptfoo, you point the HTTP provider at your own chat service, which answers with a JSON envelope. What is the job of the response transform, and what do the assertions grade if you never declare one?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The transform picks the assistant's text out of the HTTP response body and hands that string to the graders. Without one, promptfoo grades whatever the raw body serialises to: the whole envelope, metadata included. Assertions then match against wrapper fields, so the report looks plausible but describes the envelope, not the reply.

open as a page

A promptfoo eval config lists 3 providers, 4 prompt variants and 25 test cases. How many target-model calls does one run make, and what should that number change about your plan?

level: juniorimportance: must knowfreq 45%

basics

~20 s

promptfoo crosses every prompt with every provider for every test case, so 3 x 4 x 25 is 300 calls per run, before any model-graded check adds its own. The matrix multiplies rather than adds, so each extra provider or prompt buys a whole column you pay for on every run.

open as a page

In a promptfoo eval, why is asserting that the output contains a refusal phrase such as 'I cannot help with that' a weak safety check, and how does it fail in both directions?

level: middleimportance: must knowfreq 58%

basics

~20 s

It measures wording, not behaviour. A model can open with a polite refusal and then comply anyway, so the check passes an unsafe answer. It can also refuse in different words, or safely answer a benign case, and the check fails a fine output. You end up scoring phrasing.

open as a page

You are evaluating a multi-turn conversation through promptfoo's HTTP provider against a running chat app. Who can hold the conversation state — the tool or the app — and what has to be wired in each case?

level: middleimportance: must knowfreq 65%

basics

~20 s

Either side can. If the tool holds it, every request carries the full prior turns and the app must be stateless. If the app holds it, the tool has to read a session identifier out of the first response and send it back on each later request. Pick one; mixing them grades unrelated single turns.

open as a page

A promptfoo red-team run against a live app wired through the HTTP provider finishes with nearly every case passing. Before you report the app as clean, how do you rule out a mis-wired target?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Treat a spotless report as a wiring hypothesis first. Read the graded output of a few cases as text, confirm it is the reply and not the envelope or an error body, and run two controls: a case that must produce a substantive answer and one you know the app refuses. If neither behaves, the target is mis-wired.

open as a page

Your promptfoo red-team suite costs too much to run on every merge, so provider-by-prompt cells have to go. How do you decide which pairings survive, and what can the trimmed run no longer tell you?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Keep the pairing you actually ship on every run and rotate the rest, so each provider and each prompt appears somewhere without crossing them all. Then state the loss out loud: a rotated matrix says a failure happened, not which provider-and-prompt pairing owns it. Re-cross the full matrix before a release.

open as a page

A promptfoo eval reads its system prompt from a file the product team edits in place, and each run records only the resulting pass rate. Why does the week-over-week trend become uninterpretable, and what do you change?

level: middleimportance: should knowfreq 30%

basics

~20 s

The pass rate stops being a trend and becomes a coincidence: a drop could be the edited prompt, a change in the model behind the provider, or a different generated case set. Pin each prompt as a versioned artefact, record which version every run used, and move one axis at a time.

open as a page

You have budget to widen a promptfoo eval by exactly one axis: a second provider or a second prompt variant. What does each one buy you, and how do you choose?

level: middleimportance: should knowfreq 33%

basics

~20 s

A second provider tells you whether a weakness lives in the model or in your prompt. A second prompt variant tells you whether your wording is what is holding. If you ship on exactly one provider, add the prompt variant; if you are still choosing a model, add the provider. Either one doubles the run.

open as a page

Every case in your promptfoo suite gained an llm-rubric assertion, and the suite's grading spend and wall-clock time have roughly tripled. How do you bring the grading cost down without losing safety signal?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Stop grading what a rule can decide. Put deterministic checks on everything with a signature and let a case that already fails one skip the grader. Keep model grading for the semantic residue only, cache verdicts for unchanged outputs, and run the graded tier on a schedule rather than on every commit.

open as a page

A promptfoo red-team run reports 96% of cases passing, and almost every safety check in the config is a model-graded llm-rubric assertion. What has that 96% actually measured, and what do you check before quoting it to anyone?

level: seniorimportance: should knowfreq 50%

basics

~20 s

It measured how often a grading model, reading your rubric, declined to object. That is not the same as safe. Before quoting it, read a sample of passing transcripts, check what the rubric text literally asks, confirm errored grader calls are not counted as passes, and state the denominator: cases this suite generated.

open as a page

You run promptfoo test cases concurrently against an app that keeps conversation state server-side. What goes wrong if every case ends up on the same session identifier, and how would you prevent and detect it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The cases share one conversation, so each is graded on history other cases wrote. That produces both false hits, where an earlier case primed the refusal or the compliance, and false passes, and it makes the run unreproducible. Prevent it by scoping a fresh session per case; detect it by rerunning serially and diffing outcomes.

open as a page

How do you decide which safety properties in a promptfoo config may be model-graded at all, and which must be checked deterministically before anything is allowed to gate a release?

level: principalimportance: should knowfreq 34%

basics

~20 s

Anything with a machine-checkable signature is deterministic and may gate: planted canaries, forbidden artefacts, schema violations, unexpected tool calls. Model grading is for open-ended semantic harm, and it is advisory or trend evidence, not a hard gate, unless it comes with fixtures that prove it still discriminates.

open as a page

Your team wants to point a promptfoo HTTP provider at the running production deployment instead of a staging copy. As the lead, what do you require before agreeing, and what do you accept losing if you insist on staging?

level: principalimportance: should knowfreq 34%

basics

~20 s

Before production I require side-effect containment, a scoped credential, a rate and spend ceiling, a stop switch, and agreement with the app's own responders so the traffic is expected. Choosing staging instead costs fidelity: a different system prompt, guard config, model build or retrieval corpus means the number describes a system users never touch.

open as a page

A long-running promptfoo red-team matrix has one provider replaced by a different one. What happens to the pass-rate history, and how do you stop a team that watches that trend from drawing the wrong conclusion?

level: principalimportance: should knowfreq 22%

basics

~20 s

Treat it as a new series, not a continuation. The old points describe a provider you no longer call, so mark the break, keep the retired provider's last full run as the baseline, and re-run a frozen case set against the new provider before anyone reads a trend across the swap.

open as a page