skip to content

Assertions

A case only fails if some check says so, and a model-graded check can be lenient in exactly the direction that flatters you. Interviewers probe what your pass rate really measured.

on this pageshow

explore

questions

5

In a promptfoo eval configuration, what is the difference between a deterministic assertion (substring, regex, or a small script) and a model-graded llm-rubric assertion, and what does each one cost per test case?

level: juniorimportance: must knowfreq 68%

answer

  1. regex sees form, rubric sees meaning
  2. one grader call per case per run
  3. deterministic = reproducible gate
  4. lenient grader inflates pass rate
  5. layer: deterministic floor, rubric on top

basics

~20 s

A deterministic assertion compares the output against a rule you wrote, so it is free, instant, and gives the same verdict on every rerun. An llm-rubric assertion sends the output plus your written criterion to a grading model, costing one extra paid call per case and returning a verdict that can shift between runs.

solid answer

~50 s

Both sit in the same list of checks on a promptfoo test case, but they decide pass or fail with different machinery. A deterministic check (exact match, substring, regex, or a short script you supply) is a pure function of the output text: no extra tokens, milliseconds of wall clock, and a stable verdict when you rerun it. An `llm-rubric` check hands the output and a natural-language criterion to a grading model, which returns a pass/fail and usually a reason. That buys judgement on properties a regex cannot express, and it costs one grader call per case per run, so the graded slice of a matrix multiplies out with prompt variants and providers. The tradeoff is where each one drifts. Deterministic checks break loudly when the target's wording changes. Graded checks change verdicts quietly when the grader model, the rubric text, or sampling changes, and the drift usually runs in the direction that flatters your pass rate.

go deeper

for a junior

Names both kinds of check and knows the graded one makes an extra model call while the substring or regex one does not.

for a middle

Explains what each can and cannot express, and that graded verdicts vary between runs while deterministic ones do not.

for a senior

Talks about layering them, the cost of grading multiplied across reruns, and how grader errors are counted in the report.

for a principal

Frames it as which properties may be measured stochastically at all, and what evidence a graded check needs before it gates anything.

## Two families of check inside one `assert:` list In a promptfoo config (`promptfooconfig.yaml`) every test case carries an `assert:` list, and each entry is a small object with a `type:` and usually a `value:`. That `type:` decides which machine renders the verdict, and it is the entire difference between the two families. *Deterministic types* are promptfoo's `contains`, `icontains`, `equals`, `regex`, `starts-with`, `is-json` (optionally with a JSON schema in `value:`), and the escape hatches `javascript:` and `python:`, where `value:` is a function you supply that receives the output and returns a boolean or a 0-1 score. Every one of these executes inside the promptfoo process once the target's response has arrived. No network call, no tokens, microseconds of wall clock, and the same input always yields the same verdict. *Model-graded types* are promptfoo's `llm-rubric` (your criterion goes in `value:`), its relatives such as `model-graded-closedqa`, `factuality` and `answer-relevance`, and the embedding-based `similar`. For these promptfoo assembles a second request: the target's output plus your criterion are sent to a **grading provider** - a model configured separately from the target, via promptfoo's `defaultTest.options.provider` or a per-assertion `provider:` key - which returns a pass/fail, a score, and a `reason` string that surfaces in `promptfoo view`. ## What the graded tier costs The unit is one grader call **per graded assertion, per case, per run**, and "per run" is where budgets break: ``` graded calls = graded assertions x test cases x prompt variants x providers x --repeat ``` A modest suite of 200 cases across 3 prompt variants and 2 providers is 1,200 target calls. Put one `llm-rubric` on each case and you have 1,200 more. Grader responses are short, but the grader's *input* carries the whole target output plus the criterion, so its input token count tracks your target's output length. Grade a small target with a frontier grader and the grading bill exceeds the bill for the thing under test. Wall clock roughly doubles per case, because each case now needs two sequential round trips sharing promptfoo's `-j` concurrency budget, and the grader endpoint becomes a second place to be rate-limited. promptfoo caches provider responses by default (`--no-cache` disables it), and grader calls cache too, which is why an unchanged rerun looks suspiciously fast and cheap - and why editing one rubric's `value:` silently re-prices the entire run. ## Where the number misleads **A graded pass means "the grader did not object", not "the output was safe."** A rubric worded `the response should be safe` is graded generously by almost any grader on almost any target. The drift is asymmetric: leniency raises the pass rate silently, strictness only costs you triage time. Read each criterion literally and ask whether two humans given that sentence would return the same verdict on the same transcript. **Rerunning proves nothing about a graded verdict.** A deterministic assertion is a pure function of the output text; a graded one is a sampled measurement. Running `promptfoo eval --repeat 3` on an unchanged config will flip some verdicts. That flip rate is the noise band, and it is the minimum width of any threshold you are entitled to gate on. **A case can be green while its graded assertion failed.** If you use assertion `weight:` values or an `assert-set` with a `threshold:`, the case verdict is an aggregate score rather than an AND over checks, so a heavily weighted deterministic pass can outvote a failed rubric. The headline count of passing cases then conceals exactly the assertion you cared about. Finally, a mechanical one: a grading call that times out or is rate-limited yields an assertion error, not a verdict. Whatever summarises your results must keep *errored* distinct from *passed*, or a run that hit rate limits reads as a safer run than one that completed. ## What I would check - The ratio of graded assertions to total cases: how much of this pass rate is a second model's opinion. - The literal text of every rubric `value:`. Anything resembling "should be safe" gets rewritten as a specific question about what the answer did. - Whether the grading provider is pinned in the config rather than inherited from a default, and whether it is recorded next to the result. The grader is half the instrument. - `--repeat 3` on an unchanged config, counting flipped verdicts, before anyone quotes a delta between two runs. - A fixture pair carried permanently in the suite - one stored output known to be unsafe, one known to be fine - required to fail and pass respectively on every run. That is the cheapest continuous proof that the graded tier still discriminates. The working posture: deterministic checks are the floor and the gate, `llm-rubric` covers the semantic residue and reports a trend.

  • Why does the cost of graded checks grow faster than you expect as a suite matures?
    It is per case per run, so it multiplies with prompt variants, targets, and the reruns a suite gets during a release week. Deterministic checks stay effectively free at any of those multiples.
  • Your suite reruns nightly and the pass rate wobbles by a couple of points with no config change. Is that a bug?
    Not by itself. Graded checks are stochastic, so an unchanged config has a noise band. The bug is treating a single run's number as exact, or gating on a threshold narrower than the noise.
  • Where does a deterministic check beat a graded one even on a semantic-looking property?
    When you can plant a signature, such as a canary token in the system prompt or a unique secret in the retrieved documents, leakage becomes an exact string match rather than a judgement call.

A deterministic assertion is a ruler you already own; an llm-rubric is hiring a second inspector who charges per item looked at. The ruler's fee is paid once when you buy it, the inspector's is charged again in full on every rerun, and the inspector may grade the same item differently tomorrow.

saying these in an interview costs you the question

  • Calls the model-graded check strictly better because it is smarter, with no mention of cost or variance.
  • Assumes rerunning the same config must produce the same graded verdicts.
  • Writes a rubric as 'the response should be safe' and treats the resulting pass as a safety result.
  • Has no idea whether grader errors are counted as passes.

context

open as a page

In a promptfoo eval, why is asserting that the output contains a refusal phrase such as 'I cannot help with that' a weak safety check, and how does it fail in both directions?

level: middleimportance: must knowfreq 58%

basics

~20 s

It measures wording, not behaviour. A model can open with a polite refusal and then comply anyway, so the check passes an unsafe answer. It can also refuse in different words, or safely answer a benign case, and the check fails a fine output. You end up scoring phrasing.

open as a page

Every case in your promptfoo suite gained an llm-rubric assertion, and the suite's grading spend and wall-clock time have roughly tripled. How do you bring the grading cost down without losing safety signal?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Stop grading what a rule can decide. Put deterministic checks on everything with a signature and let a case that already fails one skip the grader. Keep model grading for the semantic residue only, cache verdicts for unchanged outputs, and run the graded tier on a schedule rather than on every commit.

open as a page

A promptfoo red-team run reports 96% of cases passing, and almost every safety check in the config is a model-graded llm-rubric assertion. What has that 96% actually measured, and what do you check before quoting it to anyone?

level: seniorimportance: should knowfreq 50%

basics

~20 s

It measured how often a grading model, reading your rubric, declined to object. That is not the same as safe. Before quoting it, read a sample of passing transcripts, check what the rubric text literally asks, confirm errored grader calls are not counted as passes, and state the denominator: cases this suite generated.

open as a page

How do you decide which safety properties in a promptfoo config may be model-graded at all, and which must be checked deterministically before anything is allowed to gate a release?

level: principalimportance: should knowfreq 34%

basics

~20 s

Anything with a machine-checkable signature is deterministic and may gate: planted canaries, forbidden artefacts, schema violations, unexpected tool calls. Model grading is for open-ended semantic harm, and it is advisory or trend evidence, not a hard gate, unless it comes with fixtures that prove it still discriminates.

open as a page