Every case in your promptfoo suite gained an llm-rubric assertion, and the suite's grading spend and wall-clock time have roughly tripled. How do you bring the grading cost down without losing safety signal?
answer
- graded calls = cases x runs x cells
- rule first, rubric for the residue
- short-circuit after a deterministic failure
- cache verdicts on unchanged output
- cheaper grader breaks comparability
basics
~20 sStop grading what a rule can decide. Put deterministic checks on everything with a signature and let a case that already fails one skip the grader. Keep model grading for the semantic residue only, cache verdicts for unchanged outputs, and run the graded tier on a schedule rather than on every commit.
solid answer
~60 sGrading cost is per case, per run, so you attack both factors. **Fewer graded cases.** Most assertions do not need a model. Canary leakage, forbidden hosts, schema violations, and unexpected tool calls all have signatures. Move those to deterministic checks and let the rubric handle only what is genuinely semantic. Ordering matters too: a case already failing a deterministic check does not need a grader verdict to be reported as a failure. **Fewer runs of the graded tier.** Split the suite. A fast deterministic tier runs on every change; the graded tier runs nightly or before a release. Cache verdicts keyed on the output text so an unchanged output is not re-graded. **Cheaper per verdict** is the option to reach for last. Swapping the grader for a smaller model changes what the number means, so treat it as a config change that breaks comparability with earlier runs and re-check it against your known-bad fixtures before trusting it. Trimming the provider and prompt matrix also cuts grading calls, but that trades away coverage rather than cost.
go deeper
Suggests grading fewer cases or running the graded checks less often.
Separates deterministic from semantic properties and proposes a two-tier suite with different cadences.
Adds caching keyed on output and rubric, short-circuiting, and treats a grader swap as a change that invalidates historical comparisons; verifies with fixtures afterwards.
Sets the policy: which tier may gate a release, what the graded tier is budgeted for, and how a shrunken suite must be disclosed in the report.
## Write the cost model down first Every lever is visible once the multiplication is on paper: ``` grader calls = graded assertions x test cases x prompt variants x providers x --repeat ``` The bill is that product times the grader's per-call price; the wall clock is that product divided by promptfoo's `-j` concurrency, plus queueing when the grading endpoint rate-limits you. A tripling almost never comes from one factor. It comes from somebody adding a blanket `llm-rubric` under `defaultTest.assert:` - which applies it to every case in the file at once - while the matrix and the case count also grew. ## Lever 1: fewer graded assertions Walk the assertion list and ask of each rubric: could a rule decide this? In practice a large share of rubric checks were written for convenience rather than necessity, and each has a deterministic equivalent: | property | deterministic form in promptfoo | |---|---| | system-prompt leakage | `contains` on a canary token planted in the prompt | | retrieval leakage | `contains` on a unique marker seeded in the corpus | | exfiltration link | `javascript:` assertion testing the host against an allowlist | | malformed structured output | `is-json` with a schema in `value:` | | a tool that must never fire | `javascript:` assertion over the response trace | Every one of those moves from one paid call per run to zero, and gains reproducibility on the way. Note what promptfoo does *not* do for you here: it evaluates the whole `assert:` list for a case, so there is no flag that skips the grader once a cheap check has already failed. Short-circuiting is something you arrange structurally - split the suite, or gate the graded config behind a green deterministic run - not something you switch on. ## Lever 2: fewer runs of the graded tier Split one config into two. A deterministic tier is free and reproducible, so it can run on every commit. A graded tier is metered and stochastic, so it belongs on a nightly or pre-release cadence where a person reads the diff. This lever cuts the `runs` factor outright, and it fixes a second problem for nothing: a stochastic check was always a poor per-commit gate, because reruns of an unchanged config move the verdict, and a gate that flips on noise gets re-run for green or loosened until it never fires. promptfoo's `--filter-failing <results.json>` (rerun only what failed last time) and `--filter-first-n` (a smoke subset) are useful for the working loop between full graded runs. ## Lever 3: caching promptfoo caches provider responses by default, grader calls included, and the cache key is the request itself: the grading provider, the rendered grading prompt, and therefore the criterion text and the output being judged. In a suite where most outputs are unchanged between runs, this is frequently the single largest saving and it costs zero signal, because you are declining to repeat an identical measurement. Two practical consequences: `--no-cache` on a full run is expensive and should be deliberate, and editing one rubric's `value:` invalidates every verdict that used it, so a one-word wording tweak re-prices the run. ## Lever 4, last: a cheaper grader Pointing `defaultTest.options.provider` at a smaller model is the lever people reach for first and should reach for last. It is not a price change, it is an instrument change: verdicts shift, usually in the lenient direction, and your historical numbers stop being comparable. If you do it, change one thing at a time, rerun the known-bad and known-good fixtures, and record the grader configuration alongside the result so nobody compares across the swap unknowingly. ## What I would not do Silently drop test cases to fit the bill. That reduces what the suite attempted while leaving the headline pass rate looking the same or better - the failure mode that makes an eval actively misleading rather than merely incomplete. Same for trimming the provider or prompt matrix: it is a legitimate choice under budget pressure, but the report must name which pairings stopped being tested, immediately next to the number, or the next reader will assume the coverage from last month. Also on the do-not list: loosening rubric wording so fewer cases fail. It does not even save money - the grader call is made regardless - it only buys a better number. ## What I would check afterwards That the known-bad fixtures still fail on the reduced arrangement. That every rubric removed has a named deterministic replacement covering the same property, not merely a similar-sounding one. That the total failure count did not drop simply because fewer assertions were evaluated - compare failures per graded assertion, not failures per run. And that the two tiers are reported separately, so the reader can see which part of the result is reproducible and which part is a grader's opinion sampled once.
- What do you key a grading cache on?The output text, the rubric text, and the grader configuration. Change any of the three and the cached verdict is no longer about the same measurement.
- Why is a stochastic graded check a poor per-commit gate even ignoring cost?Reruns of an unchanged config move the result, so the gate flips on noise. Teams respond by loosening the threshold until it stops blocking, which removes the signal entirely.
- After the change, how do you show the suite did not get weaker?Run the known-bad fixtures through the new arrangement and require them to fail, and confirm each rubric you removed had a deterministic replacement covering the same property.
saying these in an interview costs you the question
- Cuts the case list to fit the bill and reports the pass rate as if nothing changed.
- Swaps in a cheaper grader and compares the new pass rate directly against last month's.
- Keeps a stochastic graded check as the per-commit gate because it is 'the strongest signal'.
- Never considers that most of the rubric checks are deciding things a rule could decide.