skip to content

How do you decide which safety properties in a promptfoo config may be model-graded at all, and which must be checked deterministically before anything is allowed to gate a release?

level: principalimportance: should knowfreq 34%

answer

  1. signature beats judgement
  2. only reproducible checks gate
  3. fixtures buy gate rights
  4. semantic residue routes to humans
  5. lenient failures are silent and flattering

basics

~20 s

Anything with a machine-checkable signature is deterministic and may gate: planted canaries, forbidden artefacts, schema violations, unexpected tool calls. Model grading is for open-ended semantic harm, and it is advisory or trend evidence, not a hard gate, unless it comes with fixtures that prove it still discriminates.

solid answer

~50 s

I split the assertion set by whether the unsafe property can be given a signature. **Deterministic tier — may gate.** Leakage of a planted canary or a known secret, a link to a host outside the allowlist, a tool invocation that should never fire, a schema violation, output carrying another tenant's identifier. Reproducible, free, and a red build is defensible to whoever it blocks. **Graded tier — measures, does not gate by default.** Whether an answer materially advanced a harmful task, whether advice was dangerous in substance, whether tone crossed a line. No rule expresses these, so grading is the only automation available; I treat its output as a trend and a queue of transcripts for humans. **What buys a graded check gate rights:** a specific criterion, fixtures of known-bad and known-good outputs it must fail and pass every run, a pinned grader configuration, and a threshold outside the run-to-run noise band.

go deeper

for a junior

Knows deterministic checks are the reliable ones and model grading is for things a rule cannot express.

for a middle

Can sort a real assertion list into the two tiers and explain why the graded ones are unreliable as a hard stop.

for a senior

Adds planting signatures to convert semantic properties into deterministic ones, and pins grader configuration so results stay comparable.

for a principal

Sets the policy end to end: what may gate, what evidence a graded check needs to earn gate rights and how it loses them, and what the release record must disclose.

This is a governance question wearing a config file as a disguise. The artefact is a list of `assert:` entries in `promptfooconfig.yaml`; the decision is which of those entries is allowed to turn a build red. I answer it with three rules and one asymmetry. ## Rule 1: engineer a signature before you buy a judgement If I can arrange for the unsafe outcome to leave a machine-checkable trace, I do that first, and then I check the trace with a deterministic assertion. The moves are cheap and mostly an afternoon of work each: - Plant a unique canary token in the system prompt, so instruction extraction becomes a `contains` on an exact string. - Seed a unique marker into the sensitive documents, so retrieval leakage becomes an exact match rather than a judgement about topical overlap. - Maintain an allowlist of link hosts, so exfiltration becomes a set-membership test inside a `javascript:` assertion. - Declare the output contract, so a malformed or over-permissive response is caught by `is-json` with a schema in `value:`. - Name the tools that must never fire for a given case, and assert their absence from the trace. Each of these converts a semantic problem into a reproducible one, and removes a per-case grader call from every future run at the same time. The engineering to plant a signature is almost always cheaper than the recurring cost of grading its absence. ## Rule 2: only a reproducible check may gate A gate is a social contract with the people it blocks. A check that returns a different verdict on an unchanged input breaks that contract in one of two directions, and both end with no signal: the team learns to rerun until green, or the team widens the threshold until the gate never fires. Deterministic assertions - `contains`, `regex`, `equals`, `is-json`, a `javascript:` function - are pure functions of the output, so a red build from one is defensible to whoever it stopped. Stochastic measurement belongs on a dashboard and in a review queue, where variance is visible and a human holds the decision. ## Rule 3: a graded check can earn gate rights, and can lose them Some properties have no signature at all - whether an answer materially advanced a harmful task, whether advice was dangerous in substance, whether a persona crossed a line. Grading is the only automation available. If such a property genuinely must block a release, the `llm-rubric` behind it needs four things before it is allowed to: 1. **A narrow criterion** - specific enough that two reviewers reading it would return the same verdict on the same transcript. "The response should be safe" never qualifies. 2. **A live fixture set** - stored outputs known to be unsafe that the check must fail, and known-good outputs it must pass, evaluated on every run, not once at adoption. 3. **A pinned, recorded grading configuration** - the provider and its settings, written next to the result, because changing the grader changes the instrument and silently breaks comparability with every earlier number. 4. **A threshold wider than the measured noise band** - established by rerunning an unchanged config, for instance with `promptfoo eval --repeat 3`, and counting flipped verdicts. It loses those rights when the fixtures stop discriminating, when the grader configuration moves without re-baselining, or when you observe people rerunning to get green. Any of the three and it drops back to advisory until repaired. ## The asymmetry that drives all three rules A lenient graded check fails silently and flatteringly: the pass rate stays high, nothing lands in anyone's triage queue, and the miss is discovered from outside the organisation. A strict one fails loudly, costs triage hours, and therefore gets fixed. Automation should gate on what it can prove, and route what it can only judge to humans with the transcripts attached. ## What the release record has to say Not one blended pass rate. It should separate what was *proved* from what was *judged*: which properties were checked deterministically and their results; which were graded, under which grader configuration, and what the graded tier's fixtures did on that run; how many assertions returned no verdict because grading errored; and what the suite did not attempt at all. A decision-maker who sees only a single percentage cannot tell whether it rests on exact string matches or on a second model's mood, and those two artefacts deserve very different amounts of confidence. The cost side is part of the policy too: the deterministic tier is free, so it can run on every commit; the graded tier costs one call per graded assertion per case per run, so its cadence is a budgeted decision, and the budget belongs in the same document as the gate policy.

  • Give a concrete example of turning a semantic property into a deterministic one.
    System-prompt extraction. Put a unique canary token in the system prompt; the property 'did it leak the instructions' becomes an exact string search over the output, free and reproducible.
  • A team wants a graded check as a hard release gate. What do you require first?
    A narrow criterion, a fixture set of known-bad and known-good outputs it must get right on every run, a pinned and recorded grader configuration, and a threshold outside the measured run-to-run noise.
  • How does a graded gate lose its gate rights?
    When its fixtures stop discriminating, when the grader configuration changes without re-baselining, or when the team is observed re-running to get green. Any of those and it drops back to advisory until repaired.

saying these in an interview costs you the question

  • Lets a vague graded criterion block releases with no fixtures behind it.
  • Treats deterministic checks as too primitive for safety work.
  • Never considers planting signatures to make a semantic property checkable.
  • Reports a single blended pass rate with no split between what was proved and what was judged.

context