skip to content

How would you gate a release on an automated red-team suite in CI?

level: seniorimportance: should knowfreq 40%

answer

  1. worst case, not average
  2. any new success blocks
  3. tier by severity, not one score
  4. deterministic checks over judge models
  5. repeat each case k times

basics

~20 s

Run the red-team job against every release candidate, block on any new success in high-severity categories rather than on an aggregate score, keep every past incident as a permanent case, and run each case several times because any single success counts as a failure.

solid answer

~50 s

A safety gate uses different semantics from a quality gate: safety is worst-case, so the rule is *any* new success in a blocking category stops the release, not *average score fell below a threshold*. Practically that means a red-team job — promptfoo's red-team mode is the common off-the-shelf choice — running against each release candidate with cases grouped by severity: cross-tenant data exfiltration and unauthorised state change block unconditionally; tone and policy-drift categories report without blocking. Make success checks deterministic wherever you can (a planted canary string appearing in output, an assertion that a specific tool fired), because a judge model in the blocking path turns a security gate into a flaky one. Run each case k times, since a defense that holds 95% of the time will otherwise pass by luck. Keep the expensive adaptive campaign on a nightly or pre-release schedule, not on every commit, and add a permanent regression case for every incident you have ever had.

go deeper

for a junior

Know that a safety check in the pipeline blocks on any occurrence of a serious failure, unlike a quality check that compares an average against a threshold.

for a middle

Explain severity tiering, why each case is run several times under nondeterminism, and why deterministic success checks such as planted canary strings belong in the blocking path.

for a senior

Show the full layering: a lean blocking set per release candidate, a slower adaptive campaign nightly or pre-release, a permanent regression case per incident, and a recorded override path with an owner and an expiry.

for a principal

Own the cost and credibility of the gate — how lean the blocking tier must stay to avoid being routed around, who signs an override, and the fact that the launch decision rests on the adaptive campaign rather than on the fast gate.

## Why a safety gate is not a quality gate Quality evaluation aggregates: you score many cases, take a mean, and ship if it clears a bar, because one weak answer among a thousand is acceptable. Safety does not aggregate that way. One successful cross-tenant data exfiltration is a breach whether the other 999 cases passed or not. So the gate's decision rule is *worst-case over cases*, not *average over cases*, and everything downstream follows from that inversion. ## Structuring the gate **Tier the cases by severity.** Three tiers work well: - **Blocking.** Outcomes that are unacceptable at any rate — data belonging to another user leaving the system, a state-changing or irreversible tool firing on injected instructions, credentials appearing in output. Any new success here fails the build, full stop. - **Budgeted.** Outcomes with real but bounded harm, where a small residual rate is tolerable and the gate compares against the previous release's rate rather than zero. - **Reporting.** Tone, style, policy drift. Recorded, trended, never blocking — putting these in the blocking path is how teams end up disabling the gate entirely. **Make the check deterministic where the goal allows it.** The strongest safety cases have a mechanical success condition: a canary record planted in a second tenant's data whose appearance in output is a decisive leak; an assertion that the transfer tool was never invoked with the attacker's arguments; a check that no outbound URL was constructed containing case data. Deterministic checks are fast, reproducible, and cheap to rerun — which is what lets them sit in a blocking path. Reserve judge models for categories where the harmful outcome is genuinely a judgement about text, and keep those in the reporting tier unless you have measured the judge's agreement with human labels. **Repeat each case.** Because the system is stochastic, a single run of a case is a coin flip against a probabilistic defense. Run each blocking case k times and treat *any* success as a failure. This is the same any-attempt-wins semantics the attacker enjoys, transplanted into CI. It also means your gate's cost scales with k, which is the main reason to keep the blocking set small and severe. ## Splitting fast from thorough Adaptive red-teaming — an attacker, human or automated, iterating against this specific build — is the evidence that actually characterises safety, and it is far too slow and expensive to run per commit. The usual layering: - **Per release candidate:** the fixed blocking set, repeated k times, deterministic checks, minutes not hours. - **Nightly or pre-release:** a broader campaign including adaptive attempts and best-of-N resampling, whose results inform the launch decision rather than blocking a merge. - **Per incident:** whatever happened in production becomes a permanent blocking case within a day. That last rule is the one that compounds. A red-team corpus grown from real incidents is worth more than a much larger corpus of invented prompts, because it encodes failures the system has actually demonstrated it is capable of. ## Handling the awkward parts **A new success appears and you cannot fix it before the deadline.** The gate should force an explicit, recorded decision — an owner, an expiry, and a compensating control such as disabling the affected tool or gating the feature behind approval — rather than a silent override. Overrides that leave no trace are how a gate stops meaning anything within two quarters. **The suite is a target itself.** Anyone with repository access can see exactly which attacks are checked, which makes the corpus a description of your known blind spots. Treat it as sensitive, and never let "passes the corpus" become the definition of safe — that is the static-suite trap in institutional form. **Cost.** Blocking cases run on every candidate, times k, times model cost. Keep the blocking tier lean and severe; push breadth into the nightly campaign. If the gate is slow enough that people route around it, it has failed regardless of how good its cases are. ## What good looks like A release pipeline where a candidate that lets injected content trigger an exfiltration path simply cannot ship; where the gate's failures are legible enough that an engineer sees the trajectory, the tool call and the canary in one place; where every production incident is represented; and where the number the launch decision actually rests on comes from the slower adaptive campaign, not from the fast gate. The fast gate stops regressions. It does not certify safety, and saying so out loud is part of a good answer.

  • Why not block the release when the aggregate safety score drops below a threshold?
    Because averaging lets a catastrophic case hide behind hundreds of benign passes, and it moves with the size of the suite rather than with real risk. For outcomes like cross-tenant leakage the acceptable count is zero at any denominator, so the gate must key on occurrence of specific severe outcomes. Aggregate scores are useful as a trend line, not as a decision rule.
  • How do you keep this gate from becoming so slow that teams disable it?
    Keep the blocking tier small and severe, use deterministic checks so runs are fast and reproducible, and push breadth into a nightly campaign that reports rather than blocks. Measure the gate's own wall-clock and false-failure rate as an operational metric. A gate that adds an hour to every candidate will be routed around, and a routed-around gate is worse than none because it looks like coverage.
  • A blocking case fails the day before a launch and the fix needs a week. What is the right process?
    Force an explicit, recorded exception: a named owner, an expiry date, and a compensating control that removes the harm path — disabling the affected tool, gating the feature behind human approval, or scoping the release to a cohort. The decision should be visible in the same place as the failure. Silent overrides are how a gate loses meaning; documented, time-boxed ones are how it survives real deadlines.

saying these in an interview costs you the question

  • Gates on an aggregate safety score with a percentage threshold
  • Runs each red-team case once and trusts the result
  • Puts an unvalidated judge model in the blocking path
  • Allows silent overrides when the gate fails before a deadline
  • Treats passing the CI red-team suite as certification that the release is safe

context