A red-team engagement on a customer-facing chat assistant closes with a list of confirmed findings, and your release gate offers three dispositions: block the release, ship behind a compensating control, or ship with a named acceptance. How do you decide which disposition each finding gets?
answer
- classes, not tool hits
- consequence of one instance
- irreversible = block
- control must already ship
- acceptance needs a signer and a date
basics
~20 sDecide by finding class and consequence, not by a tool's score. Classes whose harm is irreversible or legally exposed block. A finding ships behind a compensating control only if a control already deployed demonstrably stops it. Everything else ships with a named owner accepting it in writing, with a review date.
solid answer
~50 sThe gate has to speak in **finding classes**, not raw hits from whichever scanner produced them. Many tool hits collapse into one class ("the assistant will produce operational instructions for a regulated harm"), and one hit can raise two. Then apply consequence, not frequency: what happens on a single successful instance in production, to whom, and can it be undone. Blocking is reserved for classes where one instance is unacceptable — irreversible harm, another tenant's data, an action the agent can take with real-world effect. Compensating control is available only when a control is already in the shipping configuration and the *specific* finding was retested through it and failed to reproduce; a planned control is not a control. Everything left is a named acceptance, scoped and dated, signed by someone who owns the product's risk. The tradeoff: a gate with a large blocking set gets routed around, and a gate that blocks on nothing is a mailing list.
go deeper
Knows the three dispositions exist and that findings are not all treated the same; can explain that blocking is reserved for the worst classes.
Explains that the gate's unit is a deduplicated class, decides from consequence rather than hit count, and knows a compensating control must already be deployed and retested.
Runs the ladder under launch pressure, keeps the class vocabulary stable across engagements, tracks acceptances to expiry, and records untested surfaces as untested rather than clean.
Owns whether the gate is credible at all: whether it has ever blocked, whether teams route around it, and how the blocking set is sized against what fix teams can clear.
### What the gate actually consumes Start with vocabulary, because three different things get called *a finding*. A **tool hit** is one attempt that some detector marked successful — a garak attempt whose detector scored past its threshold, a PyRIT prompt whose scorer returned true, a promptfoo test case marked fail. A **report item** is a deduplicated behaviour, written up with reproduction steps and a consequence. A **class** is the group of report items that produce the same consequence in production. A release gate is a decision about classes. It is not a decision about hits, and the number of hits is largely a property of how the suite was authored. A **disposition** is the decision attached to a class: block the release, ship behind a compensating control, or ship with a named acceptance. A **compensating control** is a mechanism already in the shipping configuration — an output classifier, an input filter, a rail rule, a narrowed tool permission, a rate limit — that bounds the consequence without removing the behaviour. A **named acceptance** is a person with launch authority signing that the product ships with the behaviour reachable and nothing standing in front of it. ### The mechanism, step by step 1. **Collapse.** Group hits into behaviours and behaviours into classes. Three hundred successful attempts out of a generated suite are routinely four or five classes; the count reflects how many seed objectives the generator expanded and how many templates it applied to each. 2. **Write the consequence sentence.** For each class: *in production, one success means X happens to Y, and it is / is not reversible.* If you cannot write that sentence, you do not have a gate item yet — you have tool output. 3. **Apply the ladder to the sentence, not to the score.** - **Block** when one instance is unacceptable: another tenant's data leaving the tenant, an agent action with irreversible external effect, a harm category the organisation has committed against. The rate is irrelevant to this decision; it matters to the fix plan, not the disposition. - **Compensating control** when the consequence is bounded *and* a control already in the shipping path was shown to stop this specific class. A planned control is not a control. - **Named acceptance** for the residual: bounded consequence, no control, a signer who could have said no, with a scope and a date. ### What it costs Collapsing a few hundred hits into classes and writing consequence sentences is roughly a senior engineer-day per engagement, and it is the step teams skip because it produces no new numbers. Each candidate compensating control adds a retest: replaying the original attempts plus variations against the build that actually deploys, which spends metered inference calls again — hundreds to a few thousand, trivial in money and expensive in coordination, because someone has to stand up that exact configuration. Each blocking class costs release delay measured in days to weeks, paid by a team that did not choose the class. Each acceptance costs recurring ledger review for as long as the product lives. A gate is therefore not free to make stricter, which is precisely why the blocking set has to be small and defended rather than large and negotiable. ### Where the number misleads The seductive number is the pass rate: *4% of 5,000 attempts succeeded.* It misleads at both ends. The denominator is authored — padding the suite with easy variants drops the rate without the product changing, so the rate tracks suite composition at least as much as product risk. The numerator is graded by a detector or judge with its own false-positive and false-negative behaviour that nobody re-measured against this product's outputs. The second misleading number is the disposition tally: *12 findings, 3 blocked, 9 accepted* reads as a 25% block rate, but if the nine collapse to two classes and the three to one, the ratio is describing report formatting. The deepest misreading is the gate's **silence**. A gate reports on the surfaces an engagement reached. A launch reviewer who sees only a verdict infers coverage nobody claimed, so untested surfaces must be printed next to the verdict rather than simply absent from it. ### What you would check Before trusting your own gate: is the class vocabulary fixed up front and stable across engagements, or invented at report time to fit what was found — an unstable vocabulary cannot be compared run to run and cannot be audited. Has the gate ever actually blocked anything? Are acceptances expiring, or quietly accumulating into permanent debt? Was every compensating control retested against the build that deploys, not a staging variant? And can you produce, next to the verdict, the list of what went untested this cycle?
- A generated suite returns 300 successful attempts. How many gate items is that?Unknown until they are collapsed. The gate's unit is a behaviour class with a consequence; 300 attempts is usually a handful of classes, and the count is a property of the generator, not of the product's risk.
- A class you would normally block on was found only once, in a contrived 20-turn setup. Does it still block?Consequence decides blocking, effort decides urgency and sometimes the compensating control. If one instance is unacceptable, it blocks; the difficulty of reaching it belongs in the write-up, not in the disposition.
- Who writes the class list?The programme owner, up front and stable across engagements, so gates are comparable run to run — not the tester at report time, which lets the vocabulary drift to fit whatever was found.
A generated suite reporting 300 successes is like a bug tracker that files one ticket per line of a stack trace: the number tells you how the reporter is configured, not how many distinct things are broken.
saying these in an interview costs you the question
- Deciding disposition from a scanner's pass/failure percentage rather than from what one instance does in production.
- Letting a planned or roadmap control count as a compensating control at the gate.
- Treating every raw tool hit as its own gate item, which floods the gate and trains reviewers to wave everything through.
- Having no blocking class at all, so the gate has never stopped a release and everyone knows it.
- Acceptances with no owner, no scope and no expiry.