At a release gate, a product team asks to ship a confirmed red-team finding behind an output classifier as a compensating control rather than blocking. What evidence do you require before accepting that control, and what goes into the gate record?
answer
- retest the finding, not the category
- shipping build, not staging
- attack the control too
- fail open or fail closed
- threshold in the record, expiry on the item
basics
~20 sRequire the original finding retested end to end with the control enabled in the exact shipping configuration, not the control's own published coverage. Ask what happens when the control is unavailable, degraded or times out, and who notices. Record the control, the retest evidence, residual risk, an owner and an expiry date.
solid answer
~50 sA compensating control is accepted on **evidence about this finding in this configuration**, never on the control's category list. Three things must be shown. First, reproduction: the transcript that produced the finding, replayed against the build that ships with the classifier in the path, no longer producing the harmful output. Second, robustness of the control itself: someone tried the same objective through variation — reformatting, another language, splitting across turns — because a control that catches only the literal transcript is a string match against one test case. Third, failure behaviour: what happens when the classifier is unreachable, slow, or low-confidence. If it fails open, the residual risk is not what you were told, and the control's availability is now part of the gate. The record must let a future reader reverse the decision: which finding, which control, which build, what residual, accepted by whom, revisited when.
go deeper
Knows the finding should be retested with the control turned on, rather than trusting that the control exists.
Requires an end-to-end retest in the shipping configuration and asks about the control's threshold and false positives.
Also pressures the control adversarially, pins the operating threshold and failure behaviour in the record, and gives the acceptance an owner and an expiry.
Decides when a probabilistic control is categorically the wrong instrument for a class, and how many live compensating controls a team can actually operate before the gate is fiction.
### The substitution this question is about An **output classifier** here is a component that inspects the model's generated text before it reaches the user and blocks or rewrites it when a harm score crosses a configured **threshold** (the confidence value above which the component treats the output as violating). As a **compensating control** it does not remove the underlying behaviour; it interposes something that reduces how often, and how easily, the consequence reaches a user. The failure mode the gate must resist is accepting the control's *description* — a documented list of harm categories, a published benchmark score — in place of a *measurement* of this finding, in this product's prompt and tool configuration, at this threshold. A category list is a claim about a distribution. The gate is deciding about one behaviour. ### What to demand, in order of how often it is missing 1. **End-to-end retest.** The original attempts replayed against the shipping build with the classifier in the request path. Not a staging build with different system instructions, not the classifier evaluated standalone on a file of saved outputs. Configuration drift between what was tested and what deploys is where most of these decisions quietly break, because the system prompt, retrieval context and tool set all change what the model emits and therefore what the classifier sees. 2. **Adversarial pressure on the control.** Once accepted, the classifier is part of the attack surface. Someone must pursue the same objective through variation — different phrasing, different formatting, split across turns, a different language — rather than replaying the recorded string. Without that, the retest establishes that one transcript is blocked, which is a far weaker claim than the one being made at the gate. 3. **Threshold and its cost.** The operating point trades misses against false positives on ordinary traffic. Record the threshold that was accepted, and record the false-positive rate measured on real traffic at that point, because the two are the same dial. If the threshold is quietly relaxed after launch because the product team dislikes the refusals, the gate's evidence has expired without anyone saying so. 4. **Failure and observability.** Unreachable, timing out, rate-limited, degraded model behind the endpoint: does the request proceed unchecked (**fail open**) or get refused (**fail closed**)? Is there a health metric and an alert on the control itself? A fail-open control converts a security decision into an availability dependency — acceptable when written down and staffed, dangerous when discovered during an incident. 5. **Scope and expiry.** What the control covers and what it does not, so the acceptance does not silently stretch across features added later. ### What it costs A credible acceptance is roughly one to three engineer-days: standing up the shipping configuration, replaying the original attempts, authoring and running a variation set, and pulling a false-positive sample from production-like traffic. In inference terms it is small — typically hundreds to low thousands of metered calls, plus one classifier call per generation. The real recurring cost is at runtime: an extra network hop on every response, added latency, and a per-call classifier charge that persists for the life of the feature. That is why teams later argue to raise the threshold or sample rather than classify every response, and both of those changes invalidate the gate evidence. ### Where the number misleads The number teams bring to the gate is the control's published accuracy or F1 on a public safety benchmark. It misleads because that score was computed on a different distribution — generic prompts, not this product's outputs — and because a headline accuracy hides the operating point: a classifier can be excellent on average and near-useless in the narrow region the finding lives in. The second misleading number is the retest tally, *40 of 40 replays blocked*. Forty replays of one string is a sample of size one from the population that matters; the correct denominator is distinct approaches to the same objective, not repetitions of the same approach. The third is any accuracy figure quoted while the control fails open, since a component that is not reached has no accuracy at all. ### What you would check Diff the retest configuration against the deploy manifest, field by field. Confirm the variation set was authored by someone other than the person who tuned the control. Get the false-positive number at the accepted threshold from real traffic, not from the vendor. Ask what a request does when the classifier endpoint returns 5xx, and whether that path is alerted and owned. Then write the threshold, the build identifier, the residual, the signer and the expiry into the record — and note honestly that a probabilistic control changes the rate and the ease of a harmful outcome, rarely its possibility. If the class was unshippable at a single instance, the honest answer at the gate is that it still blocks.
- The retest passes, but only that exact transcript was replayed. What is missing?Variation. The claim being made is about a behaviour class, so the same objective has to be pursued through rephrasing, formatting and multi-turn setup; otherwise you have shown a single string is blocked.
- The control fails open under load. Does the acceptance still stand?Only if the record says so and someone owns the control's availability with alerting. Otherwise the residual risk that was signed for is not the residual risk being shipped.
A classifier that fails open when its endpoint times out is a lock that unlocks itself whenever the building gets busy: it is absent exactly during the traffic that stresses it most.
saying these in an interview costs you the question
- Accepting a control because its documentation lists the harm category, with no retest of the actual finding.
- Retesting against a build whose configuration differs from what deploys.
- Never asking what the system does when the classifier is unavailable or slow.
- Not recording the operating threshold, so it can be relaxed after launch without reopening the decision.
- Using a probabilistic filter to clear a class the programme said was unacceptable at a single instance.