skip to content

How do you decide an LLM assistant is safe enough to launch, given ASR never reaches zero?

level: principalimportance: should knowfreq 28%

answer

  1. zero is not the bar
  2. risk is bounded, not eliminated
  3. report the curve, per category
  4. containment converts rate into tolerable harm
  5. ship with detection and a kill switch

basics

~20 s

You do not launch on a zero. You launch on a bounded worst case: a per-category attack-success curve against a stated attacker budget, evidence that a success cannot cause unrecoverable harm, and a detection and response plan sized for the residual rate you are accepting.

solid answer

~50 s

There is no industry threshold for "safe enough", so the decision has to be constructed rather than looked up. Three pieces of evidence carry it. **Measured residual risk:** attack-success rate per severity category against an adaptive attacker with a declared budget, presented as a curve against attempts rather than a headline number. **Bounded harm:** for each category where success is still possible, what the worst realised outcome is — a leak of one record versus a whole corpus, a reversible write versus an irreversible payment — because containment, not detection, is what converts a nonzero rate into an acceptable one. **Response readiness:** logging that lets you reconstruct a trajectory, a way to detect the class in production, a kill switch, and a staged rollout so early failures land on a small cohort. Spend limited pre-launch time on adaptive attacks and autonomous auditing over enlarging a static list, and name explicitly which residual rate you are choosing to carry and who owns that choice.

go deeper

for a junior

Know that safety sign-off is never a zero-risk claim, and that a passing red-team run is one input rather than the decision itself.

for a middle

Be able to describe the evidence a launch decision needs: attack-success rate per category with the attacker budget stated, and what the worst outcome is when an attack does land.

for a senior

Show how you would bound the harm rather than chase the rate — narrowing tool scope, removing exfiltration channels, staging the rollout — and how monitoring and a rollback path are part of the launch package, not follow-up work.

for a principal

Own the decision itself: which categories block unconditionally, what residual risk the organisation is accepting in plain language, who signs it, how the limited red-team budget was allocated, and where the field genuinely has no consensus to lean on.

## The question has no lookup answer As of 2026 there is no accepted attack-success-rate threshold that means safe. Anyone who quotes one is quoting a number their own team made up. That makes launch sign-off a judgement call built from evidence, and the mark of seniority is constructing that judgement transparently rather than hiding behind a percentage. Start from the honest premise: attack-success rate is a function of attacker effort, and for anything with a public interface the effort available is unbounded. A frontier assistant reported around 0.1% success on a single attempt and roughly 5–6% after a hundred adaptive attempts; the shape, not the endpoint, is the fact. So the decision is never *is the rate zero* — it is *what happens the fraction of the time it is not, and can we live with that*. ## The three things sign-off actually rests on **1. Residual risk, measured and sliced.** Produce ASR against attacker budget, per severity category, judged by checks you trust. Cross-tenant data disclosure, unauthorised irreversible action and credential exposure get absolute counts; softer categories get rates. This is the only part of the package that is a measurement, and it is worthless without the budget attached. **2. Bounded worst case.** For every category where success remains possible, write the sentence that describes the realised harm. "An injected instruction can cause the assistant to summarise the current case file back to the user" is very different from "an injected instruction can cause the assistant to email a case file to an arbitrary address", even at identical ASR. The lever that changes the second sentence into the first is architectural — narrowing tool scope, removing the outbound channel, requiring approval for irreversible actions — and it is what makes a nonzero rate survivable. A launch package where the mitigation story is entirely "our classifier catches it" should not clear the bar, because detection-based defenses fall to adaptive attackers and are graded accordingly. **3. Detection and response.** Since you are shipping with a known residual rate, you are committing to catch and handle the cases it produces. That means trajectory-level logging you can actually reconstruct an incident from, a monitoring signal for each blocking category, a rollback or kill path for the feature, and a staged rollout so the first realised failure hits a small cohort. Shipping with 3% residual ASR and no detection is a worse position than shipping with 8% and a same-day response path. ## Spending a fixed pre-launch budget Given, say, three weeks before a public launch, the allocation that produces the most decision-relevant evidence: - **Adaptive attacks on the blocking categories first.** Depth on the two or three outcomes that would actually stop the launch beats breadth across twenty categories that would not. - **Autonomous auditing agents.** Tools such as Petri explore the system on their own and surface behaviours nobody wrote a probe for. Their unique value is coverage of the unanticipated — which, by construction, is where the failure that embarrasses you lives. Their output is also noisy, so budget triage time. - **Agent-level cases in a sandboxed copy of the real tool environment**, since a model-layer result does not describe your product. - **Least of all, growing the static list.** It is the cheapest activity and the one that most reliably produces a comforting number without new information. ## Contested ground — say so Several parts of this are genuinely unsettled and pretending otherwise is a tell. There is no consensus threshold. There is no agreed way to weigh a low-probability catastrophic outcome against a high-probability minor one for consumer products. Human approval as a control degrades through approval fatigue, so "a person reviews it" is a weaker mitigation at scale than it looks on a slide. And the regulatory floor is moving — transparency and general-purpose-model enforcement duties land in 2026 with higher-risk obligations phased later — so a package that clears an internal bar may still need a compliance answer. The right posture is to state the tradeoff shape and the assumptions rather than manufacture a consensus. ## What the sign-off should look like A one-page decision with: the categories that block and their measured counts at a named attacker budget; the worst realised harm per category with the containment that bounds it; the monitoring signal and response path for each; the rollout plan; the residual risk being accepted in plain language; and a named owner who accepted it, with a review date. It should be readable by someone who was not in the evaluation, and it should be the artefact you reread after the first incident — because the eventual postmortem's most useful question is whether the failure was outside the risk you accepted or inside it and simply undetected.

  • Two candidate mitigations, same measured reduction in attack-success rate: one narrows a tool's scope, one adds an input classifier. Which do you prefer and why?
    The scope change, decisively. It is a deterministic property of the system that holds regardless of how clever the input is, whereas a classifier is a probabilistic filter that adaptive attackers bypass — the 2025 work on adaptive attacks pushed published detection defenses past 90% success. Equal measured reduction today does not mean equal reduction against an attacker who optimises against your filter tomorrow.
  • A safety failure reaches production despite sign-off. What does the postmortem need to establish?
    First, containment and scope: what data or actions were affected, and shut the path. Then the trajectory — which content carried the payload, which tool fired, which control was supposed to stop it. The decisive question is whether this failure was inside the residual risk you accepted and merely undetected, or outside it, because the first is a monitoring gap and the second is a modelling gap. Both end in a permanent regression case.
  • How much weight should an autonomous auditing agent's findings carry in the launch decision?
    High weight for coverage of the unanticipated, low weight as a pass signal. Its unique contribution is behaviours no written probe anticipated, so a serious finding should block; but a clean auditing run proves only that one explorer did not find anything within its budget, which is a weaker claim than a measured attack-success curve. Treat it as a discovery tool feeding your categories, not as certification.

saying these in an interview costs you the question

  • Claims a launch is safe because the red-team suite passed
  • Quotes an industry-standard acceptable attack-success rate that does not exist
  • Relies on a detection classifier as the sole mitigation for a blocking category
  • Ships with a known residual rate but no monitoring or rollback path
  • Treats human approval as a durable control without accounting for approval fatigue

context