How do you turn an LLM risk map into prioritized controls and accepted residual risk?
answer
- likelihood is close to one
- rank by blast radius, not by model behaviour
- two columns: deterministic versus probabilistic
- irreversible and undetectable move to the top
- the downstream owner signs the remainder
basics
~20 sRank by the blast radius of the worst action the system can actually take, not by how likely the model is to misbehave — against an adaptive attacker that likelihood is near one. Then separate deterministic boundaries from probabilistic mitigations and get the downstream owner to sign the remainder.
solid answer
~50 sStandard likelihood-times-impact scoring misleads here, because the honest likelihood of a successful injection against an adaptive attacker is close to certain. So invert it: for every privileged sink, ask what the worst realizable action is, whether it is reversible, and whether you would detect it. Rank by that. Then split every proposed control into two columns — **deterministic** boundaries that hold regardless of model behaviour, and **probabilistic** mitigations such as classifiers or prompt hardening that lower the rate but cannot bound it. A finding is only closed by something in the first column; the second column reduces frequency and buys detection time. Whatever is left is residual risk, and it must be accepted in writing by the owner of the system being acted upon, not by the AI team. Reversibility, audit trail and a kill switch are the compensating controls that make acceptance defensible.
go deeper
Know that not every risk gets fixed and that some are consciously accepted, and that the strongest fix is usually removing permission rather than adding a filter.
Explain why likelihood-based scoring is weak here and be able to sort a set of findings by blast radius, reversibility and detectability.
Demonstrate the deterministic-versus-probabilistic split on real controls, and show you would not report a finding as closed on the strength of a classifier or a prompt change.
Own the governance: who signs residual risk, what compensating controls make acceptance defensible, how approval-based controls decay at volume, and the funding order that puts scope reduction ahead of detection tooling.
## Why ordinary risk scoring breaks here Most risk registers rank by likelihood times impact. Applied to an LLM feature, the likelihood axis collapses. Published adaptive-attack work has broken a dozen deployed injection and jailbreak defenses at over 90% success, and even frontier models that sit near a tenth of a percent on a single attempt rise to several percent across a hundred adaptive attempts. If your threat actor is motivated and can iterate, the honest likelihood of eventually getting the model to emit a chosen action is not a number you can meaningfully discount. Ranking by it produces a register where everything is high and nothing is ordered. The useful axis is **consequence**. For every privileged sink the system can reach, ask three questions: 1. **What is the worst realizable action?** Not the worst imaginable — the worst the granted scopes actually permit. If the credential can only read, a hijack is an exfiltration risk; if it can write to a control system, it is a safety risk. 2. **Is it reversible?** A drafted email that a human sends is different from an email already sent; a scheduled shutdown is different from an executed one. Irreversibility is the single strongest priority multiplier. 3. **Would you detect it, and how fast?** An action with a complete audit trail and an alert is survivable at a risk level that a silent one is not. Rank the register by that triple. It almost always reorders the list away from the model-centric findings people arrive with and toward the two or three sinks that turn text into consequence. ## The two columns every control belongs in Split every proposed control: **Deterministic boundaries** run outside the model and behave the same whatever the model was persuaded to emit: authorization enforced per identity at the tool layer, scopes that make the dangerous call impossible rather than discouraged, egress restricted to an allowlist, an irreversible action that requires a second system's assent. These are the only controls that close a finding. **Probabilistic mitigations** depend on model or classifier behaviour: injection detectors, output filters, prompt hardening, instruction-hierarchy training, an LLM judge on the trajectory. They genuinely lower the attempt-to-success rate and they generate the signal you detect on. They never close a finding, and a register that shows them as "mitigated — closed" is misreporting. This split is what makes prioritization tractable: for the top-ranked sinks, spend on column one; for the long tail, column two plus monitoring is proportionate. ## Controls that look strong and are not A human-approval gate is the classic overcounted control. It is genuinely valuable on rare, high-consequence actions. On a high-volume path it decays: reviewers who approve fifty requests a day approve the fifty-first without reading it, and the fluent, confident summary the agent attaches makes that easier. If a whole risk is banked on human review, the register should record the expected degradation and cap the volume the gate is asked to carry. Likewise, "we log everything" is a detection control, not a prevention control, and only counts if someone or something reads the logs. ## Residual risk needs the right signature After deterministic boundaries are in place, something remains — usually the possibility that a legitimate, in-scope action is triggered by hostile content. That residual has to be **accepted explicitly, in writing, by the owner of the system being acted upon**. On a plant copilot that can schedule a line shutdown, the accepting party is plant operations, not the AI team, because operations owns the consequence and is the only function that can weigh it against production risk they already manage. AI teams accepting risk on behalf of a downstream system is the most common governance defect in this space, and it is what turns an incident into a blame argument. Acceptance is only defensible when it comes with compensating controls: an audit trail that attributes every action to the agent identity that took it, a documented kill switch with a named owner and a tested activation path, blast-radius caps such as rate or value limits on the action itself, and a review date. "Accepted" without those is "ignored" with paperwork. ## Sequencing the work A workable order for a team that has just produced its first map: 1. Remove authority nobody needs. The cheapest risk reduction is a scope that never existed, and it usually costs nothing in product value. 2. Put deterministic enforcement in front of the top-ranked irreversible sinks. 3. Make every action attributable and alertable, so the probabilistic layer has somewhere to report to. 4. Add probabilistic mitigations across the breadth of the surface. 5. Record what remains, with an owner, a compensating control set and a date. Step 1 before step 4 is the part teams reverse. Buying a detection product before removing an unnecessary write scope is spending on the column that cannot close the finding. ## What a good register looks like Each line names a sink, a worst realizable action, its reversibility, the paths that reach it after consuming untrusted content, the deterministic control at the crossing, the probabilistic controls layered on, the residual position, and the signing owner. It is short — usually a handful of lines that matter and a long tail marked not applicable with reasons. A register of forty undifferentiated findings is a sign the consequence ranking was never done.
- How do you argue against a team that wants to buy an injection-detection product as its main control?Accept it as useful and place it correctly. It lowers attempt-to-success rates and gives you a detection signal, which is real value. But show the adaptive-attack evidence that such classifiers fall to attackers who iterate, and point at the sink: if a successful bypass still reaches an irreversible action, the product changed the frequency and not the outcome. Fund the scope reduction first, then buy the detector as depth.
- What makes a residual-risk acceptance defensible six months later during an incident review?Four things on the record: the specific action accepted and its blast radius, the compensating controls in place at the time, the named signer who owned the downstream system, and a review date that had not lapsed. Reviews go badly when acceptance was implicit, signed by the wrong function, or written broadly enough to cover an action nobody had actually considered.
- How does this ranking change for a read-only assistant with no tools?The consequence axis shifts from action to disclosure, so the ranking runs over what the system can read and where its output can travel. The top findings become cross-tenant retrieval, sensitive content in context, and any rendering path that can carry data outward. The two-column control split is unchanged; the deterministic column is now scoped retrieval and egress restriction rather than tool authorization.
saying these in an interview costs you the question
- Ranks findings by how likely the model is to be jailbroken
- Marks a finding closed because a classifier was deployed
- Has the AI team accept risk on a system it does not own
- Counts a high-volume human-approval gate as a full control
- Buys detection tooling before removing unnecessary write scopes