The provider of a hosted model closes your submission as expected model behaviour rather than a vulnerability, but your client's application still exhibits the effect. What do you do with that item in the client's report?
answer
- ruling changes type, not truth
- property of the dependency
- rate the client's capability, not the model
- control at an owned layer
- residual risk written down
- ruling as procurement evidence
basics
~20 sIt stays in the report, re-framed. The item is no longer waiting on a fix; it is a standing property of the platform your client chose. Rate it against the client's application, describe a compensating control the client can apply themselves, and record the residual risk that is left once that control is in place.
solid answer
~1 minAn "expected behaviour" ruling changes the item's *type*, not its truth. Rewrite it from "defect pending upstream fix" to **an accepted property of the dependency, with a client-side control and a residual risk**. Concretely: keep the reproduction and the rate; move the rating onto the client's asset, since the impact is a function of what the application lets the model do, not of the model in the abstract; propose a control at a layer the client owns (narrow the tool or permission the output can reach, require a human confirmation on the sensitive action, filter the specific effect at the boundary, or remove the capability from the product); then state honestly what the control does not cover — usually it reduces frequency rather than removing the behaviour. Two further moves matter. Record the ruling itself as an artefact, with date and wording, because it is procurement evidence: if this is intended behaviour, then choosing this dependency accepts it. And schedule a re-test, since intended behaviour today can change silently and your control may quietly stop being necessary — or stop being sufficient. Do not escalate the ruling as if it were an appeal you can win on the client's behalf. You are not their counterparty.
go deeper
Knows the finding does not disappear just because the provider declined it, and that the client still needs advice.
Reframes the item as a platform property with a client-side mitigation, instead of leaving it as pending-upstream.
Moves the rating onto the client's capability, distinguishes controls that remove the outcome from those that only shift probability, and writes the residual risk and a re-test date.
Treats the ruling as procurement evidence and drives the accept-mitigate-or-change-supplier decision, including what the firm does when a whole class of findings comes back as intended.
## What the ruling is, mechanically "Expected behaviour" is a statement about the provider's bug bar, not about your client's risk. It means the observed output falls inside the envelope the provider considers the model's documented or intended behaviour, so no engineering work will be scheduled and the ticket closes. You cannot appeal it: you are not the provider's counterparty, your client may be. The finding's *truth* is untouched — the transcripts and the rate are exactly as valid the day after the ruling as the day before. What changed is the item's type: from "defect pending an upstream fix" to "a standing property of a dependency the client chose." ## Reframe the item Delete every phrase that implies something is coming. The heading becomes a property statement: under these conditions, at approximately this rate, this platform will do this. Everything below that sentence is now the client's decision. Then move the rating. "The model can be talked into describing X" is not rateable — there is no asset in the sentence. "This assistant, which is permitted to issue refunds, can be induced to issue one outside policy in roughly 12% of 200 crafted multi-turn attempts" is rateable, because severity lives in the capability the application granted, and that capability is precisely what the client can change. A provider verdict must never move the severity; severity follows impact on the client's system. ## Control classes and what each actually costs | control | cost | what it buys | |---|---|---| | remove the capability | the feature is gone; a product decision, not an engineering one | the only control whose effectiveness does not decay | | gate behind a human step | latency plus support headcount, measured in tickets per day times handling minutes | changes the outcome, but the business will push for a threshold that quietly restores most of the exposure | | constrain scope (tool, data, per-account limits) | roughly a day of engineering, no per-request cost | caps blast radius rather than frequency | | detect and block at the boundary | a second inference per request: a comparable-size judge model roughly doubles per-request cost and adds a few hundred milliseconds; a small fine-tuned classifier is single-digit milliseconds and cheap, with a worse error profile and its own threshold to tune | shifts a probability, and degrades against rephrasing | The first two change what a success is worth. The last two change how often a success happens. Say out loud which class you are offering, because a client who believes a filter closed the item will be unpleasantly surprised. ## Where the residual-risk number misleads Reports routinely say "the boundary filter reduces the observed rate from 24% to 3%", measured by replaying the same fixed suite through the filtered pipeline. Three things are wrong with reading that as a 3% residual. **Same-suite fitting.** The filter was tuned while looking at those prompts, so it has effectively been fitted to its own evaluation set. A held-out set of paraphrases written by someone who did not build the filter usually comes back materially higher. **Non-adaptivity.** The suite does not respond to the control. A person does: they observe what gets blocked and rewrite. A 24-to-3 drop against a frozen prompt list is not a claim about an adversary. **Denominator.** 3% of crafted attempts is not 3% of traffic, and a business that hears the latter has swapped the denominator in the direction that flatters the report. Write what the fraction is *of* into the same sentence as the number. One more measurement trap: a pre-control rate and a post-control rate taken weeks apart are not comparable if the endpoint drifted between them. Where you can, re-measure the unfiltered path in the same window as the filtered one and report both, so the delta is attributable to your control rather than to the provider quietly changing something. ## The ruling as procurement evidence, and the re-test Capture the ruling verbatim with its date. It converts a technical finding into a supplier input: if this is intended behaviour, then continuing on this dependency is a documented acceptance, and that is a conversation an owner can have with a vendor that you cannot have for them. Then record the item as an **accepted risk** with a named owner, the control, the residual risk, and a review date — not as closed, because closed means the risk is gone. Schedule a re-test of both the behaviour and the control. Intended behaviour today can drift tomorrow: your control may stop being necessary, or quietly stop being sufficient. A quarterly re-measure at 200 attempts with a judge is a few hundred calls and an afternoon, which is cheap next to discovering the drift from an incident. ## What to check Was the control re-measured on held-out prompts rather than the ones that shaped it? Is severity unchanged by the provider's verdict? Is the item recorded as accepted-with-an-owner rather than resolved? Is the ruling quoted with its date? And does the client understand which control class they bought?
- The client asks you to mark it closed since the provider says it is intended. What do you say?Closed means the risk is gone or accepted. Here it is accepted, so record it as an accepted risk with an owner, the compensating control, and a review date, not as a resolved finding.
- How does an expected-behaviour ruling help the client at all?It removes the option of waiting. It is dated evidence that the property is part of the dependency, which turns the decision into a procurement and design one the client can actually make.
- Which control class should you prefer when you can only pick one?Removing or gating the capability, over detecting the effect. Detection shifts a probability against an adaptive input; removing the capability changes what a success is worth.
saying these in an interview costs you the question
- Dropping or downgrading the finding because the provider declined it.
- Rating the model in the abstract instead of the impact on the client's application.
- Offering an output filter and implying it closes the item.
- Promising the client you can get the ruling reversed.
- No re-test date, on a dependency that can change silently.