Your success rate was measured by calling a model API directly with your harness's own system prompt. Production wraps that same model in a fixed template, a 300-character user field and an input classifier. What does that do to the severity you file?
answer
- a rate belongs to an entry point
- re-measure through production
- template + field length + classifier
- wrapper can add surface too
- never divide by a guessed filter rate
basics
~20 sIt means your rate describes your harness's entry point, not the product's. Say so explicitly in the finding, then re-measure through the production path before you finalise the rating. If you cannot re-measure, file the number with the measurement point stated and rate the reachability separately, rather than quietly discounting it.
solid answer
~50 sSeverity is about the deployed product, so the reachability question is: can a real user, through the real interface, get the same result? - **The fixed template and short field** shrink what an attacker can inject; many payloads will not fit or land where the template neutralises them. - **The input classifier** removes some fraction of attempts before the model sees them — but it is itself probabilistic, so it lowers the rate rather than zeroing it. - **Occasionally the wrapper makes things worse**, by concatenating retrieved documents or other user-controlled fields the raw API call never carried. So re-run the same attempt set through the production entry point and file that rate as the product's, keeping the direct-API rate as supporting evidence. What you must never do is divide the direct-API rate by a guessed filter effectiveness — that invents a measurement. If re-measuring is impossible, state the limitation and rate on impact plus a flagged reachability assumption.
go deeper
Should say the test did not go through the real product path, so the number may not apply to users.
Names the specific differences — template, field length, classifier — and asks to re-measure at the production entry point.
Files both rates with their measurement points, notes that the wrapper can add surface as well as remove it, and marks which mitigations the rating depends on.
Requires a declared measurement point on every filed rate and defines when a finding may be rated on an unverified reachability assumption.
This is the most common way an AI finding's severity goes wrong, and it goes wrong in both directions. It is a reachability question, not a statistics question: no amount of extra sampling fixes a rate measured at the wrong door. ### A rate belongs to a pair, not to a model A success rate is a property of *(attack set, entry point)*. The **entry point** is the interface an attacker actually reaches — the product's chat box, an API your customers hold a key for, a support widget. "87% against the raw model API with a permissive system prompt" and "4% against the product's chat box" can both be true of the same defect on the same day. A finding that does not name which one it measured cannot be triaged at all, because the reader cannot tell whether a customer can do this. ### What the production path changes, mechanism by mechanism - **The fixed template.** It places the product's system instruction ahead of the user's text, often with the user text sandwiched or delimited, and it replaces the harness's own permissive system prompt. This is frequently the single largest term in the difference, and it is invisible in a direct API test because the tester supplied that system prompt themselves. - **The 300-character user field.** A length limit, not a semantic control. Longer attempts arrive truncated and malformed; anything that depended on volume or on a long preamble simply stops working at the boundary. It moves the rate; it does not zero it. - **The input classifier.** Itself a probabilistic control with its own false-negative rate *on your specific attempt distribution* — a number almost certainly never measured, because vendors quote benchmark performance on their own distribution, not yours. - **An output classifier, if present.** It can suppress a response the model already produced. That attempt is a model failure where a control saved you, and it deserves its own outcome category in the finding rather than being folded into the misses. ### Where the wrapper makes the finding worse Production wrappers add content the direct call never carried: retrieved documents, tool and function outputs, prior conversation turns, uploaded files and OCR text, and in shared contexts other tenants' data. Each is attacker-influenceable text arriving inside the prompt. A product can therefore be strictly *more* exposed than the bare model, and a tester who only hit the raw API may have understated the finding rather than overstated it. Assuming the wrapper can only subtract is the second-order version of the same mistake. ### What re-measuring costs Through production you pay in wall clock and in blast radius, not in tokens. The product path enforces authentication, session limits and per-account throttles, so 200 attempts that took four minutes against the API may take hours through the app, may need several seeded accounts, and may trip abuse detection or page an on-call responder. Coordinate it, and log the exact attempt window so the alerts can be reconciled afterwards. Budget it as an engagement task with a named owner, not as "just rerun it". ### Where the number misleads The signature error is arithmetic dressed as evidence: "the classifier catches about 80%, so 87% becomes roughly 17%". That invents a measurement. The classifier's pass-through rate on *these* attempts is precisely the quantity nobody has, and a guessed divisor produces a precise-looking figure with nothing behind it. The mirror error is quoting the direct-API rate to a product owner, who will read any number in a security finding as customer-reachable, and will either panic or — after someone points out the harness prompt — discount the whole finding. ### What to file Two lines with two measurement points: the rate measured at the production entry point, with counts and date, as the rating input; and the direct-API rate as supporting evidence about the model's own behaviour. Name the mitigations in the path that are load-bearing, because the severity is conditional on them and must be re-rated if the classifier is disabled or the field length raised. If you genuinely cannot re-measure before the engagement ends, say so plainly — "measured at the model API only; production reachability not verified" — rate on impact with the reachability assumption flagged, and attach a specific re-measurement request to the owning team. An honest bound beats a precise number that describes the wrong system.
- Give a case where the production path makes the finding worse, not better.When the wrapper concatenates retrieved documents, tool output or another user's content into the prompt. Those are attacker-influenced inputs the direct API call never carried, so the product has a surface the bare model did not.
- The output classifier suppresses the response, but the model still complied. Same severity?Usually lower, because no user saw the content, but it is still a finding: the compensating control is now the only thing preventing disclosure, and the rating should say the severity is conditional on it.
- You have two hours left in the engagement and cannot re-run through production. What do you file?The direct-API rate with the measurement point stated, impact rated from a single success, reachability marked unverified, and a specific re-measurement request assigned to the owning team.
It is the difference between testing a lock on the workbench and testing it fitted in the door: the frame may hold it shut, or there may be a letterbox beside it that the workbench test never saw.
saying these in an interview costs you the question
- Filing a direct-API rate as the product's rate because "it is the same model"
- Applying an invented discount factor for the filter instead of measuring
- Assuming the wrapper can only reduce exposure and never adds injection surface
- Leaving the measurement point out of the finding entirely