A hosted moderation service returns only "allowed" or "blocked" for each prompt, with no numeric score. How do you still show a client where that guard breaks across its strictness range instead of handing over one bypass percentage?
answer
- sweep settings, not scores
- step curve, vendor-chosen resolution
- one rerun per setting = N x cost
- never paint your classifier's score on their verdicts
- allowed vs. attack-actually-worked columns
basics
~20 sSweep whatever the service does expose instead of a score: each selectable strictness level and each category toggle. Run the same corpus at every setting and report a step curve over settings rather than a smooth score curve. Pin each configuration exactly, and state that the resolution is limited to the settings offered.
solid answer
~60 sYou cannot recover a continuum the service never gave you, so change what you sweep. **Sweep the settings, not the scores.** Enumerate the discrete configurations the service exposes — strictness or severity levels, category enable/disable — and run the identical corpus at each. The result is a step function with as many points as the vendor offers, which is coarse but honest. **Keep each step reproducible.** Record the exact configuration per step, not a label, because a step curve whose steps are described in prose cannot be re-derived. **Do not fabricate the missing axis.** It is tempting to score prompts with a second classifier you control and present that as the guard's score. That measures your classifier, not the guard, and the report must never blur the two. It is legitimate as a separate ordering used to describe *which kinds* of prompts survive at each step — labelled as your measurement. **Say what the coarseness costs.** With three strictness levels you learn nothing about behaviour between them, and if the operator retunes the meaning of a level, every step moves. Both belong in the limitations paragraph, not in a footnote.
go deeper
Should recognise that with no score you have to rerun at each available setting rather than recompute from stored numbers.
Enumerates the exposed settings, keeps the corpus and judge fixed across steps, and records each configuration verbatim.
Budgets the N-fold query cost, samples where necessary, separates allowed-through from attack-actually-worked, and refuses to present a self-supplied score as the guard's own.
Weighs a verdict-only hosted guard against a score-returning one as a measurability requirement in procurement, since the first makes every future comparison cost a full rerun.
## The instrument question: what continuum do you actually have? Every "show me the curve" request reduces to one thing — which axis the guard lets you vary. A guard that returns a per-prompt **score** hands you a dense axis you can sweep in post-processing at zero extra query cost. A guard that returns only a **verdict** (allowed/blocked) has already collapsed that axis at a threshold the operator owns, and nothing in the response lets you recover it. What remains is the axis the vendor chose to expose: named strictness or severity levels, and per-category enable/disable toggles. So you sweep the settings instead of the scores, and the deliverable is a **step function** with as many points as the vendor offers — coarse, but honest and reproducible. ## The deliverable One row per exposed configuration, and two count columns, not one: | setting | prompts allowed through | judged successful attacks | |---|---|---| | level 1 (permissive) | 180 / 900 | 96 | | level 2 (default) | 74 / 900 | 41 | | level 3 (strict) | 22 / 900 | 15 | The two columns are the point. "The guard allowed it" and "the attack worked" are different events, and the gap between them is informative on its own: a guard that allows many prompts whose attacks then fail is being carried by the model behind it, not by the guard, and that changes what happens if the model is swapped. ## What it costs This is where a verdict-only guard hurts. With a score you pay one pass; without one, each point on the curve is another pass over the corpus, so an N-level sweep multiplies **guard calls** by N — 3 levels × 900 prompts = 2,700 guard calls, against 900 for a score-returning guard. The saving most teams miss: the target-model completions and the judge calls do **not** multiply. A prompt's completion depends on the prompt and the model, not on which strictness level let it through, so cache completions and judgements keyed by prompt and reuse them across settings. You end up paying the completion and judging cost once for the union of prompts ever allowed — closer to 180 than to 2,700 in the table above — while the guard calls carry the N-fold multiplier. Budget on that shape before promising a curve, and when the corpus is large, sweep every setting on a stratified sample and run the full corpus only at the setting actually in production, reporting the per-step sample size so nobody reads a sampled row as a census. ## Where the step curve misleads - **Levels are not a scale.** `medium` is not halfway between `low` and `high`; the gaps are vendor-defined, undocumented, and unequal. Any language implying interpolation between steps ("roughly 15% at a mid setting") is fabricated. - **The strictest row is a floor with nothing under it.** "Only 1.7% got through at the strictest level" reads as reassuring until you note there is no stricter setting to buy — that residual is what you ship with, and it is the number that belongs in the summary. - **A level's meaning can be retuned server-side.** With a hosted verdict-only guard you own neither the model nor the definition of the levels, so a re-measurement after an operator-side change is a *new* measurement, not a check of yours. Record the run window and any version identifier the response carries, and give the result a stated shelf life. - **The manufactured axis.** The tempting move is to score every prompt with a classifier you control and plot the service's verdicts against those scores. It produces a smooth, plausible, publishable chart that is not the guard's operating characteristic — it is your classifier's ordering with somebody else's verdicts painted on it. It is legitimate only as an explicitly labelled descriptive device ("at the strictest level, the prompts still getting through are the ones my ordering ranks mildest"), and the report must not let a reader infer the vendor exposes a score. ## What to check Before plotting the steps, verify the steps are real: run a small calibration slice at every level and confirm the verdict sets actually differ. On more than one service, the categories do the work and two adjacent levels return identical verdicts on a given corpus — in which case your step curve has fewer genuine points than it has rows, and drawing all of them implies resolution you do not have. Then confirm that each row's configuration was captured verbatim rather than by label, so a later reader can reconstruct which configuration produced which row.
- Why is plotting the guard's verdicts against a classifier of your own not the guard's curve?Because the axis is your model's ordering. It describes which prompts survive, but it is not the guard's internal score and must not be presented as one.
- The corpus is large and the service charges per call. How do you still deliver a step curve?Sweep every exposed setting on a stratified sample, then run the full corpus only at the setting actually in production, and report the sample size per step.
With a score-returning guard, sweeping the threshold is re-reading a temperature log you already have; with a verdict-only guard, every point on the curve means going back outside and taking the reading again.
saying these in an interview costs you the question
- Presenting a second classifier's scores as if they were the guard's operating characteristic.
- Promising a smooth curve for a guard that exposes only discrete settings.
- Ignoring that each extra setting costs a full rerun of the corpus and its query budget.
- Labelling steps by name only, so nobody can reconstruct which configuration produced which row.
- Reporting allowed-through counts as bypasses without checking the attack actually worked.