Your team runs garak against every model release and publishes the failure rates to other teams. The detectors' error rates ship with the tool and you never measured them. How do you set policy for how those numbers may be used?
answer
- pin probes, detectors, tool version
- bridging re-run on config change
- standing sample audit, both sides
- retire families you cannot defend
- trend signal, never a release gate
basics
~20 sTreat the rates as a trend within one pinned probe and detector set, never as an absolute safety figure. Pin that configuration across releases, hand-audit a sample of failures and passes each cycle to estimate detector error, publish the audited numbers with sample sizes, and never let a rate alone gate a release.
solid answer
~60 sThree rules do most of the work. **Pin the configuration.** A rate is only comparable to another rate produced by the same probe set and the same detector set. Version that configuration alongside the results, and when you change it, publish a bridging run on the previous release so the step change is attributable to the tooling rather than the model. **Measure your inherited error, cheaply and repeatedly.** Each cycle, hand-score a fixed random sample of failures *and* passes for the probe families you actually report on. That gives a disagreement rate per family. You are not trying to compute precision as statistics - you are trying to know whether a family's numbers are worth quoting at all, and to notice when a family's disagreement rate starts drifting because the models' output style changed. **Constrain the claim.** Publish the tool rate, the audited rate and the sample size together, name the detector class, and say in one sentence what these detectors cannot see. Bar the number from being used as a clearance: a release decision needs triaged findings, not a rate. The organisational risk is not a wrong number; it is a right number read as a guarantee.
go deeper
Understands the number is not an absolute safety score and should not be repeated without saying how it was produced.
Pins the probe and detector configuration so releases are comparable, and reports the detector alongside each rate.
Adds a recurring hand-audit of failures and passes, publishes the audited rate with its sample size, and reproduces anything asserted as a finding.
Owns the whole regime: bridging runs when configuration changes, per-family disagreement rates tracked over time, retiring families that cannot be defended, splitting the engineering and leadership views, and refusing to let a rate become a release gate because that incentivises moving the detector rather than the behaviour.
### The situation, stated plainly You are operating an instrument whose calibration you did not perform, publishing its readings to an audience that will not read your caveats, on a cadence that turns those readings into a series. Every failure mode of this programme follows from those three facts, and policy has to survive all of them. ### 1. Comparability is achievable; absolute accuracy is not You cannot fix a shipped detector's true error rate - that would mean building and maintaining a labelled benchmark of your own, which is a bigger programme than the scanning is. What you *can* control is whether two numbers mean the same thing. So pin the whole configuration: probe selection, detector set (including extended detectors), garak version, generator settings and `--generations`, plus the target build. Store that record next to the results. Treat any change to any of it as a **break in the series**, not a data point. When you must change - and you will, because probe and detector catalogues move between releases - run a **bridging scan**: re-run the previous model release under the new configuration once, and publish both configurations' numbers for that release. Now the step change is measured and attributable to tooling. The cost is one extra full scan, which is the cheapest insurance in the programme: prompts x `--generations` calls at your endpoint's price, plus the wall-clock its rate limit imposes. Budget it as a line item, because the alternative is a quarter spent arguing about whether a model regressed. ### 2. A standing audit, sized to what you actually publish Each cycle, hand-score a fixed random sample of hits **and** passes for the families you report on, against a criterion written down in advance. Score both sides: hits reveal false alarms, passes reveal misses, and a policy that only samples hits will steadily drive the published rate downward while learning nothing about coverage. Cost it honestly. At one to two minutes per attempt, 100 scored attempts is three to four engineer-hours per cycle. That is affordable, and it caps how many families you can defend - which is a feature, because it forces the choice of what you publish to be deliberate. The output is a **per-family disagreement rate**, tracked over time. After a few cycles that series is the most valuable artefact the function owns: it tells you which of the tool's families are quotable, and it flags the moment a family starts drifting because model output style changed under it. ### 3. Retire what you cannot defend If a family's disagreement rate is high and stays high, stop publishing its rate. Keep running the family as a **lead generator** whose hits go into triage, and say so. Publishing a number you know is noise costs more credibility than the apparent coverage buys, and it is the thing a hostile reviewer will find. ### 4. Split the audiences Engineering wants candidates to triage: give them the raw attempts, the hit log and the queue. Leadership wants a direction of travel: give them the audited rate, its sample size, the pinned configuration and one plain sentence about what these detectors cannot see. Never give leadership the raw tool count on its own; it will be repeated without any of the surrounding conditions within a week. ### Where these numbers mislead, and the gate you must refuse The headline risk is not a wrong number - it is a right number read as a guarantee. Three concrete misreadings to pre-empt: a rate quoted as an absolute safety level when it is a function of your own probe and detector selection; a calibration Z-score read as a correction for detector error when it only ranks you against runs judged by the same detectors, blind spots included; and a fall in the rate read as improved behaviour when the app team merely reworded the boilerplate that used to trip a text matcher. Which is why the specific thing to refuse is wiring the rate into a release gate. The team that owns the number also owns the instrument that produces it, so a hard threshold creates direct pressure to move the instrument - a laxer detector, a dropped noisy family, reworded product copy - rather than the behaviour, and nothing in the tool defends against that. Gates run on triaged, reproduced findings. The rate stays a trend signal that triggers investigation. ### How you would check the policy is working Every published figure carries its configuration and sample size; per-family disagreement rates exist and are trending rather than being recomputed from scratch; at least one family has been retired from publication for noise; a bridging run exists at each configuration change; and no release has ever been approved on a rate alone.
- You must upgrade the tool and its probe catalogue mid-programme. How do you keep the published series honest?Re-run the previous release under the new configuration once, publish both configurations' numbers for that release, and mark the series break so the step is attributed to tooling rather than to the model.
- One probe family's audited disagreement rate is persistently high. What do you do with it?Stop publishing its rate and keep running it as a lead generator into triage. A number you know is noise costs more credibility than the coverage it appears to buy.
- Why is wiring the rate into a release gate the specific thing to refuse?It rewards moving the measurement - a laxer detector, a dropped family, reworded boilerplate - rather than the behaviour, and the tool provides no defence against that.
saying these in an interview costs you the question
- Publishes rates across releases while the probe or detector set changes underneath them.
- Lets a failure rate gate a release with no triaged findings behind it.
- Reacts to a bad number by swapping in a more permissive detector.
- Never measures how often the tool and a human disagree.
- Presents a garak rate to leadership as an absolute measure of model safety.