You wrote both the probe and the detector for a custom family in the garak LLM scanner, and the run reports a 22% failure rate. What do you do before that number goes into a report?
answer
- you are the only calibration
- sample flagged and unflagged
- misses hide in the cleared pile
- two labellers, disagreement = error floor
- labelled set becomes a fixture
basics
~20 sMeasure your own detector's error. Sample flagged replies and read them by hand for false positives, and sample unflagged replies for misses. Label a fixed set, ideally with a second reader, and quote the 22% with that measured error attached rather than as a bare number nobody has checked.
solid answer
~50 sNothing calibrates a detector you wrote but you. The number is a product of two things - the target's behaviour and your ruling criterion - and only one of them has been examined. The minimum honest procedure is a two-sided sample. Read a random sample of the flagged replies: every one that is not really a failure is a false positive, and that fraction bounds how much of the 22% is your detector talking. Then read a sample of the unflagged replies, the half people skip - misses appear nowhere in the report, so the cleared transcripts are the only place to find them. Have a second person label the same set independently; where you disagree, the criterion is under-specified, and that disagreement rate is part of the error you report. Then quote the rate with its measured error and method. Re-do it whenever the target changes materially.
go deeper
Says the number should be spot-checked by reading some flagged replies before trusting it.
Samples both flagged and unflagged replies and can explain that misses never surface in the report.
Runs a two-sided labelled sample with a second reader, keeps it as a regression fixture, and reports the rate with its measured error and method.
Sets the rule that no custom-plugin number leaves the team without a validation record, and decides how such rates may be compared to shipped ones.
### What the 22% is actually a product of A garak failure rate is the output of two instruments in series: the target's behaviour, and the criterion that ruled each reply. When both the probe and the detector are yours, exactly one of those instruments has ever been examined by anyone — and it is not the detector. A shipped detector at least carries the tool's history: other people have run it, its embarrassing false positives have been filed as issues. Yours has been run once, by you, and the number it produced is the first thing anybody has seen it do. So the work before the number ships is not more scanning. It is measuring the ruler. ### The validation pass 1. **Write the failure criterion as one sentence** a colleague can apply to a reply without reading your code. If you cannot write it, the detector is not testable and neither is the rate. 2. **Sample both sides, stratified.** Draw roughly 50 flagged and 50 cleared replies at random from the run. The flagged side yields false positives, which bound how much of the 22% is your detector talking. The cleared side yields misses. Reading only the flagged side is the common half-job, and it hides precisely the errors that make a system look safer than it is. 3. **Two independent labellers on the same set.** Their disagreement rate is a floor on the precision of anything you quote: you cannot report the target more precisely than two humans can agree on the criterion. 4. **Freeze the labelled set as a fixture.** It becomes a regression test for the detector. Change the matching logic, re-score the fixture, and see what moved — in the detector, not in the model. 5. **Report the rate with its error and its method**, not as a bare figure. ### The reweighting trap, worked This is where practitioners go wrong even when they do the labelling. Suppose the run produced 1,200 outputs and flagged 264 of them, giving the headline 22%. In the 50-item flagged sample, 42 are genuine failures — precision 0.84. In the 50-item cleared sample, 3 are genuine failures — a miss rate of 0.06 across the 936 cleared outputs, about 56 missed failures. Corrected estimate: 264 x 0.84 = 222 true positives, plus roughly 56 missed, is about 278 real failures, or 23%. Note what just happened. The headline was 22% and the corrected figure is 23% — nearly identical — while the detector was wrong about one flag in six *and* missing one real failure in eighteen cleared replies. Two substantial errors in opposite directions very nearly cancelled. The total agreeing is not evidence that the ruler is right, and if the target's behaviour shifts so that one error grows and the other does not, the cancellation stops and the rate lurches for reasons nobody will attribute correctly. The second half of the trap: because the sample was stratified 50/50 over a population that is 22/78, its raw failure share means nothing about the run. It must be reweighted by each stratum's true share, exactly as above, or you are reporting a number about your sampling design. ### What this costs Labelling 100 replies carefully runs about ninety minutes; two labellers plus adjudication of the disagreements is most of a working day, and it recurs whenever the detector or the target changes materially. That is the real budget line on a custom family, and it is why "we wrote a probe over the weekend" is only ever a report on the cheap half. A 50-item stratum also carries a confidence interval of roughly plus or minus ten points at 95% — so "precision 84%" is honestly "somewhere in the seventies to low nineties", and a report that states 0.84 as a fact overstates what 50 items can support. ### The rest of what I would check The sample must come from **this** run's target: a detector validated against one deployment does not carry over to another with a different system prompt, formatting or refusal style. And a custom family's rate does not belong in the same table as shipped families' rates without a footnote — they were ruled by different, differently-validated criteria, and a shared table invites a comparison nobody has earned. The senior instinct here is to treat the detector, not the model, as the instrument under test, and to publish its calibration alongside its readings.
- Why is sampling the unflagged replies the half people skip, and what does skipping it cost?Misses leave no trace in the report, so nothing prompts you to look. Skipping it means you can only ever discover over-reporting, and the system looks safer than it is.
- Your second labeller disagrees with you on a fifth of the sample. What does that tell you?The failure criterion is under-specified. Tighten the written definition and re-label before quoting any rate; the disagreement rate bounds how precisely you can report at all.
- How would you keep the detector honest across releases?Freeze the hand-labelled set as a fixture, re-score it whenever the detector or the target changes, and treat a shift on the fixture as a detector problem until proven otherwise.
A scale that reads heavy on one pan and light on the other can still balance, and the balance proves nothing about either pan. A detector that over-flags and under-flags in roughly equal measure lands on a plausible total the same way — the headline being right is not evidence that the ruling is.
saying these in an interview costs you the question
- Only reviews flagged replies, never the cleared ones
- Quotes the bare rate with no measured detector error
- Treats one person's spot-check as validation
- Assumes a detector validated on one deployment stays accurate on another
- Puts the custom family's rate in the same table as shipped families without a caveat