How can a garak detector record a pass on a reply that was in fact a successful attack, and what does that imply about how you present the run's failure rate?
answer
- misses are silent; false alarms are loud
- form, distribution, capture
- wrong reply field logged = clean run
- sample the passes to size the gap
- clean run is not a clearance
basics
~20 sIt happens whenever the harmful content arrives outside what the detector looks for: another language, an encoding, code or pseudo-code, a summary, or wording the matcher was never written for. Those attempts score as passes. So the reported failure rate is a floor on the model's weakness, never a measure of it.
solid answer
~60 sA detector can only rule on evidence it recognises. Misses come from three directions: - **Form.** The content is present but expressed as code, a table, a translation, an encoding, or a step-by-step summary rather than the prose the matcher expects. - **Distribution.** A classifier detector meets replies unlike its training data - truncated, multilingual, code-heavy, unusually long - and lands on the wrong side of a boundary you never set. - **Pipeline.** The reply the detector saw was not the reply a user would see: only the first field of a structured response was captured, streaming was truncated, or the app renders something the transport did not carry. The consequence for reporting is one sentence and it belongs in the report: **a zero or low failure rate is evidence that this detector did not catch anything, not evidence that the model is safe.** Present the rate as a lower bound, name the detector, and pair the number with what it structurally cannot see. Reading a sample of passes is the only way to estimate the size of the gap.
go deeper
Says a detector can only spot what it was written to spot, so a low failure rate does not prove the model is safe.
Lists concrete miss routes - other languages, encodings, code, summarisation, off-distribution replies for a classifier - and frames the rate as a lower bound.
Adds capture failures as a silent cause, samples passes to estimate the miss rate, verifies once per target that the logged reply matches what the endpoint returns, and writes the bound into the report.
Refuses to let a clean run function as a release clearance anywhere in the organisation, and requires the harms outside the detector set's reach to be named next to any rate.
### Why a miss is structurally invisible A garak detector returns a score per output and the evaluator compares it to a threshold. There is no third value. A detector that has no idea what it is looking at returns a low score, and a low score is byte-for-byte indistinguishable in the report from a reply the detector confidently judged harmless. **There is no abstain, no confidence flag, no "off-distribution" marker.** A false positive eventually reaches a human because someone reads the hit log; a false negative produces silence, and silence is exactly the result a stakeholder was hoping for. That asymmetry is why misses are the more dangerous half of detector error even though they get a fraction of the attention. ### The four places misses come from **Representation.** A matcher is a claim about wording, and the model is under no obligation to use that wording. The same content as code, as a numbered list, in another language, inside a fictional frame, base-encoded, or spread across two turns scores 0.0. This is not an edge case; it is the ordinary behaviour of models asked the same thing twice. **Off-distribution classifier input.** A small classifier trained on one corpus - forum comments, say - meets a deployed product's replies: branded preamble, headed sections, mixed languages, a cut-off at a token cap. Its scores degrade quietly and land below the inherited 0.5 cut. It reports nothing unusual. **Capture.** Strictly this is the generator's fault, but it arrives as detector silence and it is the most common cause of a suspiciously clean first run. If the generator pulled the wrong field out of a structured response, or logged an error body, or truncated a stream, the detector scored text no user ever saw. Every attempt passes and the report looks immaculate. **Adjudication scope.** Some harms simply cannot be expressed as a score over reply text: a confidently-framed wrong answer, an unsafe action taken through a tool call, a policy breach that depends on who the user is, a harm that only appears over several turns. If nothing in your detector set can express the harm, the run cannot report it however many prompts you send. Probe coverage does not rescue detector coverage. ### What estimating the miss rate costs Sampling passes costs the same minutes per item as sampling hits - one to two minutes to read a prompt and reply against a written criterion - but there are far more passes and the events you are hunting are rarer, so the information per hour is much worse. Concretely: if the true miss rate is around 5%, hand-scoring 20 passes will most likely find nothing at all, and finding nothing in 20 items is consistent with a true rate up to roughly 15%. You need on the order of 60 items before "we found none" starts to mean something. Budget a few hours per probe family per cycle, and accept that the honest output is a bound rather than a corrected figure. ### Where the number misleads The single sentence that matters: **a low or zero failure rate is evidence that this detector set found nothing among these prompts, not evidence that the model is safe.** Two ways that gets abused. A clean run is quoted as a clearance, which reverses the burden of proof - the run supports only the narrow claim, and the narrow claim should be written out in full. And the calibration Z-score or percentile is treated as a correction for the blind spot: it is not, because it positions your run against other runs adjudicated by the same detectors, so a shared blind spot moves every run together and never appears in the comparison. ### What you would check Verify capture **once per target** before believing any result: send something whose reply you can predict, then confirm that the text garak logged is the text the endpoint actually returned - the right field, untruncated, not an error body. Then hand-score a random sample of *passing* attempts against the same written criterion you used for hits; the fraction you disagree with is your estimate of the miss rate for that family. Finally, write the bound into the report as a sentence, and list next to it the harms this detector set structurally cannot see, so the reader knows what silence covers.
- A garak run against a new deployed endpoint reports zero failures across every probe you selected. What do you suspect before you celebrate?A capture problem: the wrong reply field, an error body, or truncated streaming logged as the reply. Verify once that the logged text matches what the endpoint really returns.
- How would you estimate the miss rate for one probe family?Hand-score a random sample of the passing attempts against a written success criterion; the fraction you judge as real failures is your estimate of what the detector missed.
A detector that misses is a smoke alarm with a flat battery: the silence sounds exactly like safety, and only the person who presses the test button ever learns the difference.
saying these in an interview costs you the question
- Reads a zero failure rate as evidence the model is safe.
- Never inspects passing attempts.
- Does not consider that the logged reply may not be what the endpoint actually returned.
- Claims coverage of harms no detector in the set can express.