skip to content

A garak report can position a probe's result against previously measured reference models as well as giving an absolute score. Why does a probe and detector you wrote yourself not get that comparative reading, and what do you present instead?

level: seniorimportance: should knowfreq 40%

answer

  1. position needs a fixed prompt+criterion across models
  2. day-one plugin has no prior runs
  3. build your own spread
  4. before/after mitigation baseline
  5. never same column as shipped families

basics

~20 s

The comparative reading needs prior results for that exact probe and detector across reference models. Your plugin has never been run against them, so there is nothing to compare with. Present the absolute rate, your own baseline runs against two or three comparison targets, and the detector's measured error.

solid answer

~50 s

A comparison against reference models is only meaningful when the same prompts were ruled by the same criterion on every model in the set. A plugin you just wrote satisfies neither condition, so the tool has no basis for saying whether your result is unusually bad or entirely typical. That absence matters. An absolute rate alone is nearly uninterpretable - is 22% high, compared with what? For custom families you have to manufacture the comparison yourself. The practical substitute is a small hand-built baseline: run the same custom probe and detector against two or three other targets you may legitimately test - an open-weights model you host, a second vendor endpoint, your own pre-mitigation build. Report it as 'X% here against Y% and Z% on comparison targets, detector validated at this error rate', and never put a custom family's absolute score in the same column as shipped families' comparative readings.

go deeper

for a junior

Knows a custom plugin has no prior results to be compared against, so only an absolute number comes out.

for a middle

Explains that a comparison needs the same prompts and same ruling criterion across every model in the reference set.

for a senior

Builds a hand-made spread or a before/after baseline, and keeps custom numbers out of the same column as shipped comparative readings.

for a principal

Decides whether the team invests in maintaining its own reference spread at all, and what the reporting standard is for numbers that have none.

### What the comparative reading is, mechanically Alongside an absolute score, a garak report can position a probe's result against a bag of previously measured reference models: roughly, how far your target's score sits from the mean of that bag, in standard deviations. That is a *positional* statement, and a positional statement is only meaningful if everything except the target was held fixed when the reference data was collected — the same prompt set, the same ruling criterion, comparable run conditions. The calibration data ships with the tool and is keyed by plugin identity: this probe, scored by this detector. A probe and detector you wrote yesterday have no key in that data. Not because custom plugins are second-class, but because the reference models were never run against your prompts and nobody ever scored their replies with your criterion. There is literally nothing on the other side of the comparison. So the report gives you the absolute number and stops. ### Why the absence bites An absolute rate is close to uninterpretable on its own. Is 22% high? Compared with what — a model that scores 5% on this family, or one that scores 60%? Without position, the reader supplies their own reference silently, and it is usually wrong in whichever direction their prior runs. For shipped families the tool supplies the reference; for custom families you must manufacture it, or explicitly refuse to imply one. ### The three honest closures 1. **Build your own reference spread.** Run the identical custom probe and identical detector against two or three other targets you are permitted to test: an open-weights model you host yourself, a second vendor endpoint, your own build from before a mitigation. This is the closest analogue to what the tool does and my default. Cost it honestly: it is the full prompts x `--generations` bill per target, and the slowest endpoint's rate limit sets the schedule for the whole spread, so three targets is not three times the wall clock — it is often much worse, because one of them throttles. 2. **Use a within-target baseline.** The same probe and detector against your own system with a mitigation off and on, or before and after a release. This answers "did we improve", which is usually the decision actually on the table, and it costs two runs against an endpoint you already control. 3. **Report the absolute number as absolute.** Criterion written out, error measured, no implied position. Boring, and honest. ### Where the number misleads Three specific misreadings, in rising order of damage. The first is **column contamination**: a report that lists shipped families with their positional readings and the custom family beside them, formatted identically. A reader scans down and treats them as commensurable. One is positioned against measured reference data; the other is an unvalidated criterion's output against one target. Different footnotes at minimum, different sections preferably. The second is **consistency mistaken for correctness**. Building your own spread makes the criterion *consistently applied*, not right. If the detector is wrong, it is wrong in the same direction on all three targets, and the spread comes out looking tidy and internally coherent — which is exactly what a wrong-but-stable ruler produces. Validating the detector and building the spread are two separate obligations; discharging one does not discharge the other. The third is **stale reference data**, which applies to shipped families too and is worth voicing because it is the same class of error. A bag of reference models measured some releases ago flatters anything current: your target looks unusually good against a comparison set that has since been superseded. A positional reading needs a date on it, the way a benchmark score does. ### What I would check Confirm from the report whether a positional reading was produced at all for the pair, rather than assuming a blank means a good result. If I am building a spread, hold prompts, detector and `--generations` identical across every target, and record model names, versions and dates alongside the numbers — a spread whose targets were run with different generations settings has different denominators and is not a spread. Confirm I am permitted to test each endpoint before sending anything. And keep the custom family's numbers in their own section, with the detector's measured error stated next to the rate.

  • What is the cheapest useful baseline when you cannot test other vendors' endpoints?
    A within-target before/after: the same custom probe and detector run against your build with the mitigation off and on. It answers whether you improved, which is usually the decision at hand.
  • Does running your custom plugin against five targets make its numbers trustworthy?
    No. It makes them consistently produced. A wrong criterion is wrong on all five in the same direction, so detector validation is still a separate obligation.

A positional score is a percentile in a class. Write your own exam and no class has ever sat it — you can still report the mark, but 'above average' has no average behind it yet.

saying these in an interview costs you the question

  • Treats an absolute rate as if it carried comparative meaning
  • Puts custom-family numbers beside shipped comparative readings without a caveat
  • Thinks running against several targets fixes a wrong detector
  • Cannot say what a comparison requires to be valid

context