In a garak report one probe lands in a poor band on its absolute score while its calibration score sits mid-pack against the reference models, and another probe scores well absolutely but far below the reference models. You are writing the run up. Which number do you act on in each case, and why?
answer
- absolute = how bad, relative = how unusual
- bad + typical = category-wide hazard
- good + below pack = own outlier defect
- band says which verdict dominates
- no reference data = no relative score
basics
~20 sThey answer different questions. The absolute score says how bad the behaviour is in itself; the calibration score says how unusual it is versus reference models on that same probe. Poor absolute, mid-pack relative is a hazard the whole field shares. Good absolute, far below the pack is this model's own outlier defect.
solid answer
~60 s**Probe one — bad absolutely, typical relatively.** The reference models are about as exposed as yours. That is a real hazard and belongs in the report, but it is not a defect you fix by choosing a different model, because there is no evidence a different model would do better. It argues for a control outside the model: filtering, scoping, or not shipping that capability. **Probe two — good absolutely, poor relatively.** The absolute band looks reassuring, but peers demonstrably do better on the same prompts. That is the actionable finding: something specific to this target, its system prompt, its decoding settings or its fine-tune, is underperforming a bar that is known to be reachable. The band garak assigns tells you which verdict the tool itself treats as dominant for that row, and you should say in the write-up which of the two numbers you quoted. Also check whether a relative score exists at all: a probe that is not represented in the reference bundle has an absolute score and nothing to compare it against, and a blank there is not a good result.
go deeper
Can state that one number is a severity and the other a comparison, even if the case analysis is thin.
Reads both cases correctly and knows a mid-pack position does not mean safe.
Chooses the action per case, rules out configuration and sample-size explanations before escalating, and states in the write-up which number was quoted.
Decides which axis drives which kind of decision across a programme, and how a category-wide hazard is handled when no model choice fixes it.
### Two axes, computed differently The **absolute score** grades a probe's own rate: this model, these prompts, this detector, graded into a band that says how often the behaviour occurred. It is a severity-flavoured statement about your system alone. The **calibration score** takes that same rate and places it against a distribution of reference models that the maintainers measured on the same probe, using data bundled with the tool. Mid-pack means your rate sits near the reference mean; far below means your rate is worse than most of the reference population. That is a position in a population — it has no units of harm in it. Practically every mistake in this area is one of the two being read as the other. ### Case one: poor absolute band, mid-pack relative The reference models are about as exposed as yours. This is still a real hazard: the behaviour occurred, often, on your system, and it belongs in the report. What it is not is a defect you fix by picking a different model, because there is no evidence in front of you that any available model does better. So the response has to be a control that does not depend on the model improving — narrow the capability, put a filter in front of it, scope who can reach it, or accept and document the residual with an owner and a review date. The sentence you must not write is "in line with comparable models, no action required". The relative score was never a safety statement, and a non-technical reader hears it as one. If the whole field is bad at something, average is exposed. ### Case two: good absolute band, far below the pack This is the finding that gets dropped, because the absolute band looks comfortable and the eye stops there. A demonstrable gap to the reference population means the bar is reachable — other models on the same prompts clear it — and something about this specific deployment is not reaching it. That makes it the more actionable of the two rows. ### What it costs, and where the comparison misleads The calibration comparison itself costs nothing at run time: it is a lookup against data shipped with the tool, no extra calls to your target. The cost is entirely interpretive, and it is easy to underrate. The reference models were measured by the maintainers under **their** configuration — typically a bare model endpoint, no system prompt, no input filter, no retrieval context. Your target is very likely a system: a system prompt, a guard in front, decoding settings someone tuned. A flattering Z-score may simply mean you measured a different system than they did, and an unflattering one may mean the reverse. The comparison assumes a like-for-like harness that you did not necessarily run. Two further limits. The reference set is small and finite, so a position expressed as a standard-deviation offset has meaningless tails — "three sigma below the pack" out of a handful of models is arithmetic, not evidence. And the set ages: as it falls behind the current model generation, mid-pack drifts toward flattery. A blank in the relative column is a third case again — the probe is simply not represented in the bundled data, so you have an absolute number and no population to place it in. Say that, rather than letting the blank read as a pass. ### What to check before acting on either For a below-the-pack row, in order, because each is cheaper than the next: is the interval on your own rate printed and is the attempt count large enough that the gap is not sampling wobble; did the target run with a system prompt, guard or decoding setting the reference measurement did not have, and does the gap survive a re-run under the reference-like configuration; is the probe present in the bundle at all; and only then, is this a genuine property of the fine-tune or deployment worth escalating. ### What the write-up looks like One block per quoted probe: the raw counts, the absolute band, the relative position if one exists, the interval, and one explicit sentence saying which of the two numbers you are asking the reader to act on and why. The failure mode this discipline prevents is a report that quotes whichever axis reads better and never names which axis it was.
- A stakeholder says 'we are average on this probe, so we are fine.' What is your reply?Average is a position in a population, not a level of risk. If the population is uniformly weak on that behaviour, average means exposed. The absolute band is the number that speaks to risk, and it is the one to quote for that conversation.
- You find a probe well below the reference models. What do you check before calling it a defect?Attempt count and whether the interval printed, so the gap is not noise; whether the target ran with an unusual system prompt or decoding configuration; and whether the reference data covers the probe at all. Only then does it become a finding worth escalating.
The absolute score is the mark on the exam; the calibration score is the percentile among the other candidates. Scoring 40 out of 100 while sitting exactly on the class median tells you the paper was hard, not that 40 is a passing grade.
saying these in an interview costs you the question
- Writing 'in line with comparable models' as if the relative score were a safety verdict.
- Dismissing a below-the-pack probe because its absolute band looked acceptable.
- Quoting 'the score' without saying whether it was the absolute or the calibration number.
- Reading a blank relative score as a pass.
- Escalating a relative gap without first checking attempt count, interval and target configuration.