A garak report shows each per-probe rate with an uncertainty interval beside it, but for several probes that interval is blank. What does a blank interval tell you about those rows, and what would you change about the run before quoting one of those rates?
answer
- blank interval = below sample-size floor
- attempts = prompts x generations
- raise generations, then re-read
- quote raw counts, not a percentage
- attempts are the cost driver
basics
~20 sBlank means too few attempts. Below a sample-size floor the tool declines to print an interval rather than show a meaningless one, so that rate is an estimate you cannot bound. Re-run the probe with more generations per prompt, or a longer prompt set, until the interval appears, then quote it.
solid answer
~50 sThe interval is the report's statement of how firm the rate is, and it is only computed once enough attempts stand behind the row. Below that floor the tool leaves it empty — a signal, not a rendering bug. A rate with a blank interval is still a real observation (a flagged attempt is a flagged attempt), but it carries no bound: from a handful of attempts, 0% and 30% are both compatible with the same underlying behaviour. The fix is on the input side, not the report side: increase the generations per prompt, or run a probe variant with more prompts, so the attempt count crosses the floor, then read the interval. If you cannot — a metered endpoint, a slow local model, a probe with a genuinely short prompt list — say so explicitly in the write-up: quote the raw counts ("3 of 12 attempts") rather than a percentage, and do not put that row in a trend chart, because it will move on noise alone.
go deeper
Recognises the blank interval means too few data points and that the rate should not be quoted as a firm percentage.
Explains the sample-size floor, names generations per prompt and prompt-list length as the two levers, and reports raw counts when the interval will not appear.
Balances the interval against the per-call cost of attempts, targets extra attempts at the rows that will be quoted, and refuses to trend a blank-interval row.
Sets the reporting standard: which rows may appear as percentages, what evidence a trend series requires, and the attempt budget the team funds to get there.
### Why the field is empty The interval beside a per-probe rate is the report's statement of how firmly that rate is pinned down by the attempts behind it. Computed from a handful of attempts it is either absurdly wide or degenerate, so below a sample-size floor garak prints nothing rather than print something a reader would trust. A blank interval is therefore a signal, not a rendering fault, and the sentence it is speaking is "do not quote this row as a percentage". ### What sets the attempt count Two levers, and only two. The first is the probe you selected with `garak --probes`: some probe classes carry hundreds of generated prompts, others carry a dozen hand-written ones, and a short list caps the row no matter what else you do. The second is `garak --generations`, the number of completions drawn per prompt. attempts = prompts × generations, and the exact `total` for every row is written into the run's `report.jsonl` next to `passed`, so you never have to guess which of the two is small — you can read it. Generations do more than arithmetic against a non-deterministic target. Sampling at a non-zero temperature means the same prompt can be handled correctly nine times and badly once. At one generation per prompt you do not see that behaviour at all; you see one draw from a distribution and record it as if it were the distribution. Raising generations both narrows the interval and surfaces intermittent failures, and the second effect is usually the more valuable one. ### What it costs Every attempt is one call to the target, so this is a straight multiplier on money and wall clock. Taking a 60-prompt probe from 1 generation to 10 is 540 extra calls for that single row; doing the same across a hundred-probe sweep turns a run of a few thousand calls into tens of thousands, and a coffee break into most of a working day against a rate-limited endpoint. That is why "just raise generations" is not automatically the right answer. The disciplined move is to spend the extra attempts on the specific rows you intend to quote, trend, or take to a decision, and to leave exploratory probes at low generations and report them as raw counts. ### Where the number misleads The interval covers **sampling error only**. It is a statement about how much the rate would wobble if you drew the attempts again; it says nothing about whether the detector labelled them correctly. A row with 5,000 attempts and a tight interval whose detector systematically flags refusals as failures is precisely wrong with high confidence, and the tight interval makes it look more credible, not less. Sampling error and detector error are independent problems, and only one of them appears in that column. Two related misreadings. A wide but printed interval is not much better than a blank one; people relax as soon as any interval appears. And a blank interval is not a statement about the calibration comparison and not the detector expressing doubt — it is purely a count of attempts. Filling in the missing interval by rerunning with a different report prefix, or by re-reading the same data, does nothing at all. ### What to check, and how to write the row up Read `total` for the row out of the report, so you know whether the shortfall came from the prompt list or from the generations setting, and can fix the right one. If you can afford to, re-run that probe at a higher generations count and see whether the point estimate stays where it was; a rate that moves a long way is telling you the first number was noise. Spot-check a few flagged attempts by hand regardless, because that is the only check on the numerator. When the row stays blank, publish counts and not a percentage. "2 of 8 attempts flagged; too few attempts for the report to bound the rate" is defensible and self-limiting. "25% failure" from the same eight attempts is the sentence that ends up in a slide with no denominator attached and gets compared next quarter against a run of 800. And a blank-interval row does not belong in a trend series at all until its configuration changes, because quarter-to-quarter movement on eight attempts is noise wearing the costume of a signal.
- The endpoint is metered and you cannot afford more attempts everywhere. How do you decide where to spend them?Spend on the probes whose numbers will be quoted or trended, and on any row near a decision boundary. Leave exploratory probes at low generations and report them as counts. The choice belongs in the write-up so a reader knows which rows are estimated and which are indicative.
- Does raising generations per prompt do anything other than shrink the interval?Yes. Against a non-deterministic target it samples the response distribution, so an intermittent failure that one generation would miss entirely starts showing up. That is often the bigger gain.
A blank interval is a poll that stopped after eight respondents: the pollster reports the raw eight rather than a percentage, because the margin of error would be wider than the finding. Reporting 25% from those eight is not a smaller claim than the pollster's, it is a bigger one made from less.
saying these in an interview costs you the question
- Reading the empty field as a rendering bug and ignoring it.
- Quoting a percentage from a row whose interval never printed.
- Putting a blank-interval row into a quarter-over-quarter trend.
- Proposing to raise generations everywhere without acknowledging the per-call cost.
- Confusing the missing interval with a missing calibration comparison.