skip to content

Running a Scan

The knobs that set a run's bill and its blind spots: which probes go out, how often each prompt is re-sent, and what the default leaves untouched. Interviewers ask what your scan actually cost.

on this pageshow

explore

questions

10

In a garak scan, what does the generations setting (how many times each probe prompt is re-sent to the target generator) change about the run, and what does raising it cost?

level: juniorimportance: must knowfreq 72%

answer

  1. repeat count = attempts multiplier
  2. prompts x N = queries
  3. detectors score each attempt
  4. no new coverage, only depth
  5. default is not 1, read the config

basics

~20 s

Generations is garak's repeat count: how many times each probe prompt is sent to the target generator. Each reply is detected separately, so ten repeats give ten scored attempts per prompt instead of one. Cost scales with it, roughly ten times the queries, tokens and wall clock. The prompt set itself is unchanged.

solid answer

~50 s

It is a multiplier on **attempts**, not on prompts. garak takes each probe's prompt set and sends every prompt to the generator N times; each returned response is passed to that probe's detectors independently, so one prompt can produce a mix of clean and failing attempts. Everything downstream scales with N: queries against the endpoint, tokens billed, wall-clock time, rate-limit pressure, and the number of hits triage has to read. What does *not* scale is coverage of distinct prompts — you learn nothing about a family of attacks you never selected, no matter how many times you repeat the ones you did. The practical consequence is that N is the knob for **depth of sampling**, and it is the wrong knob if what you actually lack is breadth. Do not assume the default: read the repeat count out of the run's own configuration, because it is not 1 and it has not been the same value across releases.

go deeper

for a junior

Knows it re-sends each prompt N times and that the bill and runtime scale with N.

for a middle

Adds that each generation is scored independently by the detectors, so one prompt can yield both passing and failing attempts, and that coverage is unaffected.

for a senior

Talks about rate limits, retry accounting, triage load scaling with N, and reading the repeat count out of the run's recorded configuration rather than assuming a default.

for a principal

Frames the repeat count as a per-tier budget policy and insists the count travels with every rate the scan reports.

### What the setting is, and where it lives garak's unit of work is the **attempt**: one prompt, sent once, scored once. Three objects meet at that unit. A **probe** — a Python plugin class, the kind `garak --list_probes` prints — owns a fixed set of prompts. A **generator** — selected with `garak --model_type` and `--model_name` — owns the actual call to the system under test. A **detector** owns the verdict on the text that comes back. The repeat count sits between the probe and the generator: it is `garak --generations N` on the command line (short form `garak -g N`), or `run.generations` in the YAML you hand to `garak --config`. It decides how many times each of the probe's prompts is sent, and therefore how many attempts each prompt produces. Nothing else about the run changes. The prompt list belongs to the probe class; which probes run is fixed by `garak --probes`. Raising the repeat count re-sends what you already selected, and each returned response is passed to that probe's detectors independently — so a single prompt can yield a mix of clean and failing attempts. ### Why the knob exists at all The thing under test is a sampler. A hosted chat endpoint at ordinary decoding settings does not return the same text twice, and a guardrail sitting in front of it may fire on one draw and not the next. One generation per prompt is one Bernoulli draw: it reliably detects behaviour that fails every time, and is close to blind to behaviour that fails occasionally. The repeat count is the only knob that buys resolution on the second kind. ### What it costs Everything downstream of the probe is linear in N. | what scales with N | how | | --- | --- | | model calls | prompts x N, plus retries | | tokens billed | the same multiple, on both prompt and completion | | wall clock | x N, unless you also raise `garak --parallel_attempts`, which against a rate-limited endpoint mostly buys backoff | | detector work | one verdict per attempt | | triage volume | roughly x N on any probe that fires — false positives included | Concretely: a selection totalling 2,000 prompts at `-g 1` is 2,000 calls; at `-g 10` it is 20,000. At a rough 400 prompt and 200 completion tokens per call that is about 12M tokens instead of 1.2M — the difference between pocket change and a visible line item on a metered endpoint — and at a sustained three requests per second it is roughly two hours of wall clock instead of eleven minutes. None of that spend buys a single new prompt. ### Where the number misleads Two readings go wrong, and both are common. First, **absolute hit counts are not severity**. The digest and the `eval` records in garak's report carry counts as well as rates, and the counts scale with N. A probe showing 40 hits at `-g 10` is not four times worse than one showing 10 hits at `-g 10`, and it is not worse at all than a `-g 1` run showing 4 — one flaky prompt repeated ten times manufactures ten hits. Read the rate, and read the repeat count that produced it, before you rank anything. Second, **more repeats is not more coverage**. Depth and breadth are separate budgets drawn on the same wallet. A team that spends its whole query allowance raising N on the probes it already ran knows more about those prompts and exactly nothing more about the families it never selected. If the honest gap is "we have never run that probe family at all", the repeat count is the wrong knob and buying it is a way of feeling thorough without becoming thorough. ### What to check before you believe a run 1. **The count actually used.** Read `run.generations` out of the run's own recorded configuration or report header, never from memory: garak's default has not been the same value across releases, so a remembered default is a guess dressed as a fact. 2. **The generator's decoding settings.** If the target is configured to be effectively deterministic, high N returns near-duplicates at full price and the run is not sampling the variance you paid for. 3. **That attempt totals reconcile.** For a probe with P prompts the scored total should be P x N. If it is short, attempts errored, timed out or were dropped by the transport — and a dropped attempt is missing data, not a pass.

  • You raise the repeat count and the scan now takes far longer than N times as long. What is the likely cause?
    Rate limiting and retry backoff on the target endpoint. More concurrent attempts push you into 429s, and the backoff time is not proportional to the extra work.
  • The generator is configured for deterministic decoding. Is a high repeat count still worth paying for?
    Mostly no. You will get near-identical outputs at full price. Either scan at the sampling settings production actually serves, or accept a low repeat count and say so.
  • Does raising the repeat count improve the scan's coverage?
    No. Coverage is which probes and prompts you ran. Repeats deepen sampling on the prompts already selected and add nothing about the ones you skipped.

saying these in an interview costs you the question

  • Thinking a higher repeat count runs more probes or covers more attack families.
  • Assuming the repeat count is 1 by default, or quoting a default from memory instead of reading the run configuration.
  • Not connecting the repeat count to the endpoint bill and rate limits.
  • Believing repeated generations are deduplicated or cached, so extra repeats are free.

context

open as a page

When you start a garak scan against a chat endpoint and do not name any probes, what set of attacks actually runs, and why is that neither nothing nor the whole probe catalogue?

level: juniorimportance: must knowfreq 70%

basics

~20 s

garak runs a built-in default selection of its probes, not the entire catalogue. The full catalogue is far larger and would cost hours of metered calls. Some probes also only run when you name them. So a default run is a sample, and you must read the report to see which probes actually ran.

open as a page

A garak scan configured with one generation per prompt reports a probe as passing; the identical scan rerun the next day fails it. What about the tool's sampling explains the flip, and how do you pick a repeat count that makes it unlikely?

level: middleimportance: must knowfreq 64%

basics

~20 s

The target samples, so a prompt that fails only sometimes is a coin flip. With one repeat you draw once: a behaviour firing on one attempt in ten is missed nine times out of ten and filed as a pass. Pick the repeat count from the rarest failure you still want to catch.

open as a page

You have been asked to scan a metered hosted chat endpoint with garak and to predict the spend before you launch. What multiplies out to the number of requests it will send, and which part of that product does your probe selection control?

level: middleimportance: must knowfreq 60%

basics

~20 s

Requests are roughly the sum, over each probe you selected, of that probe's prompt count times the repeat count per prompt. Your selection controls which probes and therefore the prompt counts, which differ by orders of magnitude between probes. Anything model-backed in the pipeline, such as a hosted detector, adds calls on top.

open as a page

A garak run over a hand-picked subset of probes finishes with nothing flagged. How do you write that up so it is not read as "the model is safe", and what denominator do you attach when you use the word coverage?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Report it as: these named probes, at this repeat count, against this endpoint and configuration, on this date, produced no detector hits. Coverage means probes run divided by probes in the catalogue — not attack surface and not behaviours. Name the families you did not run as untested, never as passed.

open as a page

When a garak scan repeats each prompt several times, what does the per-probe pass/failure rate in its report count over, and how do a probe that fails on every attempt and one that fails occasionally look different in that report?

level: middleimportance: should knowfreq 48%

basics

~20 s

Attempts, not prompts. With the repeat count at N, each prompt yields N scored attempts and the probe's rate is over that larger total. A probe failing every time sits at the extreme of the range; one failing occasionally shows a small non-zero rate that a single-repeat run would have shown as zero.

open as a page

You are running garak against a metered hosted endpoint and the repeat count you want per prompt would blow the query budget. How do you get a usable number anyway, and which shortcuts would make the resulting rates worthless?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Spend repeats where they decide something: a high count on the few probes gating the release, a low one elsewhere, and pool several cheap runs rather than one big one. Never buy repeats by pinning the generator's decoding to be deterministic when production samples, that measures a system you do not ship.

open as a page

Your garak sweep against a rate-limited production endpoint keeps dying partway through — throttling responses, timeouts, an expiring credential. How do you restructure the probe selection so that a half-finished sweep still yields results you can report?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Split the sweep into several small runs, each with a named probe selection and its own report, ordered so the families you must report on finish first. Keep the target configuration identical across runs so results stay comparable, log the selection per run, and treat a died-mid-probe result as incomplete, never as a pass.

open as a page

Your organisation wants one standing policy for how many times garak re-sends each prompt in scans that gate a release. How would you set it, and what must you tell every reader of a rate produced under that policy?

level: principalimportance: should knowfreq 31%

basics

~20 s

Set it per tier, not once: a small count for fast pre-merge scans, a much larger one for the release gate, and freeze it so rates stay comparable between runs. Publish the count next to every rate, and say plainly that a clean scan at a low count is weak evidence, not proof of absence.

open as a page

You have a fixed query budget for a garak engagement against a paid endpoint — enough for perhaps a fifth of the probe catalogue. How do you split it between breadth across many probe families and depth within a few, and what do you tell the stakeholder the result covers?

level: principalimportance: should knowfreq 35%

basics

~20 s

Spend a first slice broadly and shallowly across families to find where signal exists, then reinvest the rest depth-first on the families that hit and the ones the deployment's threat model cares about. Tell the stakeholder exactly which families ran, which were never run, and that unrun means untested.

open as a page