skip to content

In a garak scan, what does the generations setting (how many times each probe prompt is re-sent to the target generator) change about the run, and what does raising it cost?

level: juniorimportance: must knowfreq 72%

answer

  1. repeat count = attempts multiplier
  2. prompts x N = queries
  3. detectors score each attempt
  4. no new coverage, only depth
  5. default is not 1, read the config

basics

~20 s

Generations is garak's repeat count: how many times each probe prompt is sent to the target generator. Each reply is detected separately, so ten repeats give ten scored attempts per prompt instead of one. Cost scales with it, roughly ten times the queries, tokens and wall clock. The prompt set itself is unchanged.

solid answer

~50 s

It is a multiplier on **attempts**, not on prompts. garak takes each probe's prompt set and sends every prompt to the generator N times; each returned response is passed to that probe's detectors independently, so one prompt can produce a mix of clean and failing attempts. Everything downstream scales with N: queries against the endpoint, tokens billed, wall-clock time, rate-limit pressure, and the number of hits triage has to read. What does *not* scale is coverage of distinct prompts — you learn nothing about a family of attacks you never selected, no matter how many times you repeat the ones you did. The practical consequence is that N is the knob for **depth of sampling**, and it is the wrong knob if what you actually lack is breadth. Do not assume the default: read the repeat count out of the run's own configuration, because it is not 1 and it has not been the same value across releases.

go deeper

for a junior

Knows it re-sends each prompt N times and that the bill and runtime scale with N.

for a middle

Adds that each generation is scored independently by the detectors, so one prompt can yield both passing and failing attempts, and that coverage is unaffected.

for a senior

Talks about rate limits, retry accounting, triage load scaling with N, and reading the repeat count out of the run's recorded configuration rather than assuming a default.

for a principal

Frames the repeat count as a per-tier budget policy and insists the count travels with every rate the scan reports.

### What the setting is, and where it lives garak's unit of work is the **attempt**: one prompt, sent once, scored once. Three objects meet at that unit. A **probe** — a Python plugin class, the kind `garak --list_probes` prints — owns a fixed set of prompts. A **generator** — selected with `garak --model_type` and `--model_name` — owns the actual call to the system under test. A **detector** owns the verdict on the text that comes back. The repeat count sits between the probe and the generator: it is `garak --generations N` on the command line (short form `garak -g N`), or `run.generations` in the YAML you hand to `garak --config`. It decides how many times each of the probe's prompts is sent, and therefore how many attempts each prompt produces. Nothing else about the run changes. The prompt list belongs to the probe class; which probes run is fixed by `garak --probes`. Raising the repeat count re-sends what you already selected, and each returned response is passed to that probe's detectors independently — so a single prompt can yield a mix of clean and failing attempts. ### Why the knob exists at all The thing under test is a sampler. A hosted chat endpoint at ordinary decoding settings does not return the same text twice, and a guardrail sitting in front of it may fire on one draw and not the next. One generation per prompt is one Bernoulli draw: it reliably detects behaviour that fails every time, and is close to blind to behaviour that fails occasionally. The repeat count is the only knob that buys resolution on the second kind. ### What it costs Everything downstream of the probe is linear in N. | what scales with N | how | | --- | --- | | model calls | prompts x N, plus retries | | tokens billed | the same multiple, on both prompt and completion | | wall clock | x N, unless you also raise `garak --parallel_attempts`, which against a rate-limited endpoint mostly buys backoff | | detector work | one verdict per attempt | | triage volume | roughly x N on any probe that fires — false positives included | Concretely: a selection totalling 2,000 prompts at `-g 1` is 2,000 calls; at `-g 10` it is 20,000. At a rough 400 prompt and 200 completion tokens per call that is about 12M tokens instead of 1.2M — the difference between pocket change and a visible line item on a metered endpoint — and at a sustained three requests per second it is roughly two hours of wall clock instead of eleven minutes. None of that spend buys a single new prompt. ### Where the number misleads Two readings go wrong, and both are common. First, **absolute hit counts are not severity**. The digest and the `eval` records in garak's report carry counts as well as rates, and the counts scale with N. A probe showing 40 hits at `-g 10` is not four times worse than one showing 10 hits at `-g 10`, and it is not worse at all than a `-g 1` run showing 4 — one flaky prompt repeated ten times manufactures ten hits. Read the rate, and read the repeat count that produced it, before you rank anything. Second, **more repeats is not more coverage**. Depth and breadth are separate budgets drawn on the same wallet. A team that spends its whole query allowance raising N on the probes it already ran knows more about those prompts and exactly nothing more about the families it never selected. If the honest gap is "we have never run that probe family at all", the repeat count is the wrong knob and buying it is a way of feeling thorough without becoming thorough. ### What to check before you believe a run 1. **The count actually used.** Read `run.generations` out of the run's own recorded configuration or report header, never from memory: garak's default has not been the same value across releases, so a remembered default is a guess dressed as a fact. 2. **The generator's decoding settings.** If the target is configured to be effectively deterministic, high N returns near-duplicates at full price and the run is not sampling the variance you paid for. 3. **That attempt totals reconcile.** For a probe with P prompts the scored total should be P x N. If it is short, attempts errored, timed out or were dropped by the transport — and a dropped attempt is missing data, not a pass.

  • You raise the repeat count and the scan now takes far longer than N times as long. What is the likely cause?
    Rate limiting and retry backoff on the target endpoint. More concurrent attempts push you into 429s, and the backoff time is not proportional to the extra work.
  • The generator is configured for deterministic decoding. Is a high repeat count still worth paying for?
    Mostly no. You will get near-identical outputs at full price. Either scan at the sampling settings production actually serves, or accept a low repeat count and say so.
  • Does raising the repeat count improve the scan's coverage?
    No. Coverage is which probes and prompts you ran. Repeats deepen sampling on the prompts already selected and add nothing about the ones you skipped.

saying these in an interview costs you the question

  • Thinking a higher repeat count runs more probes or covers more attack families.
  • Assuming the repeat count is 1 by default, or quoting a default from memory instead of reading the run configuration.
  • Not connecting the repeat count to the endpoint bill and rate limits.
  • Believing repeated generations are deduplicated or cached, so extra repeats are free.

context