skip to content

Your garak sweep against a rate-limited production endpoint keeps dying partway through — throttling responses, timeouts, an expiring credential. How do you restructure the probe selection so that a half-finished sweep still yields results you can report?

level: seniorimportance: should knowfreq 45%

answer

  1. unit of work smaller than unit of failure
  2. batch by probe selection, one report each
  3. order by threat model, not catalogue
  4. freeze the target across batches
  5. errored call is not a clean result

basics

~20 s

Split the sweep into several small runs, each with a named probe selection and its own report, ordered so the families you must report on finish first. Keep the target configuration identical across runs so results stay comparable, log the selection per run, and treat a died-mid-probe result as incomplete, never as a pass.

solid answer

~60 s

One monolithic sweep is a single point of failure: when it dies you have a partial report whose gaps are indistinguishable from clean results. Partitioning fixes that. Run the selection in batches — a module or a few classes per run — each writing its own report, so every completed batch is a finished, quotable unit of work. Ordering matters as much as batching. Put the families the engagement actually promised first, so the highest-value results land before you hit the wall. Hold everything else constant across batches: same endpoint, same system prompt, same decoding settings, same repeat count, otherwise the batches cannot be aggregated into one statement. Separately, fix the throttling rather than absorbing it. Reduce request concurrency to what the endpoint tolerates, and make sure errored calls are visible in the results instead of collapsing into no-hit rows — an attempt that never reached the model is not evidence of a refusal. Refresh credentials before a long batch rather than mid-run, and re-run any batch that died rather than stitching its partial output into the total.

go deeper

for a junior

Should recognise that a crashed scan gives partial results and that you re-run rather than reporting what happened to finish.

for a middle

Should split the selection into batches with separate reports and know that errored calls must not be read as clean results.

for a senior

Designs the batching and ordering around the failure mode, freezes the target configuration for aggregation, tunes concurrency to the endpoint's limits, and verifies attempt counts per probe.

for a principal

Negotiates the scanning window, quota and a dedicated credential with the platform owner up front, and sets the rule that partial batches never enter the numbers.

### The principle **Make the unit of work smaller than the unit of failure.** A garak sweep that needs six hours against a throttled production endpoint *will* be interrupted — by rate limiting, by a socket timeout, by a credential with a shorter life than the run. Designing for that at hour zero costs you a few minutes of scripting; discovering it at hour five costs the whole sweep's token spend, twice. ### Partition by probe selection Batch the selection into runs of one module, or a handful of `module.ClassName` entries, each invoked as its own `garak --probes ...` command writing its own report. Each completed batch is then a durable, self-describing artefact: it names its probes, its generator configuration and its results. Ten small reports aggregate cleanly. One truncated report does not, because a probe missing from it could have been skipped, crashed, or simply not reached — and those are indistinguishable after the fact. The overhead is real and worth naming: more process startups, more report files to merge, and some engineer time stitching a summary. Against that, a died batch costs you only that batch's tokens to re-run instead of the whole sweep's — which on a metered endpoint is usually the difference between a rounding error and re-spending a day's budget. ### Order by value, not by catalogue order Put the families in the engagement's threat model first. If the run dies at 40%, you want that 40% to be the part someone actually asked for. Catalogue order is alphabetical accident; it has no relationship to what the deployment is exposed to. ### Freeze the target across batches Aggregating batches into one statement is only legitimate if every batch hit the same endpoint and model version, the same system prompt, the same decoding configuration and guard arrangement, at the same `--generations` value. Record those with each batch. If the deployment changed mid-engagement — a prompt tweak, a model version bump, a guard rolled out — the batches before and after are two results, not one, and merging them produces a number that describes no system that ever existed. ### Treat throttling as signal, not noise Throttled and timed-out calls are the attempts most likely to be silently miscounted. An HTTP 429 or a socket timeout means the prompt never reached the model; it is evidence of nothing about the model's behaviour, and it must never be aggregated with genuine non-compliant or compliant responses. So tune concurrency down until the error rate is near zero. A slower clean run beats a fast run full of holes, and against a rate-limited endpoint raising concurrency does not buy throughput anyway — it converts requests into rejections. Then verify: for each probe, compare the attempts recorded in the report against the expected prompts x generations figure. A shortfall means that probe is **incomplete**, whatever its result row says. ### Where the number misleads This is the specific misreading to guard against: **a truncated probe's failure rate looks better than the truth.** If a probe was supposed to send 500 attempts and 60 landed before the endpoint started rejecting, the denominator in its rate is 60 — and the rate is computed over exactly the attempts that got through, which are not a random sample of the intended ones. Throttling is bursty and correlates with the run's own load, so the surviving attempts skew toward the start of the probe's prompt list. Splicing that partial batch into the totals therefore does not merely add noise; it biases the aggregate in the reassuring direction, and it does so invisibly, because the report renders a partial rate exactly like a complete one. ### Re-run, do not patch A batch that died halfway gets re-run in full at lower concurrency, and its partial output is kept only as evidence that those specific attempts happened, never as a contribution to any rate. Refresh credentials *before* a long batch rather than mid-run. ### What you would keep The catalogue listing, the selection string and generator configuration per batch, the report per batch, and a short log of which batches were re-run and why. That set is what lets you answer, six months later during an incident review, exactly what was tested and what was not — which is the only question anyone will actually ask you.

  • Why not simply raise concurrency to finish the sweep before the credential expires?
    More concurrency against a rate-limited endpoint converts requests into throttled errors. You get a faster run with more holes, and the holes are the results most likely to be miscounted as clean.
  • A batch died at roughly 60%. Can you keep its partial results?
    Only as evidence that those specific attempts happened. It cannot contribute per-probe rates, because its denominator differs from the completed batches. Re-run the batch in full.
  • What single check tells you a probe's results are complete?
    Compare the attempt count recorded for that probe against the expected prompts-times-repeats figure. A shortfall means the probe is incomplete regardless of what its result row says.

It is the difference between writing a long document with autosave on and only saving at the very end: when the machine dies at hour five, one of you has finished chapters to hand in and the other has no idea which pages survived. Each completed garak batch is a saved chapter.

saying these in an interview costs you the question

  • Running the whole catalogue as one job against a rate-limited endpoint and hoping.
  • Merging a truncated batch's results into the per-probe rates.
  • Reading throttled or timed-out attempts as non-hits.
  • Changing the endpoint, system prompt or repeat count between batches and still aggregating them.
  • Cranking concurrency up to beat a credential expiry.

context