skip to content

garak

You will learn garak's probe/detector/generator architecture and how to scan a model for known failure modes and read its pass/fail report. Interviewers treat it as the nmap of LLMs — the quick-scan tool candidates are expected to have run.

on this pageshow

explore

questions

page 2 of 2

You rerun a garak probe that builds its prompts during the run against the same endpoint after shipping a mitigation, and the failure count drops. How do you establish whether the mitigation worked rather than the probe simply having sent different prompts?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A probe that mints prompts each run does not send the same prompts twice, so one lower count is a different sample, not a result. Replay the specific saved prompts that failed before against the patched system, and rerun the generating probe several times on both builds before claiming an improvement.

open as a page

A vendor hands you a garak run in which their chat product passed every probe whose prompts come from a fixed, publicly readable list. What do you check before accepting that as evidence of robustness, and what would you ask them to run instead?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Ask which probes ran and whether their prompts are a public fixed list. A public list can be filtered or tuned against, so a pass shows those strings are handled, not that the behaviour is absent. Ask for a rerun that includes probes minting prompts at run time, with the probe selection and logs attached.

open as a page

You are running garak against a metered hosted endpoint and the repeat count you want per prompt would blow the query budget. How do you get a usable number anyway, and which shortcuts would make the resulting rates worthless?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Spend repeats where they decide something: a high count on the few probes gating the release, a low one elsewhere, and pool several cheap runs rather than one big one. Never buy repeats by pinning the generator's decoding to be deterministic when production samples, that measures a system you do not ship.

open as a page

Your garak sweep against a rate-limited production endpoint keeps dying partway through — throttling responses, timeouts, an expiring credential. How do you restructure the probe selection so that a half-finished sweep still yields results you can report?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Split the sweep into several small runs, each with a named probe selection and its own report, ordered so the families you must report on finish first. Keep the target configuration identical across runs so results stay comparable, log the selection per run, and treat a died-mid-probe result as incomplete, never as a pass.

open as a page

Your team wants a risk covered that the garak LLM scanner ships no probe for. When would you decline to write a custom probe and detector pair, and what would you do instead?

level: principalimportance: should knowfreq 32%

basics

~20 s

Decline when nobody will own the ruling criterion over time. A custom pair needs re-labelling as the target changes, and a drifting criterion quietly moves the number while looking stable. For a one-off question, test by hand or with a small held-out set; reserve custom plugins for risks you will re-run every release.

open as a page

Your team runs garak against every model release and publishes the failure rates to other teams. The detectors' error rates ship with the tool and you never measured them. How do you set policy for how those numbers may be used?

level: principalimportance: should knowfreq 38%

basics

~20 s

Treat the rates as a trend within one pinned probe and detector set, never as an absolute safety figure. Pin that configuration across releases, hand-audit a sample of failures and passes each cycle to estimate detector error, publish the audited numbers with sample sizes, and never let a rate alone gate a release.

open as a page

Before a garak engagement on a customer-facing assistant, you must choose where the generator points: the vendor model endpoint behind the app, a staging deployment of the app, or live production. How do you argue the choice, and what does each option's number actually mean?

level: principalimportance: should knowfreq 35%

basics

~20 s

Name what each choice measures. The bare model measures the model; a staging deployment measures the app's prompts, retrieval and filters; production measures the live system plus its abuse controls, at the cost of polluted logs, alerts and real spend. Default to a staging build that mirrors production, and state the choice in the report.

open as a page

A single garak sweep leaves you several hundred hits and a day to triage them. How do you turn hits into report findings without over- or under-counting?

level: principalimportance: should knowfreq 47%

basics

~20 s

Group hits by the behaviour they demonstrate rather than writing one finding per hit. Read several per group, confirm the detector was right, and keep one finding with a verified example and the count behind it. Drop groups where every sampled hit was an artefact, and state what you sampled.

open as a page

Your team publishes garak's per-probe calibration scores in a quarterly report. Between two quarters the tested model was not changed, yet several of those relative scores moved. What could explain the movement, and how would you make a quarter-over-quarter comparison of these numbers trustworthy?

level: principalimportance: should knowfreq 25%

basics

~20 s

A relative score is a position against reference data that ships with the tool, so it moves when that reference data or the probe's prompts change, even with your model untouched. Sampling noise and a silently updated hosted endpoint also move it. Pin the tool and reference bundle per series, and trend absolute scores.

open as a page

Your organisation wants one standing policy for how many times garak re-sends each prompt in scans that gate a release. How would you set it, and what must you tell every reader of a rate produced under that policy?

level: principalimportance: should knowfreq 31%

basics

~20 s

Set it per tier, not once: a small count for fast pre-merge scans, a much larger one for the release gate, and freeze it so rates stay comparable between runs. Publish the count next to every rate, and say plainly that a clean scan at a low count is weak evidence, not proof of absence.

open as a page

You have a fixed query budget for a garak engagement against a paid endpoint — enough for perhaps a fifth of the probe catalogue. How do you split it between breadth across many probe families and depth within a few, and what do you tell the stakeholder the result covers?

level: principalimportance: should knowfreq 35%

basics

~20 s

Spend a first slice broadly and shallowly across families to find where signal exists, then reinvest the rest depth-first on the families that hit and the ones the deployment's threat model cares about. Tell the stakeholder exactly which families ran, which were never run, and that unrun means untested.

open as a page

Your team has a fixed number of calls it may spend on garak scans of a product endpoint each release. How do you decide whether any of that budget should go to buffed runs — where a buff rewrites probe prompts before sending — instead of more probes?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Spend the budget on breadth first: buffs add no new behaviour, they re-send prompts you already have. Buy buffed runs for a narrow slice where input normalisation or a filter is the thing under test, or to audit whether your own detectors are brittle. Treat them as an occasional audit, not a gate.

open as a page

You own a recurring garak scan attached to a release pipeline. How would you split the probe selection between probes carrying fixed, checked-in prompt lists and probes that generate their prompts each run, and what do you give up with each choice?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Gate the pipeline on a small pinned set of fixed-prompt-list probes, because only deterministic prompts give a comparable pass/fail signal per build. Run prompt-generating probes on a slower cadence for fresh coverage, reviewed by a human. Pinned probes go stale and invite tuning; generated ones produce noisy counts that cannot block a build.

open as a page

showing 31–43 of 43