garak
You will learn garak's probe/detector/generator architecture and how to scan a model for known failure modes and read its pass/fail report. Interviewers treat it as the nmap of LLMs — the quick-scan tool candidates are expected to have run.
on this pageshowhide
explore
- Scanner Plugins25 questions
- Probes5 questions
- Detectors5 questions
- Generators5 questions
- Buffs5 questions
- Custom Probes and Detectors5 questions
- Running a Scan10 questions
- Probe Selection5 questions
- Repeat Sampling5 questions
- Reading the Report8 questions
- Absolute and Relative Scores4 questions
- Triaging Hits4 questions
questions
page 2 of 2You rerun a garak probe that builds its prompts during the run against the same endpoint after shipping a mitigation, and the failure count drops. How do you establish whether the mitigation worked rather than the probe simply having sent different prompts?
basics
~20 sA probe that mints prompts each run does not send the same prompts twice, so one lower count is a different sample, not a result. Replay the specific saved prompts that failed before against the patched system, and rerun the generating probe several times on both builds before claiming an improvement.
A vendor hands you a garak run in which their chat product passed every probe whose prompts come from a fixed, publicly readable list. What do you check before accepting that as evidence of robustness, and what would you ask them to run instead?
basics
~20 sAsk which probes ran and whether their prompts are a public fixed list. A public list can be filtered or tuned against, so a pass shows those strings are handled, not that the behaviour is absent. Ask for a rerun that includes probes minting prompts at run time, with the probe selection and logs attached.
You are running garak against a metered hosted endpoint and the repeat count you want per prompt would blow the query budget. How do you get a usable number anyway, and which shortcuts would make the resulting rates worthless?
basics
~20 sSpend repeats where they decide something: a high count on the few probes gating the release, a low one elsewhere, and pool several cheap runs rather than one big one. Never buy repeats by pinning the generator's decoding to be deterministic when production samples, that measures a system you do not ship.
Your garak sweep against a rate-limited production endpoint keeps dying partway through — throttling responses, timeouts, an expiring credential. How do you restructure the probe selection so that a half-finished sweep still yields results you can report?
basics
~20 sSplit the sweep into several small runs, each with a named probe selection and its own report, ordered so the families you must report on finish first. Keep the target configuration identical across runs so results stay comparable, log the selection per run, and treat a died-mid-probe result as incomplete, never as a pass.
Your team wants a risk covered that the garak LLM scanner ships no probe for. When would you decline to write a custom probe and detector pair, and what would you do instead?
basics
~20 sDecline when nobody will own the ruling criterion over time. A custom pair needs re-labelling as the target changes, and a drifting criterion quietly moves the number while looking stable. For a one-off question, test by hand or with a small held-out set; reserve custom plugins for risks you will re-run every release.
Your team runs garak against every model release and publishes the failure rates to other teams. The detectors' error rates ship with the tool and you never measured them. How do you set policy for how those numbers may be used?
basics
~20 sTreat the rates as a trend within one pinned probe and detector set, never as an absolute safety figure. Pin that configuration across releases, hand-audit a sample of failures and passes each cycle to estimate detector error, publish the audited numbers with sample sizes, and never let a rate alone gate a release.
Before a garak engagement on a customer-facing assistant, you must choose where the generator points: the vendor model endpoint behind the app, a staging deployment of the app, or live production. How do you argue the choice, and what does each option's number actually mean?
basics
~20 sName what each choice measures. The bare model measures the model; a staging deployment measures the app's prompts, retrieval and filters; production measures the live system plus its abuse controls, at the cost of polluted logs, alerts and real spend. Default to a staging build that mirrors production, and state the choice in the report.
A single garak sweep leaves you several hundred hits and a day to triage them. How do you turn hits into report findings without over- or under-counting?
basics
~20 sGroup hits by the behaviour they demonstrate rather than writing one finding per hit. Read several per group, confirm the detector was right, and keep one finding with a verified example and the count behind it. Drop groups where every sampled hit was an artefact, and state what you sampled.
Your team publishes garak's per-probe calibration scores in a quarterly report. Between two quarters the tested model was not changed, yet several of those relative scores moved. What could explain the movement, and how would you make a quarter-over-quarter comparison of these numbers trustworthy?
basics
~20 sA relative score is a position against reference data that ships with the tool, so it moves when that reference data or the probe's prompts change, even with your model untouched. Sampling noise and a silently updated hosted endpoint also move it. Pin the tool and reference bundle per series, and trend absolute scores.
Your organisation wants one standing policy for how many times garak re-sends each prompt in scans that gate a release. How would you set it, and what must you tell every reader of a rate produced under that policy?
basics
~20 sSet it per tier, not once: a small count for fast pre-merge scans, a much larger one for the release gate, and freeze it so rates stay comparable between runs. Publish the count next to every rate, and say plainly that a clean scan at a low count is weak evidence, not proof of absence.
You have a fixed query budget for a garak engagement against a paid endpoint — enough for perhaps a fifth of the probe catalogue. How do you split it between breadth across many probe families and depth within a few, and what do you tell the stakeholder the result covers?
basics
~20 sSpend a first slice broadly and shallowly across families to find where signal exists, then reinvest the rest depth-first on the families that hit and the ones the deployment's threat model cares about. Tell the stakeholder exactly which families ran, which were never run, and that unrun means untested.
Your team has a fixed number of calls it may spend on garak scans of a product endpoint each release. How do you decide whether any of that budget should go to buffed runs — where a buff rewrites probe prompts before sending — instead of more probes?
basics
~20 sSpend the budget on breadth first: buffs add no new behaviour, they re-send prompts you already have. Buy buffed runs for a narrow slice where input normalisation or a filter is the thing under test, or to audit whether your own detectors are brittle. Treat them as an occasional audit, not a gate.
You own a recurring garak scan attached to a release pipeline. How would you split the probe selection between probes carrying fixed, checked-in prompt lists and probes that generate their prompts each run, and what do you give up with each choice?
basics
~20 sGate the pipeline on a small pinned set of fixed-prompt-list probes, because only deterministic prompts give a comparable pass/fail signal per build. Run prompt-generating probes on a slower cadence for fresh coverage, reviewed by a human. Pinned probes go stale and invite tuning; generated ones produce noisy counts that cannot block a build.
showing 31–43 of 43