skip to content

garak

You will learn garak's probe/detector/generator architecture and how to scan a model for known failure modes and read its pass/fail report. Interviewers treat it as the nmap of LLMs — the quick-scan tool candidates are expected to have run.

on this pageshow

explore

questions

page 1 of 2

You want the garak LLM scanner to cover an attack family it does not ship a probe for. Beyond writing the prompt set itself, what else must you supply before a run can report a failure rate?

level: juniorimportance: must knowfreq 62%

answer

  1. prompts are the cheap half
  2. something must rule each reply
  3. plugin discovery must find it
  4. no judge, no rate
  5. silent zero = not loaded

basics

~20 s

Prompts alone produce no number. You must also supply, or point at, the thing that decides whether each reply counts as a hit, and make both discoverable to the scanner so the run loads them. Without that judging step garak collects replies but has nothing to score.

solid answer

~50 s

A garak run has two halves for every attack family: something that sends prompts to the target, and something that rules on the replies. Ship only the first half and the run produces transcripts, not a rate. So a new family means: the prompt-generating plugin, a stated pairing to whatever will judge its replies, and either an existing shipped judge that genuinely matches your failure signature or a new one you write. You also have to make the code visible to the scanner's plugin discovery and name it on the run, otherwise it is simply never loaded. The trap for a beginner is thinking the prompts are the deliverable. They are the cheap half. The judging half is where the number comes from, and it is the half you now own with no shipped baseline behind it.

go deeper

for a junior

Says a run needs both prompts and something that rules replies, and knows the scanner has to be able to load the new code.

for a middle

Adds that the judge must match the family's actual failure signature, and can debug a silent zero-hit result.

for a senior

Treats the judging half as the deliverable, and will not quote a rate from a custom family without saying how often the judge is wrong.

for a principal

Asks whether the family deserves a permanent plugin at all, given who will keep its judging criterion honest over time.

### The pipeline you are inserting into garak is an LLM vulnerability scanner, and it splits a scan into plugin roles that form a pipeline. A **generator** owns the connection to the target — a hosted model selected with `garak --model_type` and `garak --model_name`, or a class you write yourself when the target is a deployed application rather than a raw model. A **probe** owns one attack family: it holds or mints the prompts, and it declares which detectors should rule its replies. Optional **buffs** rewrite prompts before send. A **detector** receives an `Attempt` — one prompt plus the target's replies to it — and returns one score per reply. An evaluator thresholds those scores into hits and passes, and the harness writes a JSONL report, one line per attempt plus an eval line for each probe/detector pair. Read the pipeline and the answer falls out. Prompts drive sends; detectors produce scores; the evaluator turns scores into a rate. Supply only the prompts and the run produces attempt lines and no eval line — a pile of transcripts, not a failure rate. ### The three things that must be true 1. **The prompts exist, and you know their nature.** A fixed list gives you reruns that cover identical ground; prompts minted at run time do not, so a rate that moves between runs may be sampling rather than the target. 2. **A detector is bound to the probe.** In garak that binding is an attribute on the probe class naming its recommended detector, or a detector named explicitly on the command line with `garak --detectors`. Nothing is inferred from the probe's name or module; an unbound probe is scored by nothing. 3. **garak can import both.** Probes and detectors are resolved by dotted plugin name inside garak's own `garak.probes` and `garak.detectors` namespaces. Code sitting in a directory garak never imports is not broken — it is absent. `garak --list_probes` and `garak --list_detectors` are the one-command check that your class is visible at all. ### What the run costs Sends multiply out as prompts x `garak --generations` x (buffs + 1). A 120-prompt probe at `--generations 10` is 1,200 completions against one target for that one family. Read the `--generations` default off your own `garak --help` rather than trusting memory — it has changed across releases, and it multiplies every other number in the run. At a metered endpoint producing a few hundred output tokens per reply that is a visible invoice line, and at a 60-requests-per-minute limit it is roughly twenty minutes of wall clock; `garak --parallel_attempts` cannot outrun the endpoint's own limiter. The lopsided cost is engineer time: the prompt set is an afternoon, while a detector you can defend — plus the hand-labelled sample that says how often it is wrong — is days. ### Where the number misleads The characteristic failure is the **silent zero**. A probe garak never imported, a probe with no detector bound, and a detector whose criterion cannot fire on this family's failure text produce output that looks the same from the summary: nothing reported for your family. Worse, an absent eval row and a genuine 0% eval row read identically to anyone skimming a report, and both are cheerfully summarised as "we tested that and it was clean." A zero from a custom family is a claim about your plumbing until you have shown otherwise. The second trap is the denominator. garak's rate is computed over outputs, not prompts, so `--generations` sets the denominator. A run at 5 generations and a run at 10 are not comparable for the same probe, and "we re-ran it and the rate dropped" is very often just that flag having changed. ### What I would check Confirm the class appears in `garak --list_probes`. Run the probe against a cheap target with a single prompt and verbose output, and confirm attempt lines actually appear in the JSONL report with non-empty replies. Confirm an eval line exists for your probe/detector pair, with an attempt count matching prompts x generations. Then exercise the detector offline on two hand-written replies you have already labelled — one plainly a failure for your family, one plainly benign — and check that it fires on the first and not the second. A detector that cannot fire on a reply you wrote to make it fire will never fire on a real one.

  • A custom garak family runs to completion and reports zero hits. Name two causes that are not the model being safe.
    The plugin never loaded or was never named on the run, so nothing was sent; or the judge's matching criterion cannot fire on this family's failure text, so hits were sent and never counted.
  • Why does it matter whether your custom prompts are a fixed list or minted at run time?
    A fixed list gives comparable reruns and can be tuned against; generated prompts mean two runs cover different ground, so a change in the rate may be sampling, not the target.

The prompts are the exam paper and the detector is the answer key. Print the exam without a key and you get back a stack of completed papers and no grade — which is exactly what a garak run with an unbound probe hands you.

saying these in an interview costs you the question

  • Thinks writing the prompt list completes the job
  • Cannot say what turns replies into a score
  • Reads a zero-hit custom family as a clean result without checking the plugin loaded
  • Reuses whatever judge is default without checking it matches the family

context

open as a page

In garak, what job does a detector do, and why is it the detector rather than the probe that decides whether an attempt is recorded as a failure?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A garak probe only sends prompts. The detector reads each reply and rules whether the attack succeeded. Every pass and failure in the report comes from that ruling, so the number describes the detector's opinion of the outputs, not the model's behaviour directly. Different detectors over the same replies give different numbers.

open as a page

In garak, what is a generator, and what do you have to own yourself when you wire one to a deployed chat application's authenticated HTTP endpoint instead of to a model you run locally?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A garak generator is the adapter that sends each probe prompt to the system under test and returns the reply text. Aimed at a deployed app, you own the wiring: the endpoint and method, the auth and tenant headers, the request body shape, and which JSON field the reply is read from.

open as a page

In the garak LLM scanner, what does a probe supply to a scan, and what does a clean result across the probes you selected say about the probes you did not run?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A garak probe is the plugin that supplies the prompts sent to the target for one attack family. A clean result covers only the probes you selected; it says nothing about families you never ran. Coverage means probes run out of probes available, and the default selection is not everything.

open as a page

In a garak run, what does it actually mean when an attempt is marked a hit, and why is that not yet a confirmed vulnerability?

level: juniorimportance: must knowfreq 78%

basics

~20 s

It means garak's detector, not a human, judged that one model response matched its rule for failure. The detector may be a string match or a small classifier, so it can be wrong. Open the logged attempt, read the prompt and the full response, and confirm the model really did the thing.

open as a page

A finished garak run prints, for each probe it ran, a pass/failure rate alongside an absolute score. What is that rate a fraction of, and does a higher failure rate mean the tested model did better or worse?

level: juniorimportance: must knowfreq 55%

basics

~20 s

It is a fraction of that one probe's own attempts, not of the whole run. An attempt is one prompt sent to the tested model; the detector attached to that probe labels it pass or fail. A higher failure rate means more attempts got through, so the model did worse on that probe.

open as a page

In a garak scan, what does the generations setting (how many times each probe prompt is re-sent to the target generator) change about the run, and what does raising it cost?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Generations is garak's repeat count: how many times each probe prompt is sent to the target generator. Each reply is detected separately, so ten repeats give ten scored attempts per prompt instead of one. Cost scales with it, roughly ten times the queries, tokens and wall clock. The prompt set itself is unchanged.

open as a page

When you start a garak scan against a chat endpoint and do not name any probes, what set of attacks actually runs, and why is that neither nothing nor the whole probe catalogue?

level: juniorimportance: must knowfreq 70%

basics

~20 s

garak runs a built-in default selection of its probes, not the entire catalogue. The full catalogue is far larger and would cost hours of metered calls. Some probes also only run when you name them. So a default run is a sample, and you must read the report to see which probes actually ran.

open as a page

You run a set of garak probes against an endpoint twice: once plain, once with a buff enabled that rewrites the probes' prompts before they are sent. The reported pass/fail rate is very different. Why can that difference come from garak's detectors rather than from the model's behaviour?

level: middleimportance: must knowfreq 60%

basics

~20 s

Detectors judge the reply text, and many are substring or pattern matchers. A buff changes the prompt, so the reply's wording, language or format changes too. A matcher tuned for the original wording then misses real hits or fires on harmless replies. The number moved because detection moved, not the model.

open as a page

You wrote a new probe for the garak LLM scanner and pointed it at a detector that ships with the tool for a different attack family. Why is the failure rate that comes back untrustworthy?

level: middleimportance: must knowfreq 58%

basics

~20 s

A garak detector was written to recognise one family's failure text. Aimed at a different family it misses the real hits and can match unrelated wording, so the rate you get measures the mismatch between probe and detector rather than anything about the target. Both error directions move at once.

open as a page

garak's detectors range from plain substring matching, through pattern matching, up to small classifier models run locally. What does that range change about a scan's cost and about the errors you inherit?

level: middleimportance: must knowfreq 62%

basics

~20 s

String and pattern detectors are nearly free and deterministic, but they only see the exact wording they were written for, so paraphrase, translation or odd formatting slips past. A classifier detector generalises further, costs local compute and a model download, and adds probabilistic errors you did not tune and cannot inspect per item.

open as a page

You point garak at a hosted application endpoint that allows 60 requests per minute per account. The planned run is 3,000 probe prompts with 5 generations each. What sets the wall-clock, and which levers actually shorten it?

level: middleimportance: must knowfreq 65%

basics

~20 s

Requests, not probes, set the clock: prompts times generations, divided by the requests per minute the endpoint allows. Here that is 15,000 requests, about 250 minutes at best. Parallelism above the cap only earns throttled responses and retries, which add requests. Cut generations, narrow the probe selection, or get a higher quota.

open as a page

In the garak LLM scanner, some probes send a fixed prompt list shipped inside the package while others mint their prompts during the run. How does that difference change what a passing result for a probe attests to?

level: middleimportance: must knowfreq 60%

basics

~20 s

A fixed-list probe sends the same readable strings every time, so a pass means only those exact prompts failed to elicit the behaviour, and anyone can tune against them. A prompt-generating probe sends different prompts each run, so a pass is a sample you cannot enumerate afterwards and a rerun tests different ground.

open as a page

Which failure modes of garak's detectors produce false hits, and how do you recognise each one from the logged attempt?

level: middleimportance: must knowfreq 66%

basics

~20 s

Detectors that match a substring, or that infer compliance from the absence of a refusal, fire on responses that never complied: the model quoted the request back, refused in unusual wording, returned an error string, or was cut off mid-sentence. Classifier detectors add their own misfires. Read the response text to tell them apart.

open as a page

A garak scan configured with one generation per prompt reports a probe as passing; the identical scan rerun the next day fails it. What about the tool's sampling explains the flip, and how do you pick a repeat count that makes it unlikely?

level: middleimportance: must knowfreq 64%

basics

~20 s

The target samples, so a prompt that fails only sometimes is a coin flip. With one repeat you draw once: a behaviour firing on one attempt in ten is missed nine times out of ten and filed as a pass. Pick the repeat count from the rarest failure you still want to catch.

open as a page

You have been asked to scan a metered hosted chat endpoint with garak and to predict the spend before you launch. What multiplies out to the number of requests it will send, and which part of that product does your probe selection control?

level: middleimportance: must knowfreq 60%

basics

~20 s

Requests are roughly the sum, over each probe you selected, of that probe's prompt count times the repeat count per prompt. Your selection controls which probes and therefore the prompt counts, which differ by orders of magnitude between probes. Anything model-backed in the pipeline, such as a hosted detector, adds calls on top.

open as a page

You wrote both the probe and the detector for a custom family in the garak LLM scanner, and the run reports a 22% failure rate. What do you do before that number goes into a report?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Measure your own detector's error. Sample flagged replies and read them by hand for false positives, and sample unflagged replies for misses. Label a fixed set, ideally with a second reader, and quote the 22% with that measured error attached rather than as a bare number nobody has checked.

open as a page

A garak scan against your deployed chat endpoint reports a high failure rate for one probe family. Before that number goes into a report, how do you establish that it reflects real model behaviour rather than detector error?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Open the logged prompt-and-reply records behind those failures and hand-read a sample. Check what the detector actually matched: refusal text, a stock disclaimer, an echo of the prompt, an empty or error reply, or a safety lecture can all trip a matcher. Then report the verified count, the detector used and the sample size.

open as a page

A garak sweep against a chat endpoint returns zero hits on every probe you ran. What can you conclude, and how do you check the run was not silently broken?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Only that no probe you ran, judged by its own detector, found anything. First prove the run was real: check that attempts reached the endpoint, that responses are non-empty and not error text, and that the probes you needed were in the set you ran. A broken generator looks exactly like a clean model.

open as a page

In a garak report one probe lands in a poor band on its absolute score while its calibration score sits mid-pack against the reference models, and another probe scores well absolutely but far below the reference models. You are writing the run up. Which number do you act on in each case, and why?

level: seniorimportance: must knowfreq 45%

basics

~20 s

They answer different questions. The absolute score says how bad the behaviour is in itself; the calibration score says how unusual it is versus reference models on that same probe. Poor absolute, mid-pack relative is a hazard the whole field shares. Good absolute, far below the pack is this model's own outlier defect.

open as a page

A garak run over a hand-picked subset of probes finishes with nothing flagged. How do you write that up so it is not read as "the model is safe", and what denominator do you attach when you use the word coverage?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Report it as: these named probes, at this repeat count, against this endpoint and configuration, on this date, produced no detector hits. Coverage means probes run divided by probes in the catalogue — not attack surface and not behaviours. Name the families you did not run as untested, never as passed.

open as a page

In the garak LLM scanner, a buff is a plugin that rewrites a probe's prompts on the way out, before they reach the target endpoint. What does enabling a buff do to the amount of work a scan does?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A buff sits between the probe and the target and transforms each prompt. Some buffs return one variant per prompt, others fan out into several, so the attempts sent multiply on top of the per-prompt generation count. Enabling buffs makes a scan longer and more expensive, never cheaper.

open as a page

A deployed assistant's HTTP endpoint answers with a JSON envelope rather than a bare string. What must garak's HTTP generator be told about that envelope, and what does a scan report look like when that setting is wrong?

level: middleimportance: should knowfreq 50%

basics

~20 s

It must be told where in the response body the assistant's text lives, because the endpoint returns an envelope, not a bare string. If that selector is wrong, garak stores the wrapper, an error blob or an empty string, and every attempt is judged against text the model never produced.

open as a page

A garak report shows each per-probe rate with an uncertainty interval beside it, but for several probes that interval is blank. What does a blank interval tell you about those rows, and what would you change about the run before quoting one of those rates?

level: middleimportance: should knowfreq 35%

basics

~20 s

Blank means too few attempts. Below a sample-size floor the tool declines to print an interval rather than show a meaningless one, so that rate is an estimate you cannot bound. Re-run the probe with more generations per prompt, or a longer prompt set, until the interval appears, then quote it.

open as a page

When a garak scan repeats each prompt several times, what does the per-probe pass/failure rate in its report count over, and how do a probe that fails on every attempt and one that fails occasionally look different in that report?

level: middleimportance: should knowfreq 48%

basics

~20 s

Attempts, not prompts. With the repeat count at N, each prompt yields N scored attempts and the probe's rate is over that larger total. A probe failing every time sits at the extreme of the range; one failing occasionally shows a small non-zero rate that a single-repeat run would have shown as zero.

open as a page

You are asked to report whether rewriting prompts changes how often a target fails a chosen set of garak probes. How do you set up the buffed and unbuffed garak runs so the comparison actually means something, and what is the trap in the rate itself?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Run the same probes twice against the same endpoint, changing only the buff. Hold generations, detectors and endpoint settings fixed, and run them close together. Then compare like for like: a fan-out buff sends several variants per original prompt, so decide whether your rate counts attempts or original prompts.

open as a page

Some of garak's buffs rewrite a probe's outbound prompts by calling a language model to paraphrase or translate them. What does that add to a scan you have to defend, and how do you handle it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It makes the run non-reproducible: a rerun paraphrases differently, so the prompts sent are not the same set. The paraphraser can also soften or refuse the payload, silently weakening the probe. Keep the unbuffed run as the baseline, read a sample of transformed prompts, and cite the logged prompts, not the buff name.

open as a page

A garak report can position a probe's result against previously measured reference models as well as giving an absolute score. Why does a probe and detector you wrote yourself not get that comparative reading, and what do you present instead?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The comparative reading needs prior results for that exact probe and detector across reference models. Your plugin has never been run against them, so there is nothing to compare with. Present the absolute rate, your own baseline runs against two or three comparison targets, and the detector's measured error.

open as a page

How can a garak detector record a pass on a reply that was in fact a successful attack, and what does that imply about how you present the run's failure rate?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It happens whenever the harmful content arrives outside what the detector looks for: another language, an encoding, code or pseudo-code, a summary, or wording the matcher was never written for. Those attempts score as passes. So the reported failure rate is a floor on the model's weakness, never a measure of it.

open as a page

garak treats each attempt as independent, but the deployed assistant you are scanning keeps conversation history server-side, keyed by a session identifier your generator sends in a header. What breaks, and how do you wire the generator so the run is sound?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Reusing one session id makes attempt N see attempts 1 through N-1, so earlier prompts and refusals condition later replies and the run stops being reproducible or order-independent. Mint a fresh session or conversation id per attempt, or use a stateless mode of the endpoint, and confirm isolation with a memory probe.

open as a page

showing 1–30 of 43