You own a recurring garak scan attached to a release pipeline. How would you split the probe selection between probes carrying fixed, checked-in prompt lists and probes that generate their prompts each run, and what do you give up with each choice?
answer
- gate = deterministic, blocks builds
- exploration = generated, scheduled, triaged
- graduate findings into pinned cases
- pin selection, attempts, package version
- re-pin on a schedule, announce the break
basics
~20 sGate the pipeline on a small pinned set of fixed-prompt-list probes, because only deterministic prompts give a comparable pass/fail signal per build. Run prompt-generating probes on a slower cadence for fresh coverage, reviewed by a human. Pinned probes go stale and invite tuning; generated ones produce noisy counts that cannot block a build.
solid answer
~50 sTwo jobs, two cadences. **The gate** is deterministic. A pinned selection of fixed-prompt probes, plus the concrete prompts graduated from past findings, runs on every build. It is fast, its result is comparable across releases, and a red is actionable because the same input failed. What you give up is discovery: this set only ever finds what it already knows, and over time the team optimises against it — the classic effect of turning a measure into a target. **The exploration** is scheduled, not blocking. Prompt-generating probes run nightly or weekly with a larger attempt budget, and their output goes to a human for triage rather than to a build status. What you give up is a clean signal: counts move run to run, so nothing here can fail a build honestly. The connective tissue is graduation — every triaged finding from the exploration lane becomes a pinned deterministic case in the gate, so the gate grows with what you have actually learned.
go deeper
Says fixed-prompt probes repeat and generated ones do not, so only the first kind gives a stable pipeline check.
Puts deterministic probes on the gate and generating probes on a schedule, and names the staleness and noise costs.
Pins selection, attempt counts and package version, persists artefacts, and graduates triaged findings into permanent deterministic cases.
Owns the whole programme: cadence, budget and endpoint-owner agreement, the re-pinning ritual and its broken trend line, and the standing message that a green gate is a regression signal, not assurance.
**The principle.** A pipeline gate needs a signal whose *change* means the system changed. A garak probe that mints prompts at run time does not have that property: its count moves on an unchanged target because each run is a new sample. Put it on a blocking check and you get one of two outcomes — flaky reds that the team learns to re-run until green, or a threshold loosened until it never fires. Static, checked-in prompt lists do have the property, but they decay: published strings get absorbed into filters and training data, and the set stops discovering anything. Design around both facts instead of picking one. **Lane 1 — the blocking gate.** A pinned selection of fixed-prompt probes, named class by class in a config file under version control (`garak --config <file>` rather than a hand-typed `--probes` argument in a CI script), with pinned attempt counts, a pinned `--generations`, and a **pinned garak version in the lockfile** so a catalogue or detector change arrives as a deliberate upgrade rather than as a mystery red on someone's unrelated PR. Add every prompt graduated from a real triaged finding; that is the part carrying institutional memory. Budget it honestly: 300 pinned prompts at `--generations 3` is ~900 generator calls per build, which at fifty builds a week is ~45,000 calls a week charged to whoever owns the endpoint — worth agreeing in advance, and worth noting that dropping to `--generations 1` cuts the bill threefold but stops sampling a stochastic target more than once, so intermittent failures become intermittent gate results. *Cost accepted:* it only ever measures known ground, and teams optimise against a visible check. Mitigate by refreshing on a schedule, and by never letting a green gate be quoted as a safety claim. **Lane 2 — scheduled exploration.** Generative probes, larger budgets, on a cadence the metered endpoint and its owner can absorb — nightly or weekly, never per-build. Their output is a triage queue, not a build status. Persist prompts and replies so any hit can be reproduced and handed to the owning team. *Cost accepted:* variance, real human triage effort per run, and no authority to block anything. **Lane 3 — the deliberate re-pin.** Periodically rebuild the gate set: drop cases the product has structurally resolved, add graduated ones, refresh against the current catalogue, upgrade the pinned version, re-baseline. Announce it, because the trend line breaks at that point and comparing across the break is a mistake people make quietly and confidently. **Where the pipeline's numbers mislead.** - *A green gate is a regression signal, not assurance.* It says the pinned prompts still fail to elicit the behaviour. It cannot say anything about prompts not in the set, which is most of them. - *Goodhart, on a schedule.* Once the pinned list is visible in the repository, the cheapest way to keep it green is to handle those exact strings. The metric improves; the property may not. - *A pinned version freezes the denominator.* That is the point — but it also means "coverage" stops growing while the catalogue and the threat landscape do not, so the gap widens silently until the re-pin. - *Aggregates hide the shape.* A build-status badge is one bit. Release review should read per-probe lines, not the badge. **What is really being asked.** Whether you know that a scanner in CI is an instrument for detecting regressions on known ground, and that assurance comes from somewhere else: the exploration lane, human red-teaming against the actual product surface, and a review that reads the per-probe results. The honest write-up at release says which of those a green pipeline did — and which it did not.
- Why not simply set a failure-count threshold for the generating probes and gate on that?Because the count varies run to run on an unchanged system. A tight threshold flakes and gets ignored; a loose one never fires. Neither is a gate, and both erode trust in the pipeline.
- What breaks when you refresh the pinned gate set, and how do you handle it?The historical trend is no longer comparable. Re-baseline explicitly, date the change in the record, and never plot a line across the break as if it were continuous.
saying these in an interview costs you the question
- Puts prompt-generating probes on a blocking gate and then loosens thresholds until it never fires.
- Treats a green pipeline scan as a safety assurance claim.
- Never graduates triaged findings into pinned deterministic cases.
- Leaves the package version floating so a catalogue change silently moves the baseline.
- Re-pins the gate without telling anyone, then compares the trend across the break.