Your team wants a risk covered that the garak LLM scanner ships no probe for. When would you decline to write a custom probe and detector pair, and what would you do instead?
answer
- maintenance, not difficulty
- criterion drifts with the target
- raters disagree = do not automate
- one-off risk = manual set
- named owner, fixture, cadence
basics
~20 sDecline when nobody will own the ruling criterion over time. A custom pair needs re-labelling as the target changes, and a drifting criterion quietly moves the number while looking stable. For a one-off question, test by hand or with a small held-out set; reserve custom plugins for risks you will re-run every release.
solid answer
~50 sThe build decision is not about difficulty - prompts and a matcher are a day's work. It is about a recurring obligation: whoever owns the plugin owns its false-positive rate forever, and that rate is not a constant. It is a property of the criterion **and** the current target, so a deployment change can invalidate it without anyone touching the code. I build when three things hold: the risk recurs on a schedule so automation pays back; a stable failure signature exists that I can write in one sentence and hand to a second labeller; and a named person will re-validate on a labelled fixture each release. I decline when the question is one-off, or when two experienced people disagree on the same reply - an unreliable automated number is worse than none, because it will be cited. The middle path: test manually first, then encode what you learned.
go deeper
Recognises that writing a probe and detector is ongoing work, not a one-off task.
Weighs how often the risk is tested against the cost of maintaining and re-validating the criterion.
Uses inter-rater agreement and target stability to decide, and runs the risk manually first to learn the real failure signature.
Sets an entry bar for the team's plugin kit — named owner, written criterion, labelled fixture, re-validation cadence — and accepts an honest absence over a confident wrong number.
### The axis this decision actually turns on There is nothing to buy here, so framing it as build-versus-buy misleads. The choice is between an automated criterion you will maintain and a human process you will repeat. Both cost, they cost differently, and the costs land on different people at different times — which is why the person writing the code is rarely the person who should decide. ### The commitment a custom pair creates Writing it is the small part: a prompt set is an afternoon, a detector another day or two. What follows it is permanent. A hand-labelled fixture that must be re-scored whenever the criterion or the target moves. A named person who can defend the criterion in a report review when somebody asks why a reply counted. A slot in the reporting standard for numbers that have no positional baseline behind them, because a custom family never gets one. Multiply that by every family somebody wants covered and the team owns a small internal library with a maintenance budget nobody costed. Put arithmetic on it, because the arithmetic is often surprising. Say the check runs quarterly. Manual: reading a 1,200-output run at fifteen seconds a reply is about five hours, so twenty hours a year. Automated: three days to build and validate, plus half a day of re-validation per run, is roughly five working days in year one and two days a year thereafter, on top of the query bill each run incurs either way. At four runs a year automation is not obviously cheaper; at monthly, or at a volume where manual reading is genuinely infeasible, it clearly is. The number of runs per year is the variable that decides, not how interesting the attack family is. ### Signals that say build The risk is re-tested every release, so the fixture cost amortises. The failure signature is mechanical — a marker that leaks, a specific structural property of the reply — so two independent labellers agree almost always. The volume is high enough that hand-reading is not a real option. The result feeds a gate somebody will actually act on. ### Signals that say do not build The question is asked once, for one launch. The failure is contextual enough that the criterion would encode your team's taste rather than a property of the reply — that shows up as inter-rater disagreement, and a criterion two experienced people cannot apply consistently produces a rate nobody should quote. The target is about to change substantially, so today's calibration expires before it is used. Or the family is genuinely covered by something already shipped, and "covered" can be demonstrated on a labelled sample rather than assumed. ### Where the numbers mislead if you build anyway Two failures are specific enough to name. First, **an unmaintained criterion keeps producing numbers**. It does not fail loudly when the target's formatting or refusal style shifts underneath it; it keeps emitting a rate, and the drift is read as a model regression. That is the expensive outcome, because it triggers work on the wrong system. Second, **coverage inflation**: adding custom probes raises "probes run" without raising attack surface reached, and a report that quotes the first as if it were the second is describing effort, not exposure. ### The alternatives worth naming A small hand-curated set read by a person, reported as qualitative findings with quoted examples — often more persuasive to an engineering team than a percentage. Reusing a shipped family after demonstrating on a labelled sample that its criterion matches yours. Or running the risk manually for a release or two, keeping every transcript, and using that corpus as the labelled fixture if you later automate — which hands you the one-sentence criterion for free, because you had to apply it by hand to produce the corpus. ### The position I would hold Any custom plugin entering the team's kit arrives with four things: a named owner, a written criterion, a labelled fixture, and a re-validation cadence. Anything that cannot meet that bar stays a manual exercise, and is reported as one. This is not process for its own sake. An automated criterion nobody re-validates still produces confident numbers, and a confident wrong number is more expensive to unwind than an honest absence — the absence prompts a question, the number ends an argument.
- What is the single strongest signal that a risk should not be encoded as an automated criterion?Two experienced labellers disagree on the same replies. If humans cannot apply the criterion consistently, no matcher will, and the resulting rate encodes taste rather than target behaviour.
- You inherit five custom plugins with no owners. What is your first move?Re-validate each on a fresh labelled sample against the current target, retire the ones whose criterion no longer holds, and refuse to publish numbers from any that lack an owner and a fixture.
saying these in an interview costs you the question
- Decides purely on how easy the code is to write
- Ignores that the criterion's accuracy depends on the current target
- Automates a judgment two experienced people cannot agree on
- Adds plugins to the team kit with no owner or re-validation plan
- Prefers any automated number over an honest qualitative finding