Your organisation wants one standing policy for how many times garak re-sends each prompt in scans that gate a release. How would you set it, and what must you tell every reader of a rate produced under that policy?
answer
- tier the count, don't globalise it
- derive N from tolerated occurrence rate
- freeze for comparability, version changes
- publish the detection floor
- count under change control, not the gated team
basics
~20 sSet it per tier, not once: a small count for fast pre-merge scans, a much larger one for the release gate, and freeze it so rates stay comparable between runs. Publish the count next to every rate, and say plainly that a clean scan at a low count is weak evidence, not proof of absence.
solid answer
~50 sA single global number is the wrong shape, because the counts serve different decisions. Define tiers: a cheap pre-merge scan whose only job is catching near-deterministic regressions, and a release-gate scan whose count is derived from the smallest firing rate the organisation says it cares about. Then freeze the count per tier. The value of a standing policy is comparability — the same scan run monthly only produces a trend if the sample size is constant. Changing the count is a versioned decision that resets the trend line, not a knob an engineer turns to fit a sprint. What the policy must publish with every number: the count, the attempt total, and the detection floor it implies — that is, the per-attempt rate below which this scan is effectively blind. Without the floor stated, readers will keep reading a clean gate scan as proof the behaviour is absent, and the policy becomes a manufacturer of false assurance.
go deeper
Suggests a single number and knows more repeats cost more.
Separates a cheap CI scan from an expensive gate scan and connects the count to catching intermittent failures.
Derives the gate count from a tolerated occurrence rate, freezes it for comparability, and reconciles it against triage capacity and endpoint budget.
Owns the policy as governance: tiers, versioned change control away from the gated team, a published detection floor with every rate, and an explicit narrowing of the gate when the derived count is unaffordable.
**Why one global number fails.** The repeat count trades query spend for resolution on intermittent behaviour, and the two scans in a typical pipeline are buying different things. A pre-merge scan has to finish in minutes and only has to catch things that broke outright — a regression that fails on nearly every attempt. A release gate is allowed to be expensive and has to resolve *rates*, because the decision it feeds is whether a known intermittent behaviour is rare enough to ship. Forcing both to one value either bankrupts CI or blinds the gate; there is no single number that is honest for both. So the policy defines **tiers**: a named count per scan tier, each attached to the decision it serves. Two or three tiers is usually the whole design. **How to derive the gate's count.** Do not ask engineers to pick a number. Ask the risk owners a question they can actually answer: *what per-attempt occurrence rate would we consider unacceptable to ship?* Convert that rate p and a tolerated miss chance m into a count with `N = ln(m) / ln(1 - p)` — for a 5% miss chance, about 29 repeats at p = 0.10 and about 149 at p = 0.02 — then price it. A 2,000-prompt gate at N = 30 is 60,000 calls per run: tens of millions of tokens and, at a few requests per second, the better part of a day. If that is unaffordable, the honest response is not a quieter number. It is a **narrower gate**: keep the count on the probes tied to risks you must resolve, and state explicitly which probes are now screened rather than gated. **Freeze it, and version changes.** Comparability is the entire value of a standing policy. A monthly rate is only a trend if the denominator is constant, so the count belongs in a versioned configuration next to the probe selection — changed deliberately, announced, and with the trend line **broken** at the change rather than continued through it. Anyone who raises the count mid-series and reads the resulting rise as a regression has mixed a real change with a measurement change in a proportion that cannot be recovered afterwards. **Publish the detection floor.** Every rate the gate emits should travel with three things: the repeat count, the attempt total, and one plain sentence of the form *"this scan is effectively blind to behaviours occurring below roughly X% of attempts"*. That sentence is the single highest-value output of the whole policy, because it pre-empts the most damaging organisational failure mode: a clean gate scan being cited, months later and by someone who never read the configuration, as evidence a behaviour does not occur. Without it, a scanning programme becomes a manufacturer of false assurance at industrial scale — and the more disciplined and routine the scan looks, the more weight the false assurance carries. **Second-order effects a lead owns.** - **Triage capacity.** Raising the count raises hit volume roughly linearly on anything that fires, detector false positives included. A count whose output the triage team cannot drain produces a backlog, and a backlog produces blanket waivers — which is a worse outcome than the lower count would have been. - **Endpoint cost and rate-limit contention.** The gate shares a quota with other teams; a gate that starves production experiments will be quietly disabled. - **Gaming.** If the count is a knob the gated team can lower to get through the gate, it will be lowered, on a deadline, with a good reason. Put it under change control owned by someone other than the team being gated. - **Non-comparability across targets.** The same count against a different endpoint is still a different measurement. The policy governs the instrument; it does not make two targets' rates equivalent, and a dashboard that ranks models by gate rate will imply otherwise unless the counts and targets are shown. **What good looks like.** Two or three named tiers, each with a fixed count and a stated detection floor; a versioned change process that breaks the trend line at every change; a rule that no rate leaves the tool without its count and attempt total attached; and a periodic check that the gate's hit volume still fits the triage capacity that has to absorb it.
- The derived gate count is unaffordable across the whole prompt set. What do you do?Narrow the gate rather than weaken it: keep the count on the probes tied to the risks you must resolve, and state explicitly which probes are now screened rather than gated.
- Someone raises the count mid-quarter and the failure rate moves. How do you read the trend?You do not read it across the change. Break the series at the configuration change, and re-baseline; part of the move is measurement, and it cannot be separated after the fact.
A repeat count is the mesh size of a net. A clean haul with a wide mesh proves nothing about small fish, so the policy has to publish the mesh alongside the catch.
saying these in an interview costs you the question
- Picking one round number for every scan tier because it is simple.
- Letting the gated team lower the repeat count to pass.
- Continuing a trend line across a change in the repeat count.
- Publishing rates without the count or attempt total.
- Allowing a clean low-count scan to be recorded as evidence a behaviour does not occur.
- Setting a count whose hit volume exceeds what triage can drain.