skip to content

You are running garak against a metered hosted endpoint and the repeat count you want per prompt would blow the query budget. How do you get a usable number anyway, and which shortcuts would make the resulting rates worthless?

level: seniorimportance: should knowfreq 44%

answer

  1. two-pass: screen shallow, measure deep
  2. pool independent runs, don't blend
  3. never buy repeats with deterministic decoding
  4. same target, same count, or it's not a comparison
  5. errored attempts are missing, not passes

basics

~20 s

Spend repeats where they decide something: a high count on the few probes gating the release, a low one elsewhere, and pool several cheap runs rather than one big one. Never buy repeats by pinning the generator's decoding to be deterministic when production samples, that measures a system you do not ship.

solid answer

~50 s

Treat the repeat count as a budget you allocate, not a global constant. Run the broad sweep at a low count to find where anything moves at all, then re-run just those probes at a high count to measure the rate properly. That two-pass shape gets most of the resolution for a fraction of the queries. If the endpoint is rate-limited rather than purely metered, several independent low-count runs spread over time can be pooled — and they have the side benefit of sampling different endpoint states rather than one five-minute window. The shortcuts that destroy the number: dropping decoding temperature so responses stop varying (you now characterise a configuration you do not serve); pointing at a cheaper sibling model and reporting the result as the shipped one; and quietly changing the repeat count between baseline and retest so the before-and-after rates are not comparable. Also watch retried and errored attempts, which shrink the denominator without anyone noticing.

go deeper

for a junior

Suggests running fewer prompts or fewer repeats to fit the budget.

for a middle

Proposes the two-pass screen-then-deepen shape and knows repeats multiply queries linearly.

for a senior

Allocates the repeat count per probe against the decision it feeds, pools independent runs, reconciles dropped attempts, and names the shortcuts that void the number — deterministic decoding, substituted targets, drifting counts.

for a principal

Sets the budget policy and the disclosure rules so that rates produced under different allocations remain interpretable across teams.

**Allocate the count, do not flatten it.** One repeat count applied uniformly across a whole sweep is simultaneously the most expensive option and the least informative: you pay maximum depth on probes that returned nothing at all and get minimum useful precision on the handful that did. The workable shape is two-pass. 1. **Screen wide, shallow.** A low repeat count across the full probe selection. Its only job is to find probes with any non-zero rate. Say out loud, in the write-up, that this pass cannot see low-rate behaviour — that admission is the price of the cheap pass. 2. **Measure narrow, deep.** Re-run only the probes that moved, at a count derived from the firing rate you need to resolve. This is where the budget goes, and it lands on a small fraction of the prompts. The arithmetic is why this works. Suppose 2,000 prompts and a desired count of 30. Uniformly that is 60,000 calls. Screening all 2,000 at `-g 2` (4,000 calls) and then deepening the 8% of prompts that fired at `-g 30` (about 4,800 calls) is under 9,000 calls — roughly a seventh of the spend, with the depth landing where a decision actually turns on it. At a rough 600 tokens per call and three requests per second, that is the difference between most of a working day and about an hour. **Pool independent runs instead of one giant one.** Where the binding constraint is a rate limit or a per-day quota rather than total spend, three runs of `-g 10` beat one run of `-g 30`: the same attempt count, plus samples of different endpoint states — load, cache warmth, a backend rollout mid-week. Keep the runs separate in the record so an outlier run is visible as an outlier instead of being blended into an average that hides it. **Shortcuts that void the number.** Each of these makes the run cheaper and the result meaningless, and each is reached for under budget pressure. - **Removing the randomness you were trying to measure.** Setting the generator to deterministic decoding makes repeats cheap and pointless: N near-identical outputs sample one point of a distribution production does not sit on. The bias is not even in a predictable direction — the problem is the mismatch itself. If you genuinely serve deterministic decoding, then say so and use a low count deliberately. - **Substituting a cheaper target.** A smaller or differently tuned endpoint has its own rates. The number belongs to whatever the garak generator pointed at, not to what you ship, and the two are not related by any scaling factor you can apply afterwards. - **Silently changing the count between baseline and retest.** The delta then mixes a real change with a measurement change, and nobody can separate them after the fact — including you, next quarter, when the delta is being cited as a fix. - **Truncating the prompt set to afford repeats without recording it.** That is a coverage cut disguised as a depth gain, and the resulting rate is over a different population than the baseline's. - **Counting errored or throttled attempts as passes.** Under a tight budget these become common — you are pushing the endpoint harder — which is exactly when the shortfall most needs reporting. **Where the resulting number still misleads, even done well.** A two-pass result carries two different counts, so a probe-level rate from the screen and one from the deep pass are not comparable, and an aggregate over both is meaningless. Report them as two populations. And the screening pass's clean probes are *screened*, not cleared — the write-up has to say which probes were only screened, or readers will treat the whole sweep as if it had the deep pass's resolution. **Accounting hygiene while the budget is tight.** Reconcile scored totals against prompts times repeats on every probe. Track spend per probe so the next run's allocation is evidence-based rather than a fresh guess. And attach the repeat count to every rate you hand anyone, because under an allocated budget the counts differ between probes and a bare rate has stopped meaning anything at all.

  • Why prefer several pooled runs over one large run when the endpoint is rate-limited?
    Same total attempts, but spread across time you also sample different endpoint states — load, cache warmth, backend rollouts — and you can spot an outlier run instead of blending it in.
  • You must cut the budget in half. Depth or breadth?
    Depends what the run is for. A first look at an unknown target needs breadth; a release gate on a known risk needs depth. Whichever you cut, record it, because both change what the number means.

Pinning the generator to deterministic decoding so you can afford more repeats is like weighing yourself ten times on a scale that is stuck: the readings agree beautifully, and none of them is about you.

saying these in an interview costs you the question

  • Turning off sampling in the generator to make repeats cheap, without noticing production still samples.
  • Scanning a cheaper model and reporting the rate as the shipped system's.
  • Changing the repeat count between baseline and retest and calling the delta a regression or a fix.
  • Folding throttled or errored attempts into the passing count.
  • Spending the whole budget on uniform depth across probes that returned nothing.

context