skip to content

You point garak at a hosted application endpoint that allows 60 requests per minute per account. The planned run is 3,000 probe prompts with 5 generations each. What sets the wall-clock, and which levers actually shorten it?

level: middleimportance: must knowfreq 65%

answer

  1. requests = prompts x generations
  2. endpoint RPM is the floor
  3. parallelism above cap buys throttling
  4. retries inflate cost, not coverage
  5. cut generations or narrow probes

basics

~20 s

Requests, not probes, set the clock: prompts times generations, divided by the requests per minute the endpoint allows. Here that is 15,000 requests, about 250 minutes at best. Parallelism above the cap only earns throttled responses and retries, which add requests. Cut generations, narrow the probe selection, or get a higher quota.

solid answer

~60 s

**Do the arithmetic first.** Total requests = prompts x generations per prompt, summed over the probes you selected. 3,000 x 5 = 15,000 calls. At 60 per minute that is 250 minutes of pure floor, before latency variance, retries and any warm-up. **The binding constraint is remote.** Your machine, your thread count and the generator's parallelism setting do not move that floor once you are at the cap. Pushing concurrency higher converts successful calls into throttled ones; if the generator retries them, you have *increased* total requests for the same coverage. **Levers that genuinely work**, roughly in order of honesty: 1. Fewer generations per prompt — linear saving, at the cost of a noisier per-prompt estimate for stochastic replies. 2. A narrower probe selection — but say in the report which families were dropped. 3. A raised quota or a dedicated non-production instance — the only lever that buys throughput without buying it from coverage. Decide before the run, not after, because a run stopped halfway has a coverage denominator nobody can quote.

go deeper

for a junior

Multiplies prompts by generations and divides by the rate limit to get a floor.

for a middle

Explains why parallelism beyond the cap backfires and which levers trade coverage for time.

for a senior

Knows the retry behaviour in use, watches observed rate versus cap during the run, and negotiates quota or a dedicated instance instead of quietly truncating.

for a principal

Turns the request count into a budget and schedule commitment up front and sets the coverage the report will be allowed to claim.

### Why the question is asked Candidates habitually size a scan by counting probes. The endpoint counts *requests*, and the two differ by a multiplier that is easy to forget: garak's `--generations` flag asks for that many replies per prompt, because model replies are stochastic and one sample per prompt is a coin flip rather than an estimate. Its default is greater than one (10 in most releases), so a plan quoted in probes can be off by an order of magnitude before anyone has typed a command. ### The arithmetic ``` requests = sum over selected probes of (prompts_in_probe * generations_per_prompt) floor_time = requests / allowed_requests_per_minute real_time = floor_time + retries + latency stalls + credential refreshes ``` For the case in the question: 3,000 prompts times 5 generations is 15,000 requests; at 60 requests per minute that is 250 minutes — a little over four hours — of pure floor, before anything goes wrong. Only `real_time` is observable, and only `requests` is under your control. ### The binding constraint is remote Your CPU count, your thread pool and garak's own `--parallel_attempts` / `--parallel_requests` settings do not move that floor once you are pinned at the endpoint's cap. Pushing concurrency higher converts successful calls into throttled ones. If the generator retries them — garak's REST config names the throttle statuses in `ratelimit_codes` — you have *increased* total requests and total spend for identical coverage. There is exactly one case where more parallelism is the right lever: when observed throughput is well below the cap and no throttles are coming back, which means the bottleneck is per-request latency or a serialised generator rather than quota. ### Levers that genuinely shorten it 1. **Fewer generations per prompt.** Linear saving, paid for in sampling precision: at one generation, a rare failure mode that surfaces in one reply out of five is a coin flip to observe at all, and run-to-run comparisons get noisy. 2. **A narrower probe selection.** Also linear, and it must be declared — the report has to name which families were dropped, or its rate has no denominator anyone can interpret. 3. **A raised quota, or a dedicated non-production instance.** The only lever that buys throughput without buying it from coverage, and the only one worth a conversation with the operator. Everything else — faster hardware, longer `request_timeout`, a bigger machine — moves nothing. ### What it costs beyond time 15,000 completions at realistic prompt and reply lengths is a real bill against whichever account issued the credential in the generator's `headers`, plus four hours of an engineer's attention because a sweep this long will hit a token expiry, a deploy, or a throttle storm. Bring the *request* count to the budget conversation, not the probe count; the request count is the number that maps to both money and hours. ### Where the number misleads Three readings to distrust. **A truncated run.** Kill a sweep halfway and you have results for an unstated subset; quoting a rate from it silently changes the denominator, and it is not comparable with any full run. **Silent retries.** They make the bill and the clock grow while coverage stays fixed, so a run that "took longer than planned" can mean nothing was added. **Throttled attempts stored as replies.** If throttle bodies land in the corpus rather than being retried, they are scored as ordinary non-violating replies and dilute the failure rate downward — the run looks safer precisely because it was rate-limited. ### What you check Before: run a 50-request pilot, measure achieved requests per minute, and extrapolate rather than trusting the documented quota. During: watch observed rate against the cap, the throttle-response share, the error share, and elapsed time against the projected floor; a large gap in either direction is a finding about your wiring, not about the model. After: report the request count, the wall-clock, the generations setting, and the probe families included and excluded — those four make the number reproducible, and nothing else does.

  • Observed throughput is a third of the endpoint's allowed rate and there are no throttled responses. What does that tell you?
    The bottleneck is per-request latency or a serialised generator, not quota. Here raising parallelism up to the cap is the correct lever.
  • Why is dropping generations per prompt to one not free?
    Replies are stochastic. A single sample per prompt makes each per-prompt result a coin flip, so rare failures are missed and run-to-run comparisons get noisier.
  • Halfway through, the run has to stop. What can you report?
    Only results for the probe families that completed, with the request count and the excluded families named. A partial sweep has no clean coverage denominator otherwise.

The endpoint's meter counts every completion, not every prompt, so a generations setting of five means the turnstile clicks five times for each visitor you thought you were sending through. Budget the clicks, not the visitors.

saying these in an interview costs you the question

  • Estimates the run from probe count and ignores generations per prompt.
  • Proposes raising parallelism as the fix when already at the endpoint's cap.
  • Cannot say whether the generator retries throttled calls.
  • Kills a run part-way and quotes the resulting numbers without stating the reduced coverage.

context