You have been asked to scan a metered hosted chat endpoint with garak and to predict the spend before you launch. What multiplies out to the number of requests it will send, and which part of that product does your probe selection control?
answer
- sum over probes, not one product
- prompt counts differ by orders of magnitude
- completion tokens dominate billing
- model-backed detector adds calls
- pilot one cheap probe, then extrapolate
basics
~20 sRequests are roughly the sum, over each probe you selected, of that probe's prompt count times the repeat count per prompt. Your selection controls which probes and therefore the prompt counts, which differ by orders of magnitude between probes. Anything model-backed in the pipeline, such as a hosted detector, adds calls on top.
solid answer
~60 sThe cost model is a sum, not a single product: for each selected garak probe, prompts-in-that-probe x repeats-per-prompt, summed across the selection. Selection is the axis you own here, and the mistake is treating "number of probes" as the unit. Prompt counts per probe span orders of magnitude — one probe may carry a dozen handwritten prompts, another may expand a dataset into thousands — so ten probes can be cheaper than one. Two things sit outside that product and are routinely forgotten. Anything that expands the prompt set before sending multiplies it again. And a detector that is itself a model costs something: if it runs locally, compute and wall-clock; if it calls a hosted model, more billed requests, on the *response* side rather than the prompt side. Also budget for output tokens, which the request count alone does not capture — a probe that elicits long completions bills more per attempt than the prompt length suggests. So: list the catalogue, get the per-probe prompt counts, price a single cheap probe end to end first, then extrapolate.
go deeper
Should know that the run costs real money per request and that you check the probe list and estimate before pointing it at a paid endpoint.
Should give the sum-over-probes arithmetic, note that per-probe prompt counts vary hugely, and know that completion tokens, not request counts, drive the bill.
Pilots one probe to get measured per-attempt cost, accounts for model-backed detectors and prompt expansion, and puts a hard spend cap on the engagement key.
Owns the policy: what scan spend needs approval, whether scanning runs against a dedicated billing account, and how scan cost is traded against the assurance it buys.
### The arithmetic, written out Estimating before you launch is what separates a red-team engagement from an incident with your own finance team. For a garak run the model is a **sum over probes**, not one product: ``` requests ~= SUM over selected probes ( prompts_in_probe * generations ) cost ~= requests * (avg_prompt_tokens + avg_completion_tokens) * unit_price + any model-backed detector or prompt-buff calls ``` `generations` is garak's `--generations` flag: it re-sends every prompt N times so that an intermittent behaviour has a chance to appear more than once. It is a clean multiplier over the entire selection. `prompts_in_probe` is the count carried by each garak probe class, and **that is the term people get wrong.** ### Why probe count is the wrong unit Probes are not uniform units of work. One garak probe may hold a dozen handwritten prompts; another expands a public dataset into thousands. The spread is orders of magnitude, so "we selected ten probes" carries no cost information at all — ten small probes can be an order of magnitude cheaper than one dataset-backed probe. Get the real counts from `garak --list_probes` and from the probe classes themselves rather than assuming, and price the largest probe in your selection separately: it usually dominates the total, and if you have to cut, it is the one line item worth arguing about. ### Requests are not the billing unit either Hosted endpoints bill tokens, not calls, and completion tokens are typically the more expensive half. Two probes with identical prompt counts can differ several-fold in cost because one elicits a two-line refusal and the other elicits a long compliant essay. A request-count estimate is therefore a floor, not a forecast. Use a **measured** average tokens-per-attempt from a pilot rather than an assumed one, and remember that the probes most likely to matter — the ones that actually elicit output — are the expensive ones. ### What sits outside the product Two components routinely go missing from the estimate: | component | what it adds | |---|---| | a prompt-expansion step (garak's buffs, which paraphrase or re-encode prompts before sending) | multiplies the whole sum again by the expansion factor | | a model-backed detector | local compute and wall-clock if it runs on your machine; **additional billed calls on the response side** if it is a hosted model scoring each reply | | retries on throttling or timeout | silently inflates real request count above the planned figure | The detector line is the one that catches people: they budget the prompt side carefully and forget that judging the answers is itself inference. ### Where the estimate misleads Three characteristic failures, in rough order of how often they bite: 1. **Extrapolating from a cheap pilot probe.** If the pilot elicits short refusals and the real selection elicits long completions, the request count is right and the token bill overshoots badly. Pilot on a probe whose output length resembles the bulk of your selection. 2. **Averaging away the heavy tail.** A mean tokens-per-attempt hides a probe that emits maximum-length output every time. Estimate the largest contributor separately and add it in. 3. **Treating wall-clock and spend as the same constraint.** A rate-limited endpoint fixes your throughput, so a run can be affordable and still take three days. Estimate both numbers; stakeholders care about different ones. ### What you would check before launching Run one small probe end to end at low `--generations` against the real target. That single pilot gives you a measured per-attempt token cost, confirms the generator wiring, and shows you how the endpoint behaves under throttling. Then extrapolate with the real per-probe prompt counts, present the figure **with the selection attached** — a cost number without its selection is unfalsifiable — and put a hard stop outside garak: a spend limit or rate limit on the engagement's own API key, plus splitting the sweep into batches you can stop between. The point of the cap is that when the estimate is wrong, and it will sometimes be wrong, the run fails safe instead of quietly billing all night.
- Why is 'we selected ten garak probes' a useless cost statement?Because a probe is not a unit of work. Per-probe prompt counts span orders of magnitude, so ten small probes can be far cheaper than one dataset-backed probe.
- You measured cost with a pilot probe and extrapolated. Where does that extrapolation typically go wrong?On completion length. If the pilot probe elicits short refusals and the real selection elicits long outputs, the token bill scales past the estimate even though the request count is right.
- What protects you if the estimate is simply wrong?A hard cap outside garak — a spend limit or rate limit on the API key used for the engagement — plus splitting the sweep into small runs you can stop between.
A probe is not a unit of work, the same way a box is not a unit of weight: ten boxes tell you nothing about the shipping bill until someone says what is inside them. The dataset-backed garak probe is the box full of bricks, and --generations is being charged for the whole shipment N times over.
saying these in an interview costs you the question
- Treating probe count as the cost unit.
- Estimating from prompt tokens only and ignoring completion tokens.
- Launching a full sweep against a metered endpoint with no cap and no estimate.
- Forgetting that a model-backed detector or a prompt-expansion step multiplies the run.