skip to content

You plan to run an automated adversarial-prompt campaign of tens of thousands of requests against a client's production LLM endpoint, billed to a metered model-provider account. Beyond permission to test, what do you agree before you start?

level: seniorimportance: should knowfreq 44%

answer

  1. volume is a scope term
  2. whose key, what spend cap
  3. quota contention with the live product
  4. abuse enforcement hits the account
  5. mark the traffic so it can be filtered out

basics

~20 s

Agree whose provider account and key pay for it, a spend cap, a request-rate ceiling and a stop condition. Warn that sustained adversarial traffic can trip the provider's abuse detection and throttle or suspend the account the live product uses, so ask for a separate key and a named contact who can unblock it.

solid answer

~60 s

Volume is a scope term, not an implementation detail, because on a hosted model it converts directly into money and into enforcement risk. Agree four things in writing: 1. **Whose money.** Which provider account and API key the traffic bills to, and a hard spend cap. Long adversarial prompts with long generations are far more expensive per request than typical product traffic. 2. **Rate.** A requests-per-minute and concurrency ceiling you will not exceed, chosen so you do not consume the quota the live product needs. 3. **Isolation.** A key or project separate from the one serving customers, so a throttle or suspension does not take the product down. 4. **Stop and escalation.** What halts the campaign — a spend threshold, a sustained error rate, an enforcement notice — plus a named contact on the client side reachable within a stated time. The non-obvious part is that the enforcement risk is the provider's, not the client's: automated abuse systems act on the account, and no amount of client authorisation prevents that.

go deeper

for a junior

Knows that hosted model calls cost money per request and that someone must agree to pay for the test traffic.

for a middle

Agrees a spend cap and a rate ceiling, and understands that multi-turn attempts and a scoring model multiply the billed call count.

for a senior

Isolates the campaign onto a separate key or project, sets stop conditions and escalation contacts, and anticipates provider abuse enforcement and quota contention with the live product.

for a principal

Decides how the firm prices and structures high-volume AI campaigns so that provider enforcement and cost overruns are never a client's surprise, including when to insist on locally-run infrastructure.

**What "volume" converts into on a metered endpoint.** Against a self-hosted service, a heavy scan costs CPU you already own. Against a hosted model it costs three separate things, and only the first is obvious: money billed per token, quota consumed from a pool the live product shares, and standing with an automated abuse system that can act on the whole account. None of those is a consequence the client agreed to unless you asked. That is why request volume is a scope term and belongs in writing, alongside the endpoints and the dates. **The arithmetic, and why the naive estimate is always low.** Billed calls are not attempts. They are roughly ``` billed_calls ~= attempts x variants_per_attempt x turns_per_attempt x (1 + judge_calls) billed_tokens ~= billed_calls x (mean_input_tokens + mean_output_tokens) ``` Every one of those multipliers has a named owner in real tooling. garak's `--generations` flag sets how many completions are requested per prompt, so it multiplies the call count of every probe linearly. promptfoo's `numTests` config key sets how many test cases each plugin generates, and its *strategies* axis then expands each generated case into further variants — two multiplicative axes, not one. PyRIT's multi-turn red-teaming orchestrators loop until their scorer is satisfied or a maximum-turns limit is hit, so the call count for a given attempt is *not known in advance*; the turn cap is the only thing that bounds it, and if you did not set one, you have no upper bound on the bill. Then price it: multiply billed tokens by the provider's current published per-million-token rate for that model, separately for input and output, because they are priced differently and output is the expensive side. **Where the estimate misleads.** Run a calibration batch of a few hundred attempts and you will get a mean cost per attempt that is confidently wrong in a predictable direction. Three reasons. First, the distribution is tail-heavy: most attempts end quickly, and the ones that matter run to the turn cap, so the mean is set by a tail the calibration batch under-samples. Second, the judge is usually a larger and more expensive model than the target, so a campaign that looks cheap because the target is cheap can have most of its bill on the scoring side — the "cost per attack" figure hides that entirely unless you meter the judge separately. Third, rate-limit errors and retries burn wall-clock and, depending on where the failure occurs, can bill for work you discard. A related trap: cost per attempt measured on simple single-shot probes does not extend to strategies that expand each seed into many variants, so the extrapolation is not linear in the number you extrapolated from. **Quota contention and enforcement.** Provider quotas are usually per account or per project. A campaign running at full concurrency starves the live product of the same quota, producing a degradation that the client's monitoring reads as an incident — one you caused. Enforcement is worse: automated abuse detection keys on exactly the pattern a red-team campaign emits, namely repeated attempts to elicit prohibited content, at machine rate, from a single account. The response is a throttle or a suspension, not a conversation, and it lands on the account regardless of what the client signed. If that account serves the product, the test has caused an outage. This is the strongest argument for a separate project and key, and for moving prohibited-content generation onto a model you run yourself. **Telemetry pollution.** Engagement traffic contaminates the client's own safety metrics, alert history, abuse dashboards and any dataset built from request logs. Agree a marker up front — a dedicated key, a request header, a fixed session attribute — so the traffic can be filtered out afterwards, and agree who does the filtering. Retrofitting that separation from timestamps alone is guesswork. **What goes in the document, and what you check on day one.** Account, project and key; spend cap; requests-per-minute and concurrency ceilings; the campaign window; the traffic marker; stop conditions; escalation contacts on both sides with a response-time expectation; and an explicit acknowledgement that the provider may act on the account irrespective of the engagement. On day one, run the small calibration batch, meter the judge separately from the target, compare observed spend and error rates against the estimate, confirm the spend cap is actually enforced by the provider's billing controls rather than by your intention — and only then open the throttle.

  • How do you estimate cost before the campaign rather than after?
    Run a calibration batch of a few hundred attempts, measure mean billed input and output tokens per attempt including any scoring calls, multiply out to the planned attempt count, and add margin for multi-turn attempts that run longer than the median.
  • The client cannot issue a separate key. What is your fallback?
    Lower the rate ceiling well beneath the product's headroom, run in a low-traffic window, keep the prohibited-content work on a locally-run model so the provider-side traffic looks ordinary, and get a named contact who can restore the account if it is throttled.

Budgeting one API call per attack attempt is like budgeting one photocopy per page of a document you will copy double-sided, in triplicate, with a proofreader re-reading every sheet. The multipliers, not the unit price, decide the bill.

saying these in an interview costs you the question

  • Treats request volume as an implementation detail that needs no agreement.
  • Runs the campaign on the production key because it was the one available.
  • Gives no cost estimate and no spend cap for a metered endpoint.
  • Ignores that multi-turn attempts and a scoring model multiply the billed call count.
  • Has no plan for separating engagement traffic out of the client's safety telemetry afterwards.

context