You are sizing a red-team probe run against a hosted content-moderation service that bills per request and rate-limits your API key. How does that shape the corpus and the run design, and which outcomes must never be scored as a clean result?
answer
- two meters: bill and rate limit
- explore shallow, confirm with repeats
- throttle or timeout = no verdict
- store raw bodies, re-threshold free
- name the coverage denominator
basics
~20 sEvery probe costs money and a slot in the rate limit, so build a small stratified corpus instead of a huge random one, deduplicate identical payloads and cache responses. Never score a rate-limit rejection, timeout or error as clean: those are missing verdicts, and counting them as not-blocked inflates your apparent bypass rate.
solid answer
~50 sTwo meters run at once: the bill, and the key's request allowance. Naive concurrency turns a run into retries — wall-clock time stops tracking work done and the retries are billed too, so self-throttle below the documented allowance rather than discovering it with throttle rejections. Design the run in **two phases**. Exploration: one call per payload family, broad and shallow, to find which families are even interesting. Confirmation: repeat only the interesting ones several times, because a hosted service can return different values for the same input and a single call cannot tell a bypass from noise. The scoring rule matters more than the sizing. A transport error, a throttle rejection, a truncated response and a timeout are all **no verdict**. Bucket them separately and re-drive them; folding them into "not blocked" is how a run reports a bypass rate that is really an availability problem. Store the full response body so re-analysis never costs another call.
go deeper
Knows the calls are billed and rate-limited, and that errors should not be counted as results.
Sizes the corpus against an explicit call budget, self-throttles below the key's allowance, repeats confirmed cases, and stores raw responses.
Budgets the re-drives, control set and re-verification pass too, and designs the run so a reviewer's later question never requires re-buying the data.
Decides what the engagement buys with a finite metered budget across this and other layers, and sets the team's reporting convention for denominators and no-verdict handling.
**The two meters.** A probe run against a hosted moderation service is throttled by two independent things at once, and teams usually plan for only one. The first is money: services bill per request or per unit of text submitted, so the corpus size is a line item. The second is the key's request allowance — a per-minute or per-second ceiling above which the endpoint returns a throttling status instead of a verdict. Exceeding the second does not merely slow you down; it converts work into retries, and retried calls are still calls. **Write the multiplication out before the run.** The call count is not `payloads`. It is: ``` families x variants-per-family x repeats-per-variant + control-set calls (each session) + re-drives of every no-verdict outcome + the end-of-engagement re-verification pass ``` The first term is the one everyone budgets. The last three routinely double it. Decide up front which term you cut when the budget does not stretch — normally **variants** first, because breadth across harm families is what makes a coverage statement possible; **rarely repeats**, because they are what separates a bypass from noise; and **never the control set**, because without it a clean run is uninterpretable. **Two phases.** Exploration runs one call per payload family, broad and shallow, to find which families are even worth spending on. Confirmation revisits only the survivors, several calls each. A single clean response is a weak observation about a networked service whose internals you cannot pin; a handful of identical calls tells you whether the value is stable, and stability is what makes a finding survive the customer trying to reproduce it. **The scoring rule that matters more than the sizing.** A transport error, a throttling rejection, a truncated body and a timeout are all **no verdict**. None of them carries any information about the content. Fold them into "not blocked" and your reported bypass rate is partly an availability metric wearing a security label — the run will show a bypass surge exactly when the service was busiest, which is the opposite of a finding. Bucket them separately, back off, and re-drive them. Dropping them silently is the subtler version of the same error: it shrinks the denominator without announcing it, so every rate in the report quietly changes and nobody can tell by how much. **Concurrency in practice.** Cap in-flight requests yourself below the documented allowance rather than discovering the ceiling by collecting throttle rejections. Honour any retry-after the service returns. Use exponential backoff **with jitter**, because a pool of parallel workers that all back off by the same amount resynchronises into another burst and rediscovers the limit together. A useful health signal mid-run: the share of calls that are retries. If it is climbing, you are paying twice for the same probe and your wall-clock time has stopped tracking work done — reduce concurrency, which will finish the run sooner as well as cheaper. **Store raw, decide later.** Persist request, full response body with its per-category values, and a timestamp, per call. The round trip is the metered, rate-limited, irreversible part; everything downstream is free. A change of mind about the cut-off, a reviewer asking for a different stratification, a customer disputing one case — all answered offline against the stored corpus for zero additional spend. **Say which denominator you mean.** "Coverage" on a run like this can mean probes sent out of the corpus you built, harm families represented in that corpus, or categories the service exposes that you exercised at all. Those give three different percentages from the same run, and the largest is usually the least meaningful. A bare coverage figure with an unnamed denominator is the fastest way to have an otherwise sound run dismissed in review — state the denominator in the sentence that carries the number, not in a footnote.
- Your run's response times climb and the failure count rises mid-run. What do you change?Reduce in-flight concurrency, honour retry-after, add jitter to the backoff, and re-drive the affected payloads. Then check that none of those failures were scored as results.
- The budget only covers half the corpus. What do you cut?Variants within families first, keeping breadth across families, and keep both the repeats on confirmed-interesting cases and the control set. Then state the reduced denominator in the report.
- How do you decide how many repeats a confirmed-interesting payload needs?Enough to see whether the response is stable at all. A handful usually separates a consistently clean payload from one that flips; anything more should be justified against the per-call cost.
The payload count is the ticket price; the control set, the repeats and the re-drives are the booking fees that are only visible at checkout. Budget as if the fees are half the bill again, because they usually are.
saying these in an interview costs you the question
- Counting throttled or errored calls as "not blocked", which inflates the reported bypass rate.
- Firing maximum concurrency and treating throttle rejections as normal, paying for retries and reporting wall-clock as throughput.
- One call per payload with no repeats, so nondeterminism is indistinguishable from a bypass.
- Persisting only pass/fail, forcing a second billed run for any re-analysis.
- Quoting a coverage percentage without saying what the denominator is.