skip to content

You are red-teaming a chat endpoint that sits behind a safety classifier — a guard that is itself a model, such as Llama Guard or ShieldGemma, which reads text and returns a verdict. Why does each test case in that run cost more than testing the chat endpoint alone, and how should that shape the size of your test set?

level: juniorimportance: should knowfreq 55%

answer

  1. guard is a model = extra inference
  2. input screen + output screen
  3. samples multiply cost
  4. time the guard separately
  5. no guard calls = wiring bug

basics

~20 s

Because the guard is itself a model, every case runs extra inferences: one to screen the input, usually another to screen the reply, on top of the chat call. Each case therefore costs roughly two to three times the tokens, money and latency. Budget for that and run a smaller, deliberately chosen set.

solid answer

~50 s

A pattern or deny-list filter is effectively free to exercise: matching costs microseconds and no tokens. A guard that is a model is a second (and, if you also screen the response, a third) inference on every case, so cost and wall-clock scale with the number of prompts, not with their complexity. Practical consequences: - **Size the set to the budget, then order it.** With a fixed number of calls you get, put the harm categories you actually care about first, not the largest generated pack you can find. - **Repeats multiply.** If the chat model is nondeterministic and you want three samples per prompt, you pay the guard three times too. - **Latency is a separate finding.** Time the guard call apart from the chat call; an inline model guard can dominate p95 for the product, and that is worth reporting even when its verdicts are good. - **Self-hosted is not free either** — it is GPU time and throughput contention rather than a per-token bill.

go deeper

for a junior

Says the guard is another model call, so each test case costs extra tokens and time, and the test set has to be smaller.

for a middle

Adds the multiplier from output screening and repeated samples, and separates guard latency from model latency in the report.

for a senior

Plans the run against a call budget and rate limits, caches and deduplicates during debugging, and verifies the guard was actually invoked before trusting any result.

for a principal

Decides whether the inline guard's latency and spend are acceptable for the product at all, and what the team should measure continuously versus in a one-off engagement.

## What this layer actually is, mechanically A safety classifier is not a deny-list. It is a model. The self-hosted ones — Llama Guard, ShieldGemma, Granite Guardian, WildGuard, Prompt Guard — are small language models (roughly 1B to 8B parameters) that are handed a prompt template listing a fixed set of hazard categories together with the text under review, and asked to emit a short verdict: a safe/unsafe token, usually followed by the codes of the categories it believes apply. The hosted ones — the OpenAI moderation endpoint, Azure AI Content Safety, Perspective API — do the same thing behind an HTTP call and return per-category scores. Everything about the economics of testing one follows from that single fact: producing a verdict is an *inference*, so it consumes tokens or GPU time, it is billed or queued, and it occupies wall-clock time inside the request path. Contrast the alternative your instinct is calibrated on. A regex or deny-list guard matches in microseconds, costs nothing per call, and can be exercised a million times over lunch. Against a pattern filter, test-set size is limited only by how many strings you can think of. Against a model guard, it is limited by budget. ## Write the arithmetic before you write prompts For a set of N distinct prompts, K samples per prompt, with input screening and output screening both enabled: ``` chat inferences = N x K guard inferences = N x K (input screen) + N x K (output screen) total ~ 3 x N x K ``` So 200 prompts at 3 samples is 600 chat calls and about 1,200 guard calls — 1,800 inferences, not 600. Add a multi-turn probe where the guard rescreens the growing conversation on every turn and the guard term grows with turns as well as samples. Add a paraphrase cluster of eight wordings per harm and you have multiplied the whole thing by eight. ## What it costs, concretely On hosted infrastructure the cost is a bill plus a rate limit, and the rate limit usually binds first: two providers, two quotas, and the slower of them sets your wall-clock. On self-hosted infrastructure the guard is not free either — it is GPU residency and throughput contention. A small guard model serving short classifications will do order tens of requests per second on one modern accelerator, which sounds ample until you notice it is sharing that accelerator with the thing you are testing. The third and most underrated cost is engineer time: guard verdicts arrive as text that must be parsed, mapped to your labels and triaged, and triage does not scale the way generation does. ## Where the number misleads **The pack-size illusion.** A downloadable pack of 100,000 harmful prompts is quoted as coverage. At three inferences per case it is 300,000 inferences, which you will not run. What actually happens is that someone samples two per cent of it, and the coverage denominator silently becomes the sample while the report still names the pack. **"It is a small model, so it is cheap."** Cheap per call, expensive per run, and — the part that hurts — *serial*. Input screening sits ahead of the chat call, so it adds directly to time-to-first-token; it does not amortise away under load. A guard that is inexpensive on the invoice can still be the dominant term in the product's p95 latency, and that is a legitimate finding even when every verdict it returns is correct. **Retries and fail-open.** A rate-limited or timed-out guard call that the harness retries is paid for twice, and if the harness records both attempts it is counted twice in the results. Worse, many integrations fail open on timeout: the fastest, cheapest-looking cases in your log may be exactly the ones where the guard never ran, which quietly inflates the clean rate. ## What you would check Reconcile counted guard calls against expected calls — if you sent 600 samples with two-sided screening, you should see about 1,200 guard invocations, and a shortfall is a wiring or fail-open bug, not a security result. Log the guard's tokens and duration separately from the chat call's, so latency and spend can be attributed. Pull the error and retry counts and subtract retried duplicates before computing any rate. Confirm spend against the plan mid-run rather than at the end. And before the real measurement pass, turn off any verdict cache you used while debugging, so you are measuring the guard rather than your own dictionary.

  • You have budget for 500 guard calls. How do you spend them?
    Reserve some for known-positive controls that prove the guard fires at all, then split the rest across the harm categories in scope, with a few paraphrases per category rather than one prompt each.
  • Does caching guard verdicts risk hiding a real behaviour?
    Only if the guard is nondeterministic or the cache key ignores something the guard sees, such as prior turns. Key on the exact text handed to the guard, and disable the cache for the final measurement run.

Adding a model guard turns every test case into a three-course meal rather than one dish: the input screen, the model call and the output screen are each billed. A pack of 100,000 prompts is a 300,000-inference bill, so the pack size you quote as coverage is never the run you can actually afford.

saying these in an interview costs you the question

  • Assuming the guard is free because it is 'just a filter'.
  • Planning a six-figure prompt pack with no token or rate-limit arithmetic behind it.
  • Reporting one end-to-end latency number with no split between the model call and the guard call.
  • Never checking that the guard was actually invoked for every case.

context