skip to content

An automated red-team run against a hosted assistant finishes, but roughly 30% of its attempts ended in transport errors, rate-limit rejections or empty responses rather than reaching a scored verdict. How should the coverage statement in the report treat those attempts?

level: seniorimportance: should knowfreq 44%

answer

  1. attempted / scored / hit — three numbers
  2. errors land on the safe side of the fraction
  3. losses are not a random sample
  4. empty body: refusal or failure?
  5. throttle = coverage loss and possibly a control

basics

~20 s

Unscored is not passed. Attempts that ended in transport errors, rate-limit rejections or empty replies never reached a verdict, so they belong in neither the hit nor the clean column. Report attempted, scored and hit as three separate numbers, and re-run the lost attempts before drawing conclusions.

solid answer

~50 s

Most harnesses compute their headline rate as hits over attempts, so an errored attempt silently lands on the safe side of the fraction. Thirty percent errors therefore does not mean a slightly noisier result; it means your success rate has an inflated denominator and your clean areas may be clean only because nothing arrived. Two things to do. First, split the reporting into **attempted / scored / hit**, and compute every rate over *scored*, with the unscored count printed beside it. Second, ask whether the losses are random. They usually are not — long payloads time out more often, multi-turn attacks lose more sessions to a mid-conversation error, and a moderation layer that returns an empty body looks identical to a network failure. The attempts you lost skew towards the aggressive end, which is precisely the end you were testing. Then re-run: lower concurrency, add backoff, and re-drive the lost attempts before you write a conclusion about those areas.

go deeper

for a junior

Notices that errored attempts are not evidence of safety and says they should be re-run.

for a middle

Separates attempted from scored, computes rates over scored, and prints the unscored count next to every rate.

for a senior

Classifies the losses by cause, argues that they are biased towards the strongest attempts, and marks affected clean results provisional; distinguishes an empty-body refusal from a transport failure using raw logs.

for a principal

Sets a threshold above which a run is not reportable, and requires unscored-attempt accounting in the house template so no engagement ships a rate whose denominator is unexamined.

An unscored attempt is dangerous because it is *quiet*. A failed attack and a failed HTTP call both land in the results as a row with no hit, and every harness that computes hits over attempts puts both on the safe side of the fraction. Thirty percent unscored is therefore not a noisier result; it is a rate whose denominator contains 1,800 events that were never judged, and clean areas that may be clean only because nothing arrived. **Where the number comes from, mechanically** `garak` writes one record per attempt into `*.report.jsonl` and copies only detector hits into `*.hitlog.jsonl`, so its per-probe score is hits divided by attempts generated — an attempt whose generator raised is still an attempt that was generated. `promptfoo` distinguishes an `error` result from a `fail`, which is the right shape, but any summary line that reads “N passed of M” has already made a decision about which bucket the errors went to. The first thing to establish about your own harness is which of those two behaviours you have; the answer is in the per-attempt records, not the summary. **Classify the losses before interpreting them** Transport timeouts, connection resets, provider rate-limit rejections (HTTP 429), quota exhaustion, harness parse errors on a malformed response, a policy refusal returned as an empty body, and a multi-turn session that died mid-conversation are all different events wearing one symptom. Only some are yours. An empty body that was really a moderation refusal is a *result* — the target declined — and filing it as an error understates the control; a genuine timeout filed as a refusal overstates it. If the harness does not separate them, the separation has to come from status codes, latencies and any moderation signalling in the raw logs. **The sample you lost is not random** Error rate correlates with exactly the attributes that make an attempt strong: length, encodings that inflate token counts, tool-heavy flows, and long multi-turn campaigns that have more turns in which to die. The lost 30% is weighted towards the attempts most likely to have produced a finding, so the surviving remainder is not a smaller sample of the same distribution — it is a biased one, biased in the direction that flatters the target. **What it costs to fix** Cheap, and the arithmetic is worth doing out loud. Re-driving 1,800 lost attempts against a metered endpoint at roughly a cent of tokens each is tens of dollars; the binding constraint is wall clock, because the fix is to *slow down*. At a sustainable two requests per second with backoff, 1,800 attempts is about fifteen minutes of machine time plus an hour of engineer time to lower concurrency, add retry with jitter, and re-run the lost subset. Compare that with the cost of a wrong clean verdict on an area the client then ships. There is almost no engagement in which re-driving is not the cheaper branch — which is why “we ran out of time” is rarely the real reason it does not happen; nobody looked at the unscored count. **Rate limiting is two claims, not one** The throttle cost you coverage, and it may also be a control a real attacker faces. Those are independent statements needing independent support. “We could not complete X because of provider throttling” is a coverage loss and belongs in the declaration. “The throttle limits an attacker to N attempts per minute” is a finding, and it is only true after you have established whether the limit is per key, per IP or per account, and how cheaply it parallelises. Never let the second sentence retire the first. **What goes in the report** ```text Area attempted scored hits unscored (cause) Direct injection 2,400 2,290 18 110 (429 x86, timeout x24) Indirect via retrieval 1,800 1,020 0 780 (session drop x612, parse x168) PROVISIONAL Tool misuse 1,800 1,740 6 60 (timeout) ``` Three numbers per area, unscored broken out by cause, and an explicit provisional marker anywhere the unscored share is material. Then note which attack families lost the most, because that is the sentence that tells a reader which way the bias runs. **What I check** Does the harness fold errors into passes? Is the printed rate computed over scored rather than attempted? Does the unscored share differ by attack family — and if so, is the family that lost most also the aggressive one? Were empty bodies separated from transport failures using status codes rather than assumption? And has every area with a material unscored share either been re-driven or labelled provisional in the text a reader will actually see?

  • The target returns an empty response body for some attempts. Is that an error or a result?
    It depends what produced it, and you cannot tell from the empty body alone. Check status codes, latency and any moderation signalling in the raw logs: a policy refusal is a result, a truncated or reset connection is a coverage loss.
  • Provider throttling stopped you completing a third of the plan. Does that make throttling a mitigation you can report?
    Only as a separate, separately tested claim. Coverage loss is yours; whether the throttle constrains a real attacker depends on whether it is per-key, per-IP or per-account and how cheaply it is parallelised.

It is an exit poll where a third of the people walked off before answering — and the ones who walked off were disproportionately the ones with strong opinions. Counting them as "no opinion" gives you a clean-looking result from a sample that lost precisely the voices you were trying to hear.

saying these in an interview costs you the question

  • Reporting a success rate over total attempts when a large share never got a verdict
  • Treating errors as a random sample of the test set
  • Reading a rate-limit rejection as evidence the attack failed
  • Concluding an area is clean without re-driving the attempts that were lost in it

context