skip to content

You reran the same 20-attempt red-team script against the same hosted chat endpoint the next day: the first run succeeded twice, the rerun six times. How should the finding's severity handle a measured success rate that moved from 10% to 30%?

level: middleimportance: should knowfreq 45%

answer

  1. 20 attempts, huge interval
  2. pool: 8 of 40
  3. noise or target changed?
  4. quote a range, not a point
  5. 4x samples to halve the width

basics

~20 s

Do not rerate on it. At 20 attempts both results sit inside the same wide confidence interval, so the swing is what sampling noise looks like, not evidence the target got worse. Report both runs with their counts, widen the band you quote, and if the rate really drives the score, collect a much larger sample.

solid answer

~50 s

Two things could produce that swing, and the finding should say which you ruled out. **Sampling noise, the default explanation.** Twenty attempts gives an interval roughly spanning single digits to the high thirties around either estimate. The two runs are statistically indistinguishable; treating 30% as "three times worse" is reading noise as signal. **A real change in the target.** Hosted endpoints get reversioned and guardrails get rolled out without notice, so confirm the model identifier, the decoding settings, the system prompt and the wrapper were identical on both days before you call it noise. For the rating: quote a range, not a point, and cite the pooled counts (8 of 40). If your rubric has a band that 10% and 30% straddle, that is a rubric problem — either the bands are too fine for your sample sizes, or the run is too short to resolve them. Never quietly file the higher number because it makes the finding look better.

code

python · 12 lines
python
from math import sqrt

def wilson(successes, n, z=1.96):
    p = successes / n
    d = 1 + z * z / n
    centre = (p + z * z / (2 * n)) / d
    half = z * sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return centre - half, centre + half

for s, n in ((2, 20), (6, 20), (8, 40)):
    lo, hi = wilson(s, n)
    print(f"{s}/{n}: {s / n:.0%}  interval {lo:.0%}-{hi:.0%}")

go deeper

for a junior

Should recognise that 20 attempts is a small sample and that the two numbers may not really differ.

for a middle

Pools the runs, quotes an interval, and knows to check whether the endpoint or its configuration changed between the runs.

for a senior

Decides how much sample the rating actually justifies, states the limitation in the report, and rates on impact and retry cost when the sample cannot resolve the band.

for a principal

Sets rubric bands coarse enough to survive realistic sample sizes and requires counts plus an interval on every filed rate.

The failure mode here is treating a proportion measured on a tiny sample as a stable property of the system. Before arguing about what changed, do the arithmetic, because the arithmetic usually ends the argument. ### The numbers first A binomial confidence interval — the Wilson form is the standard one for small samples — gives: | run | observed | 95% interval | |---|---|---| | Monday | 2 of 20 (10%) | roughly 3% – 30% | | Tuesday | 6 of 20 (30%) | roughly 15% – 52% | | pooled | 8 of 40 (20%) | roughly 10% – 35% | The two intervals overlap heavily, and a two-proportion test on 2/20 versus 6/20 is nowhere near significance. Nothing in these counts *requires* the target to have changed. That sentence, with the numbers behind it, is what belongs in the finding. ### But do not stop at "it is just noise" Three things can produce the swing, and only the first is free to assume. 1. **Sampling noise** — the default, given the interval widths above. 2. **The target moved.** Hosted endpoints are reversioned, guardrails ship, a system prompt is edited by the product team, traffic is routed to a differently-quantised deployment or a different region. None of that is announced. 3. **The measuring instrument moved.** This is the one people forget. If the judge is an LLM grader running at non-zero temperature, or its rubric prompt was edited, or the detector was upgraded between runs, part of the variance sits in the scorer rather than in the model. The check is cheap and decisive: re-score Monday's stored transcripts with Tuesday's judge. If Monday's success count changes, the instrument moved and the two rates were never comparable. If the model identifier, decoding settings, template, classifier configuration or judge version differ between the runs, then the two runs measured two systems and pooling them is simply wrong. ### What resolving the question would cost Interval width shrinks with the square root of n, so roughly four times the sample halves the width. Separating a true 10% from a true 30% with confidence takes low hundreds of attempts per arm, not tens; separating 15% from 20% takes thousands. Against a raw model API, hundreds of attempts is minutes of wall clock and small money — but the judge cost scales with it, and human validation of several hundred transcripts is most of a day. Against a *production* path the cost is different in kind: throttles turn hundreds of attempts into hours, several seeded accounts may be needed, and a run of thousands is likely to trip abuse detection and consume the engagement's goodwill. So the usual operator decision is not to buy the precision at all. "Sample too small to distinguish 10% from 30%; rated on impact and retry cost" is an honest, defensible line in a report, and it is better than a number that implies resolution the run never had. ### Where the number misleads - **Filing the higher run** because it supports the severity you already wanted. This is how a rate gets talked upward, and it is invisible unless both runs are on the record. - **Declaring a regression.** "The endpoint got measurably weaker overnight" is a strong claim built on 40 samples; it needs version and configuration evidence, not counts. - **Averaging into a bare point estimate.** 20% is the right centre, but printing it without an interval reasserts precision the sample does not support. - **Rubric bands finer than the sample can resolve.** If 10% and 30% straddle a band boundary, the score is being decided by noise, and every triage meeting will re-litigate it. That is a rubric defect, not a measurement defect. ### What to check, and what to file Compare the model identifier or version string, decoding settings (temperature, top-p, and a seed if the endpoint honours one), the system prompt and template, any input or output classifier, the deployment region, and the judge version. If all match, write the exploitability line as "8/40 across two runs on those dates, 95% interval roughly 10–35%", score impact from a single success as always, and note the limitation explicitly. If any differ, report the two runs separately with what differed, and do not pool them.

  • How many attempts would you need to distinguish a 10% rate from a 30% rate with confidence?
    Low hundreds per arm, not tens. Interval width shrinks with the square root of n, so roughly four times the sample halves the width — budget for that before you let the rate drive the score.
  • What would convince you the target actually changed rather than the sample?
    A version or model identifier that differs between runs, a changed system prompt or wrapper, a new classifier in front, or a swing far outside the pooled interval on comparable sample sizes.

Twenty attempts is a bathroom scale that reads to the nearest ten kilos. The two readings differ, but by less than the instrument's own coarseness, so nothing has yet been learned about the weight.

saying these in an interview costs you the question

  • Declaring the endpoint "got weaker overnight" from 2/20 versus 6/20
  • Filing the higher of the two runs because it supports a higher severity
  • Averaging the runs into a point estimate with no interval
  • Never checking whether the model version, decoding settings or wrapper changed between runs

context