skip to content

After the product team edits the system prompt, you rerun your red-team suite against the same sampling LLM endpoint. One attack prompt now succeeds on 3 of 20 attempts where the filed run recorded 0 of 20. How do you establish whether this is a real regression before you file it?

level: seniorimportance: should knowfreq 42%

answer

  1. 0/20 is not zero — rule of three
  2. overlapping intervals, not a regression
  3. re-decide old transcripts, judge drift
  4. paired old-vs-new, not versus a stale number
  5. severity can outrank significance

basics

~20 s

Treat both numbers as estimates. Zero of twenty never meant zero; its upper confidence bound is well over ten percent, so 3 of 20 is barely separable. Raise the trial count, pin decoding, re-decide the old transcripts with the current judge, and if possible run the old configuration alongside the new one.

solid answer

~50 s

Three things could produce this, and you have to separate them. **Sampling variance.** Both figures are estimates from 20 draws. A 0/20 observation is consistent with a true rate up to roughly 14 percent at 95 percent confidence, and 3/20 is 15 percent with a wide interval of its own. The intervals overlap heavily, so this evidence alone cannot claim a change. Pin decoding where the API allows and raise trials — hundreds, not tens. **Judge drift.** If the component deciding whether an attempt counted has changed, or is itself a model, the rate moves with the target untouched. Re-decide the archived transcripts of the old run with the current judge. **A real change.** Only then does the prompt edit become the explanation, and the honest test is running the old and new prompts back to back rather than against a months-old number. Then decide by severity, not only significance, and report the interval.

go deeper

for a junior

Says to rerun it more times before believing a difference from so few attempts.

for a middle

Computes or cites an interval on both rates, notes the overlap, and pins decoding parameters before rerunning with more trials.

for a senior

Separates sampling variance, judge drift and genuine change; designs a paired old-versus-new run to cancel provider drift; states cost; files on severity with an honest interval.

for a principal

Sets the programme rule — minimum trials per severity class, the interval reporting format, and who funds high-trial confirmation runs — so this argument is not relitigated per finding.

The mistake this question is built to catch is treating a red-team success count as a measurement rather than as a sample. A sampling endpoint returns a distribution over outputs; twenty attempts are twenty draws from that distribution, and a count of zero is the least informative observation available. ### Quantify before you argue For 0 successes in 20 trials, the one-sided 95% upper bound on the true rate is about 14% — the rule of three, 3/n. That is the number that belonged in the original report: not "the attack fails", but "the rate is probably below about one in seven, and we cannot see below that with twenty draws". Against that bound, 3 of 20 is unremarkable. Its own interval is wide too — a point estimate of 15% with a 95% interval running from roughly 5% to 36% — and it overlaps the old bound heavily. So the two observations are consistent with no change at all, and they are equally consistent with a real one. The evidence in hand does not decide it. Sample size follows from what you want to detect. Distinguishing a 2% rate from a 10% rate with any confidence needs trials in the hundreds per prompt, not tens. That is a budget sentence, not a statistics sentence: 400 trials on the one suspicious prompt is 400 requests — trivial — while 400 trials across a 60-prompt suite is 24,000 requests plus a judging call each, which is a real spend and a real queue against the endpoint's rate limit. Confirmation runs should therefore be narrow and deep, not wide and deep. ### The multiplicity trap If you reran a whole suite, ask how many prompts you compared. Across 60 prompts at 20 trials, several low-rate behaviours crossing from 0 to 2 or 3 successes is the *expected* outcome of chance, not a signal. A finding selected because it moved is a finding selected on noise, which is why a narrow high-trial confirmation on the specific prompt is the required next step rather than a nice-to-have. ### Control what can be controlled Pin `temperature` and any `seed` the endpoint honours, and record whether it actually was. Do not assume `temperature: 0` buys determinism: batching, mixture-of-experts routing, non-associative floating-point reduction across varying batch shapes and silent serving-stack updates all reintroduce variation. Plan for repetition, not reproducibility. ### Run the comparison, not the memory The strongest design is paired. If the previous system prompt can still be deployed to a staging copy of the application, run old and new configurations interleaved in the same window — same decoding, same prompt set, same judge, same hour — and compare those. Interleaving cancels provider-side drift, which is the whole point: comparing today's run against a number filed three months ago silently attributes every intervening model repoint, guardrail bump and routing change to the prompt edit. That attribution error is the specific way this number misleads, and it is nearly always in the direction of blaming whatever change you happened to know about. ### Check the instrument, always Confirm the attack prompt set hash is byte-identical, the harness version is unchanged, and the deciding component is the same version. Then re-decide the archived transcripts of the original run under the current judge as a control. If the old, unchanged transcripts now score differently, the judge moved and the target may not have. ### Then use judgement Statistical caution is not a licence to sit on a serious result. If those three successes are a severe harm class, file now — described honestly as an observed 3 of 20 with a wide interval, on a configuration that changed, with a higher-trial confirmation run pending and its cost stated. Severity can outrank significance; that is a decision the programme should have made in advance, per severity class, so it is not relitigated per finding. What you must not file is "regressed from 0% to 15%". That sentence overstates both endpoints, and the first person to rerun it will get a third number.

  • Roughly what does an observation of 0 successes in 20 attempts tell you about the true success rate?
    Only that it is probably below about 14 percent — the rule-of-three upper bound, 3 divided by the number of trials. It does not establish that the attack fails.
  • Why is running the old and new system prompts interleaved in one window better than comparing to the filed number?
    Because it cancels everything that drifted in between — a model swap, a guard bump, provider-side routing. Comparing to a stale number attributes all of that to the prompt edit.

Twenty attempts is a keyhole. Seeing nothing through it does not tell you the room is empty — anything rarer than roughly one in seven could have been standing there the whole time.

saying these in an interview costs you the question

  • Reporting the change as zero percent to fifteen percent with no interval and no trial-count discussion.
  • Assuming a temperature setting of zero makes a hosted endpoint deterministic.
  • Attributing the difference to the system-prompt edit without checking whether the model or guard also changed in the intervening months.
  • Dismissing three successes on a severe harm class purely because the result is not statistically significant.

context