A developer runs the reproduction from your AI red-team finding once, gets a polite refusal, and closes the ticket as unreproducible — but the behaviour is real and intermittent. What should the evidence package have contained to prevent that outcome?
answer
- one attempt is a sample
- count over trials, with denominator
- settings stated, not assumed
- loop script ships with the ticket
- zero in N is a bound, not a fix
basics
~20 sState up front that the target is sampled, so one attempt proves nothing. Ship the observed count out of total attempts, the sampling settings you used, a script that repeats the request and tallies how many responses met the criterion, and a written verification rule such as: any hit in twenty attempts is still a failure.
solid answer
~60 sThe package failed because it presented a probabilistic behaviour in a deterministic format. A single transcript is an existence proof, not a reproduction procedure. What should have shipped: - **The observed frequency with its denominator** — "7 of 20 attempts" — so the reader expects misses. - **The generation settings used**, stated rather than assumed: a reader re-running at a different temperature or system prompt is not running your test. - **A repeat harness that is not your harness**: a short loop that issues the request N times and prints how many responses met the criterion. That loop is the reproduction; a single call is a sample from it. - **An explicit verification rule**: at least one qualifying response in N attempts means unfixed; zero in N means the rate is below a stated bound, which is not the same as impossible. - **What you hit and when**, so a provider-side change can be told apart from a fix. The rule matters most. Without it, "I tried it and it refused" reads as disproof rather than as one sample.
go deeper
Understands that the model does not answer identically each time, so one refusal does not disprove the finding.
Ships the rate with its denominator, the settings, and a repeat loop, and writes the acceptance rule into the ticket.
Also handles the reverse error — treating zero hits in a small sample as proof of a fix — and separates a provider-side change from a real remediation.
Sets the organisation's standing convention for how probabilistic findings are reported, verified and closed, so teams are not renegotiating it per ticket.
Sampled generation breaks the contract that ordinary defect reporting rests on. In classic bug work, steps-to-reproduce either work or they do not, so one failed re-run is a real signal. Here the same input yields a refusal and a compliance from the same system minutes apart, and both parties can honestly report opposite results. ## Why the same request gives different answers The decoder draws each token from a probability distribution shaped by temperature and by nucleus (top-p) or top-k truncation; two identical requests take different paths through it. Setting temperature to zero narrows this but does not remove it on a hosted endpoint: batching, expert routing in mixture-of-experts models, and floating-point non-associativity across differing hardware all perturb logits slightly, and a near-tie between two tokens flips. Most hosted chat endpoints also expose no usable seed, so you cannot pin the draw. The practical consequence is that "same input, same output" is simply not a guarantee you can hand over, and a package that implies it is misrepresents the target. ## Turn the reproduction into a measurement The unit of evidence is not a transcript, it is a count over trials. Package it that way and the developer's re-run returns a number they can compare with yours rather than a yes/no they can disagree with. Concretely: the observed count with its denominator ("7 of 20"); the decoding settings and system prompt stated rather than assumed, since a reader re-running at a different temperature is running a different experiment; a short repeat loop that is not your harness — plain requests, fixed trial count, tallied against the written criterion; and an explicit acceptance rule agreed in advance, such as *twenty attempts at the stated settings; any qualifying response means the finding stands*. Without that rule in writing, "I tried it and it refused" reads as disproof. ## What it costs Verification costs N model calls per check, multiplied by the turn count when the attack is a conversation: twenty trials against a six-turn attack is a hundred and twenty calls. On a hosted endpoint that is small money but real wall-clock once a per-minute limit throttles the loop, and it recurs on every re-verification and every pipeline run if the loop becomes a regression test. This is exactly why N is negotiated up front rather than left to "run it a lot" — an unstated N invites the developer to stop at the first miss, and invites you to keep going until you get a hit. ## Where the number misleads Both directions of this go wrong, and the second is the costlier one. Reading the rate as precise: 7 of 20 is 35%, but at that sample the interval spans roughly one in five to more than one in two. A later run showing 5 of 20 is not evidence of improvement; the two are indistinguishable. Quoting 35% upward as "one in three users hits this" also swaps the denominator — attempts by an operator deliberately steering are not user sessions. Reading zero as fixed: this is the common overclaim at closure. With zero events in N trials, the rough 95% upper bound on the true rate is about 3/N — the rule of three — so twenty clean attempts bound the rate only below roughly 15%. A behaviour firing on one request in ten can plausibly show zero in twenty. Say so in the package before the developer says the opposite on your behalf: zero in N means *below a stated bound at these settings*, not impossible, and it says nothing at all about the variant cases that were never re-run. The third misleading zero is provider drift — a hosted alias reversioning gives you a clean run that is a model change, not a remediation, which is why the deployment identifier and timestamps belong at both ends. ## What to check before you send Run your own loop from the ticket text alone, on a clean shell, with the harness uninstalled; if it does not tally, the developer's will not either. Confirm the criterion is mechanically checkable — greppable, or a short predicate — so the tally is not a human reading twenty transcripts. Confirm the settings block is complete, including the system prompt. And write the closure rule into the ticket, including what a zero run does and does not establish. ``` for i in $(seq 1 20); do send_request >> out.jsonl # same body, same settings, no red-team tooling done grep -c "<criterion match>" out.jsonl # compare with the reported 7/20 ```
- The developer re-runs twenty times and sees zero hits. Is it fixed?Not proven. Zero in twenty only bounds the rate loosely; you can say the behaviour is now below a stated frequency at those settings, and you should re-check with more trials or a different variant before closing.
- Why does the package have to state the sampling settings even when they are the application defaults?Because defaults change and staging often differs from production. Written settings make the developer's run the same experiment as yours rather than a similar-looking one.
- The endpoint gives no seed. Does that make the finding unreportable?No — it makes the reproduction a repeated trial instead of a single call. Report the rate with its denominator and say explicitly that no seed was available to pin.
Re-running a sampled model once to check a finding is like flipping a coin once to decide whether it is biased. The answer you get is a sample, not a verdict.
saying these in an interview costs you the question
- Reporting an intermittent behaviour as a single transcript with no attempt count.
- Accepting one clean re-run as proof the issue is fixed.
- Leaving sampling settings unstated so the re-run is a different experiment.
- Quoting a rate without saying how many attempts it came from.
- Blaming the developer for closing the ticket rather than fixing the package.