A red-team scan sends each attack case to the endpoint ten times because responses vary, and three of one case's ten attempts were flagged as hits. When you collapse those ten attempts into one report item, what must the collapsed record keep, and what breaks if you keep only the first hit?
answer
- trials always collapse, variants do not
- summary, not selection
- keep numerator and denominator
- attach a refusal as well as a hit
- record what decided the hit
basics
~20 sKeep both numbers: three hits out of ten attempts, plus the exemplars and the non-hit responses. Keeping only the first hit destroys the denominator, so nobody downstream can say how often the failure happens, and a re-run after a fix has no baseline to compare against. One hit and ten of ten read identically.
solid answer
~60 sRepeated trials are the one axis where collapsing is unambiguously correct — ten sends of the same case against a stochastic endpoint are not ten defects. But the collapse must be a **summary**, not a **selection**. The record needs: attempts made, attempts flagged, which object flagged them, at least one hit exemplar with its full exchange, and at least one near-miss where the endpoint declined. The near-miss matters more than people expect: it is the evidence the behaviour is intermittent rather than absent, and it is what stops a developer who re-ran once, saw a refusal, and closed the item as unreproducible. What breaks with first-hit-only is the denominator. Downstream, a rate cannot be computed, a re-run cannot be compared to a baseline, and an item that fired three times in ten looks the same as one that fired ten in ten. It also inflates confidence: a single attached transcript reads as deterministic, and the fix gets validated by one clean run. How the rate then feeds a severity rating is a separate judgment; deduplication's job is only to make sure the numbers survive the collapse.
go deeper
Should say the ten sends are one finding, not three, and that the count of hits ought to be written down.
Should keep both numerator and denominator and explain that a rate cannot be recovered once the denominator is dropped.
Should also attach a non-hit exemplar, record what object decided a hit, and explain why re-run comparisons are meaningless without published attempt counts.
Sets the trial count as a run-level standard so items are comparable across runs, and challenges any report whose intermittent items ship a single transcript.
### Two kinds of duplication, opposite treatment One scan produces two duplications and they are not the same problem. **Variant duplication** — many differently-worded attempts at the same underlying weakness — is a genuine judgment call, and the grouping key is where that judgment lives. **Trial duplication** — the same case re-sent because the endpoint samples its output rather than answering deterministically — is not a judgment call at all. Ten identical sends are never ten defects, so they always collapse. The mistake is never *whether* to collapse; it is *how*. The collapse must be a **summary** of what happened across the ten, not a **selection** of the one that looked best. ### What the collapsed record has to carry ```text case_id, attempts=10, flagged=3, decided_by=<the object that marked a hit: detector, classifier or judge model>, hit_exemplars=[full request/response exchange, ...], non_hit_exemplar=<full exchange where the endpoint declined>, decoding_settings_if_you_set_them ``` Every element is there for a reason a reviewer can name. **Both numbers, not one.** Three is a numerator and it is meaningless alone. The denominator is what makes it a rate, and a rate is what everything downstream compares against. **What decided a hit.** A detector with a false-positive habit turns "3 of 10" into a statement about the detector rather than about the model. Naming the deciding object is what lets a reader discount it, and what lets a second analyst re-judge the same transcripts with a stricter check. **A non-hit transcript, not only a hit.** This one is workflow rather than completeness. Intermittent findings are the ones most often bounced back as "cannot reproduce": a developer re-runs the case once by hand, sees a refusal, and closes the item. An attached refusal transcript pre-empts that entire round trip by showing on the face of the item that a refusal is an expected outcome of the same case. **Decoding settings if you controlled them.** Temperature and sampling parameters move the rate. A rate reported without them cannot be compared against a re-run that used different ones. ### What it costs Trials are the cheapest thing in the run to buy and the most expensive to buy carelessly. Ten sends instead of one multiplies the endpoint bill and the wall clock of the whole scan by ten, and against a metered chat endpoint that is the difference between a run you start and forget and a run you have to budget. It also multiplies the rows an analyst reads. Ten is a common default because it is enough to notice that something is intermittent; it is nowhere near enough to state a rate precisely, and it costs almost nothing to say so in the report. Storing the members — all ten exchanges, not the three flagged — is the other cost, and it is small: text transcripts, kept for the life of the engagement, are what make every later question answerable without re-running anything. ### Where the number misleads This is the paragraph that matters. - **Dropping the denominator makes intermittent and total failure identical.** An item that fired three times in ten and an item that fired ten times in ten read exactly the same when both ship with one transcript and no counts. One is a control that mostly works; the other is a control that never works. Nobody downstream can tell them apart, and severity gets set by the transcript's tone rather than by frequency. - **A post-fix zero is not evidence unless its denominator is stated.** Three-in-ten falling to zero-in-ten is weak: if the true rate after the fix were still one in five, a ten-send re-test misses it more than a tenth of the time. Three-in-ten falling to zero-in-two-hundred is strong. The two look identical in a report that writes "re-tested, no longer reproduces", and that sentence is the most common way a fix gets validated by a run too small to detect the failure. - **Three of ten is a small-sample number and reads as a precise one.** Written as "30 percent" it acquires a confidence the sample does not support; the honest form keeps the raw fraction visible so a reader sees the sample size. - **Trial collapsing quietly absorbing variants.** If the ten sends differed in wording at all, they were not ten trials of one case. Folding them together buries the variant axis inside a number that looks like sampling noise, and the report loses the fact that several distinct framings worked. How that rate then feeds a severity rating is a separate judgment. Deduplication's only job here is to make sure the numbers survive the collapse intact. ### What you check Before publishing an intermittent item: does it carry both numbers; does it name what decided a hit; does it attach at least one refusal alongside the hit; and were the sends genuinely identical? Before accepting a fix: what attempt count was the re-test run at, and is that count large enough to have detected the original rate?
- Why attach a non-hit transcript to an intermittent item?It shows the reader that a refusal is a normal outcome of the same case, which stops the item being closed as unreproducible after one clean manual re-run.
- The ten sends were not identical — the wording differed slightly. Does anything change?Yes. Those are variants, not trials. Folding them together hides the variant axis behind a number that looks like sampling noise; group them by the variant key first and count trials within each.
Reporting three hits without saying there were ten attempts is like reporting that a test batch had three defective units without saying whether the batch was ten units or ten thousand. The numerator alone cannot tell a reader whether the process is broken or merely imperfect.
saying these in an interview costs you the question
- Reporting the item with one transcript and no counts.
- Calling three hits in ten attempts three findings.
- Treating differently-worded sends as trials of the same case.
- Declaring a fix successful from one clean re-run with no attempt count.
- Not recording which object decided that an attempt was a hit.