Your custom PyRIT scorer reports 12 objective-achieved hits across a 400-turn run. What do you do before that number goes into the engagement report?
answer
- the count measures the scorer first
- read all 12 hits, sample the 388
- stratify the negatives, do not sample uniformly
- report miss rate next to hit count
- calibrate on this run's transcripts
basics
~20 sHand-label a sample and compare. Sample from the 388 the scorer called non-hits, not only the 12 hits, because misses are where uncounted findings sit. Read the hits too. Then report the count together with the measured miss rate and false-alarm rate on that sample, so a reader knows what the 12 excludes.
solid answer
~1 minThe 12 is a property of the scorer, not of the target, until you have measured the scorer. Two labelling passes, both against transcripts from this run: **The hits.** Read all 12. Every one that does not survive a human read is a false alarm, and the surviving count is what actually goes in the report. Twelve is small enough to read exhaustively, which is the cheap half. **The non-hits.** This is the half people skip and the half that matters. Draw a stratified sample from the 388 — not uniform, because most of them are obvious refusals. Over-sample the long responses, the ones without refusal language, the ones in an unexpected format or language, the ones in a modality your scorer may not parse. Label those by hand. From the two passes you get the two numbers that make the 12 interpretable: of the responses a human called a hit, what share the scorer caught, and of the scorer's hits, what share survived. Report both alongside the count, plus the sample sizes. If agreement on the negatives is poor, the honest move is to fix the scorer and re-score the stored transcripts before the number ships, not to caveat it in a footnote.
go deeper
Says the number should be checked by reading transcripts by hand before it is trusted.
Splits the check into hits and non-hits and explains that only the non-hits reveal misses, since triage already covers the hits.
Stratifies the negative sample on leak-correlated features, writes the success criterion down first, uses a second labeller on an overlap, and reports both error rates with sample sizes.
Makes 'no hit count ships without its calibration sample and scorer version' a standing rule, and budgets labelling time into the engagement plan rather than finding it at the end.
### What the 12 currently measures Twelve is a property of your scorer, not of the target, until the scorer has been measured. Stated honestly, the run so far says: across 400 scored turns, this bespoke rule fired twelve times. Every step from there to "the target is vulnerable in these twelve ways" is an inference the calibration pass has to support. ### Two labelling passes, both on this run's transcripts **The hits — read all twelve.** Exhaustive and cheap: at a few minutes each, under an hour. Every hit that does not survive a human read is a false alarm; the survivors are what goes in the report. This half is the one triage would have done anyway. **The non-hits — sample the 388.** This is the half people skip and the only half that speaks to coverage. Do not draw uniformly. Most of the 388 are plain refusals, so a uniform sample spends the labelling budget confirming easy cases and rarely reaches the responses where a bespoke scorer actually leaks. Stratify on features that correlate with leakage: - response length above the median for non-hits - absence of the refusal phrasing the rule keys on - non-text modalities, or text in a language other than the one the scorer was written against - truncated, errored or empty responses — an exception string scored as a non-hit is a hole, not a refusal - final turns of conversations that ended on budget exhaustion, where a late success hides behind the stop reason ### What the calibration costs Budget it into the engagement plan rather than discovering it at the end. Sixty transcripts at three to five minutes each is three to five hours of senior time, plus an overlapping slice for a second labeller. That is real engagement time — and it is the cheapest option available, against either shipping an unvalidated count or re-running 400 turns of live traffic to get a second opinion. ### Labelling discipline Write the success criterion down before you start reading, because "objective achieved" drifts once real transcripts are in front of you: the tenth borderline case gets judged against a standard the first one never faced. Have a second person label an overlapping slice. If two humans disagree materially, the scorer is not the first problem — the criterion is underspecified, and no scorer built on it can be right in a stable way. ### Where the number misleads **The denominator.** "12 in 400" invites a reader to hear a 3 percent attack success rate, but 400 is *turns*. If those twelve hits came from three of fifty conversations, the per-conversation figure is 6 percent; if the engagement carried four objectives and all twelve hits sit under one of them, the per-objective story is different again. Say which unit you divided by, every time you print a rate. **Accuracy as a summary.** On a corpus that is mostly genuine refusals, a scorer that never fires still posts high accuracy. Report the two error rates apart: the share of human-labelled hits the scorer caught, and the share of its own hits that survived a read. **Calibrating on the wrong corpus.** Tuning the scorer against curated demo transcripts, then meeting a real target that complies in a summary, in a code fence, or in another language. Calibrate on the artefacts this run produced, and re-calibrate whenever the target, the objective wording or the scorer changes. ### What the report line should contain A reader should be able to reconstruct the claim: the triaged hit count, the size of both labelled samples, the share of hand-labelled hits the scorer caught, the share of its hits that survived, the unit of the denominator, and the scorer version that produced the verdicts. "The scorer missed roughly a third of the successes in a labelled sample of 48 non-hits" is a number a reader can weigh. A bare 12 is not. And if agreement on the negatives comes back poor, the honest move is to fix the scorer and re-score the stored transcripts before the number ships. A footnote under a figure you know to be wrong is not calibration; it is disclosure of a defect you had both the evidence and the time to repair.
- You have budget to hand-label 60 transcripts. How do you split them between the 12 hits and the 388 non-hits?Read all 12 hits, since that is exhaustive and cheap, and spend the remaining 48 on a stratified draw from the non-hits weighted toward long, non-refusal, unusual-format and unparsed-modality responses.
- Your labelled sample says the scorer missed a large share of successes. What do you do with the original 12?Treat it as provisional, fix the scorer, re-score the stored transcripts, and report the new count with the corrected scorer identified. A caveat under a known-wrong number is not a substitute for re-scoring evidence you already have.
Twelve hits in 400 turns is a defect rate quoted per bolt when the client asked about the machines. The same twelve findings read as 3 percent of turns, 6 percent of conversations, or most of one objective, depending only on which denominator you print next to them.
saying these in an interview costs you the question
- Labelling only the hits and calling the scorer validated.
- Sampling the non-hits uniformly and concluding the scorer is fine after reading forty refusals.
- Shipping the raw count with no sample sizes and no scorer version.
- Calibrating against curated demo transcripts instead of the run's own output.
- Treating a poor agreement result as a footnote rather than as a reason to fix and re-score.