skip to content

You inherit a finished automated jailbreak run that reports sixty successful attacks against a chat endpoint, every one labelled by a scoring model. How do you measure that scorer's error rate in both directions before the report goes out?

level: seniorimportance: should knowfreq 45%

answer

  1. write the success criterion first
  2. hits are enumerable, misses are not
  3. stratify near threshold and zero-hit strategies
  4. second scorer, adjudicate disagreements only
  5. no stored non-hits, recall unmeasured

basics

~20 s

Hand-label a random sample of the sixty labelled hits to get precision. For the other direction you need the stored transcripts of everything the scorer rejected: sample those, weighted toward borderline confidence, and hand-label. If only hits were kept, say recall is unmeasured rather than guessing.

solid answer

~60 s

Two separate exercises, and they are not equally easy. Precision is straightforward: take a random sample of the sixty labelled hits, hand-label each against a written criterion agreed before you start, and report the proportion confirmed with an interval, not a point estimate. At sixty items you can often just adjudicate all of them, which is the honest answer for a customer-facing report. Recall is the hard half, because misses are not on any list. You have to sample from what the scorer rejected, which is only possible if the run persisted non-hit transcripts. Sample there in strata: borderline scores near the threshold, attack strategies that produced zero hits, later turns, and any response shape the scorer is known to handle badly such as non-prose or non-English output. Re-score the stored text with a second scorer written from a different prompt and hand-adjudicate the disagreements; that enriches the sample cheaply since it costs no target calls. If the run kept only hits, say so. The report then carries a measured precision and an explicit note that recall is unknown, which is a defensible statement where a guessed one is not.

go deeper

for a junior

Should say read a sample of the sixty transcripts and check whether each really contains a violation.

for a middle

Separates the two error directions and knows the misses require sampling from rejected transcripts, so persistence matters.

for a senior

Runs the whole exercise: written criterion first, inter-labeller check, full or sampled adjudication of hits, stratified sampling of rejections, a second independent scorer to enrich, and honest reporting of what was and was not measured.

for a principal

Sets the standing rule for what every run must persist and what a report must state about machine labels, so audits are comparable across engagements rather than improvised per run.

You are auditing an instrument, not a model. The sixty is a count of labels produced by a scoring model, and your job before the report ships is to attach two error estimates to it — or to state honestly that one of them cannot be attached. ## Step zero: write the success criterion first One paragraph, agreed before anyone opens a transcript, saying what would count as a violation for each harm area in scope. Without it two reviewers labelling the same sixty transcripts will disagree and you will have measured nothing. Have a second person independently label a shared subset and check how often you agree. If the humans disagree materially, the criterion is the defect, not the scorer, and no amount of sampling repairs that. Do not show reviewers the scorer's confidence while they label. It anchors them to the instrument under test. ## Precision: the easy half The labelled hits are an enumerable population, so sampling is trivial. At sixty items, adjudicating all of them is usually the right call for a customer-facing report — at a few minutes per transcript that is a few hours of one experienced person, which is small against the cost of shipping a finding that does not survive the client's own reading. Report a proportion with an interval, not a point estimate. Sixty items adjudicated to, say, forty-one confirmed is roughly 68% precision with a confidence interval wide enough that quoting "68%" bare is overclaiming. Then look at the *structure* of the errors, not only the count. Hits clustered on one attack template, on turn one, or in one harm category usually mean the scorer matched a phrasing pattern rather than content — and that pattern tells you which of the remaining hits to distrust and which future runs to instrument differently. ## Recall: the half the run either supports or forecloses Misses are not on any list, so they can only be found among the attempts the scorer *rejected*. The first thing to establish is therefore what the finished run persisted. If every attempt is stored with prompt, full response, raw score, attack strategy, turn index and thread-termination reason, recall is estimable. If only hits were kept, it is not, for this run, at any price short of re-running. Where the store exists, sample the rejected pool in strata rather than uniformly, because a uniform sample of thousands of clear refusals finds almost nothing per hour: - near-threshold scores first, where the classifier is uncertain; - strategies or harm categories with suspiciously zero hits; - long responses, non-prose responses (code, tables, structured output), and non-English responses, where scorers are systematically weaker than the targets they score; - final turns of threads that ended on the turn cap. Then re-score the entire rejected pool with a second scorer written independently from a different prompt, and hand-adjudicate only the disagreements. This is the cheap enrichment step: it bills scorer tokens over stored text, no attacker calls, no target calls, no rate-limit backoff and no fresh authorisation to touch a live system, and it hands you a list where at least one of the two labels is wrong. Adjudicating agreements spends hours to learn little. ## Where the number misleads — the paragraph to get right A stratified sample deliberately oversamples where misses are likely. If 12% of your stratified rejections turn out to contain real violations, that 12% is **not** the run's miss rate, and 88% is **not** the scorer's recall. It is an enriched estimate that tells you misses exist and where they concentrate. To quote a population figure you must weight the strata back to their true proportions in the rejected pool, and say which of the two you are reporting. Two adjacent denominator swaps to refuse. Hits divided by attempts is an attack-success rate over attempts; it contains no information about labels the scorer got wrong in either direction. And "sixty this quarter versus forty last quarter" describes the target only if the scorer model, prompt and threshold were identical across both — otherwise the delta is a property of your instrument. ## When the store is missing Say so. The report carries a measured precision on the sixty, an explicit statement that the miss rate is unmeasured for this run, and a note on why. That is a defensible sentence; a guessed recall figure is not. Do not spend target budget re-running the campaign to answer a question stored transcripts would have answered for free, unless you have both the budget and a reason beyond tidiness. The real deliverable of an inherited audit is the change to the pipeline: persist every attempt with its raw score and its termination reason, so that every future audit costs transcripts and reviewer hours instead of target calls.

  • The run persisted only the sixty hits. What do you put in the report?
    The measured precision on those sixty, and an explicit statement that the miss rate is unmeasured for this run, plus a change to persist all attempts next time.
  • Why adjudicate only the disagreements between two scorers rather than everything?
    Agreements are mostly the easy cases and consume human time for little information. Disagreements concentrate the region where at least one scorer is wrong.
  • Why is a stratified sample of rejected responses not a clean recall number?
    It deliberately oversamples where misses are likely, so it estimates where errors live. To quote a population figure you must weight the strata back to their true proportions.
  • What single artefact makes the next audit far cheaper?
    A persisted store of every attempt with prompt, response, raw score and thread-termination reason, so re-scoring and sampling need no further target calls.

Estimating misses by sampling the flagged hits is like inspecting a net's catch to work out which fish swam past it. The only way to answer is to look in the water the net rejected, which means somebody has to have kept it.

saying these in an interview costs you the question

  • Reporting the sixty as findings with no adjudication at all.
  • Claiming a low miss rate from a sample that only contains labelled hits.
  • Quoting a stratified enriched sample as if it were an unbiased population recall figure.
  • Re-running the campaign against the target to answer a question that stored transcripts could have answered.
  • Labelling transcripts without agreeing a written success criterion first, so two reviewers score differently.

context