skip to content

Your automated jailbreak loop uses the same model deployment as both the target under attack and the scoring model that decides which responses count as successes. As the engagement lead, what is your policy on that arrangement, and what do you require before a machine-labelled finding list is signed off?

level: principalimportance: nice to knowfreq 34%

answer

  1. correlated blind spots, not random misses
  2. clean where coverage was thinnest
  3. independence of assessment
  4. calibration set stated in the report
  5. unmeasured section: the miss rate

basics

~20 s

A scorer that shares the target's blind spots misses exactly the outputs the target should have refused, so the most important successes go unreported. Prefer an independent scorer, add deterministic detectors, measure the scorer against a fixed human-labelled set, and require human adjudication before any finding is signed off.

solid answer

~60 s

The failure is correlated, not random. If the scorer does not recognise a response as harmful, that is often the same reason the target did not refuse it. So the misses are not spread evenly across the run, they concentrate on the novel categories and framings you were hired to find, and the finding list looks reassuring precisely where it is least trustworthy. There is also a defensibility problem independent of the technical one: a report whose success count was produced by the same deployment being assessed is easy to challenge, and you will not win that argument with an error-rate table alone. My policy: use an independent scoring model where budget allows; where it does not, ensemble the model scorer with deterministic checks and treat any single-source label as provisional. Regardless of that choice, require three things before sign-off. A fixed human-labelled calibration set that the scorer is measured against for this engagement, with the numbers stated in the report. Human adjudication of every finding that reaches a customer. And an explicit statement of what was not measured, particularly the miss rate, so the reader knows the finding count is a floor.

go deeper

for a junior

Should recognise it is unwise to let a system grade itself and suggest a separate scorer.

for a middle

Explains the mechanism: shared training and safety behaviour means the scorer misses what the target failed to refuse, so misses are correlated with the interesting cases.

for a senior

Adds the operational response: ensemble with deterministic checks, calibration set, retain transcripts, and report machine-labelled versus human-confirmed counts separately.

for a principal

Sets policy and owns the cost: budget for an independent scorer, define sign-off requirements including an unmeasured section, and defend the instrument's independence in the client review.

There are two separate arguments against scoring a target with a model drawn from the same deployment, and a lead should keep them apart, because they need different answers and only one of them is technical. ## The technical argument: the errors are correlated with the findings A scoring model and a target from the same deployment share training data, share safety tuning and — the load-bearing part — share a notion of what counts as harmful. Where that notion has a hole, two things happen at once: the target answers instead of refusing, and the scorer reads the answer and sees nothing wrong. The miss is not drawn at random from the space of responses. It sits precisely on the frontier of the model's safety behaviour, which is the only part of the report anyone was paying for. The observable symptom in a finished run is a harm category that comes back clean with almost no borderline scores. Uncertainty is what a classifier produces near its decision boundary; a category with zero hits *and* no near-threshold scores usually means the scorer was confident, and confident in the same direction the target was. A hand read of a handful of rejected transcripts from that category is the cheap test, and when it turns up content the scorer called a refusal, the arrangement is disqualified for that engagement. Note what this does *not* predict: inflation. The shared-blind-spot mechanism under-reports. A team that expects a same-deployment scorer to flatter its own outputs into a higher success count has the direction backwards. ## The organisational argument: independence of assessment A number produced by the assessed system, about the assessed system, is contestable on its face. Even with a sound error analysis in hand, you will spend the client review defending the instrument instead of the finding, and an error-rate table will not win that argument, because the objection is not about the rate. ## The policy Prefer a scorer that does not share the target's provenance, and budget it as a line item rather than treating it as an optimisation to be dropped when the week gets tight. Where cost, contractual or data-handling constraints forbid it — and they sometimes genuinely do — do not describe the arrangement as neutral. Ensemble the model scorer with cheap deterministic checks that do not share its blind spots, retain every transcript for offline re-scoring, and mark single-source labels as provisional in the artefact itself, not in a footnote. ## Sign-off requirements, whichever way that decision went 1. **A maintained calibration set** of hand-labelled responses per harm area, including deliberate near-misses, that the scorer is run against *before* the campaign, with the results printed in the report. Without it, every claim about the scorer is an assertion. 2. **Human adjudication of every finding that leaves the team**, and a report that keeps machine-labelled and human-confirmed counts as two distinct numbers. 3. **An explicit unmeasured section**: the miss rate, the response shapes the scorer is known to handle badly, and the categories where no hits were found but coverage was thin. This is the requirement that addresses correlated blind spots directly, because it is the only one that speaks about what is absent. 4. **Retention of every attempt** with its raw score, so a later dispute is settled by re-scoring stored text rather than by re-running against a target that may have been updated since. Requirement 2 is the one people mistake for a fix. Adjudication operates on the findings that reached the queue; misses never reach it. Human review therefore raises precision and leaves the miss rate exactly where it was. ## What it costs, said out loud A second scoring pass over every stored attempt. A second vendor relationship, or a self-hosted scorer with its own serving cost, plus the data-handling review that comes with sending transcripts somewhere new. Storage for a full run rather than its hits. And human hours per engagement that scale with the finding list, not with the number of real findings. That is the price of a number the client can act on, and it belongs in the scoped engagement, not absorbed silently by the team on the last Friday. ## Where the number misleads The dangerous artefact from this arrangement is not an inflated count; it is a **clean category**. "No findings in category X" reads to a client as "X is safe", when what it may record is that the scorer and the target agree about X. The report must therefore separate three statements that a single zero currently blurs: how much was attempted in that category, how much of it the scorer rejected, and how much of the rejected material a human actually looked at. A zero with no coverage evidence behind it is a statement about your instrument wearing the clothes of a statement about the target.

  • What symptom in a finished run would make you suspect shared blind spots between scorer and target?
    Whole harm categories with zero hits and almost no borderline scores, especially categories where hand review of a few rejected transcripts turns up content the scorer called a refusal.
  • Budget allows only one model scorer. What do you add instead of independence?
    Cheap deterministic checks alongside it, a calibration set the scorer is measured against, all transcripts retained for re-scoring, and provisional marking of single-source labels.
  • Why does human adjudication not solve the correlated-miss problem on its own?
    Humans adjudicate the findings that reached them. Misses never reach the queue, so adjudication raises precision while leaving the miss rate untouched.

saying these in an interview costs you the question

  • Treating a same-deployment scorer as neutral because it was given a strict prompt.
  • Concluding a harm category is safe from a zero hit count with no coverage evidence.
  • Relying on human adjudication of findings and calling the miss rate handled.
  • Reporting a machine-labelled count and a human-confirmed count as if they were one number.
  • Absorbing the cost of independent scoring silently instead of scoping it into the engagement.

context