skip to content

Ongoing Assurance

The first report is the cheap part; the programme decides what blocks a release, which model changes force a rerun, and which numbers justify the budget. Interviewers ask because that is what they are buying.

on this pageshow

explore

questions

14

An AI red-team programme reports a quarterly "findings" number obtained by summing every hit its automated LLM scanners emitted. Why is that number not the number of findings in the report, and what unit should be counted instead?

level: juniorimportance: must knowfreq 62%

answer

  1. hit = one attempt, not one issue
  2. probe x seed x sample x rerun
  3. dedup to behaviour + cause + owner
  4. count moves with the catalogue
  5. publish attempts and confirmed separately

basics

~20 s

A scanner hit is one prompt attempt its detector marked as failed. Many hits are the same behaviour repeated across prompt variants, probe classes and reruns, and some are detector false positives. The report's unit is a triaged, deduplicated issue with a cause. Count those, and report attempts separately as volume.

solid answer

~50 s

The tool's unit and the report's unit are different objects, and conflating them is the most common way a programme metric goes wrong in its first quarter. An automated LLM scanner emits one record per attempt that its detector judged a failure. A single underlying weakness — say, one weak refusal boundary — produces a hit for every paraphrase in the probe's seed set, for every temperature sample, and again on every scheduled rerun. Multiply that by however many probe classes touch the same behaviour and the count moves with the catalogue, not with the system. The reportable unit is a triaged item: deduplicated to one behaviour, attached to a cause and an owner, and confirmed by a human against the raw transcript because detectors misjudge in both directions. So publish two numbers with distinct names: attempts (volume, a cost and coverage signal) and confirmed issues (the risk signal). Never let one label carry both.

go deeper

for a junior

Says a hit is one attempt, that the same problem repeats across variants and reruns, and that the report should count deduplicated confirmed issues.

for a middle

Adds the cross-product that generates the hits, that detectors have false positives in both directions, and publishes attempts and confirmed issues as separately named numbers.

for a senior

Adds precision of the automated stage as its own metric, explains why the count moves with catalogue configuration, and refuses quarter-on-quarter comparisons across a config change.

for a principal

Frames it as a measurement-validity problem: any headline number whose error rate nobody has sampled cannot be defended when it is challenged in a review.

**What a "hit" is, mechanically.** An automated LLM scanner has three moving parts, and the headline number falls out of the third. A *probe* is the scanner's unit of test generation: a class that emits seed prompts for one behaviour it wants to elicit. A *target wrapper* sends each prompt to the system under test and captures the reply. A *detector* — the scanner's automated judge, which may be a string or regex rule, a small trained classifier, or another model asked to grade the reply — reads the response and returns pass or fail. A hit is one (prompt, sample, response) triple that the detector marked fail. The scanner writes one line per attempt to its run log and one line per failure to its hit log, and the number that gets quoted upward is, almost always, the hit log's line count. **The arithmetic that manufactures the number.** Hits are the product of a cross-product and a rate: `hits ~= probe classes enabled x seed prompts per probe x samples per prompt x reruns x observed fail rate` Concretely: 20 probe classes at 25 seed prompts each, sent 5 times apiece, is 2,500 attempts in one run. Swap the curated subset for the scanner's whole catalogue and the first factor triples. Raise the scanner's samples-per-prompt setting from 5 to 20 and everything downstream is multiplied by four. Add a second scheduled run in the quarter and it doubles again. None of those edits touched the system under test. Nothing in the pipeline deduplicates by *behaviour* either, because the scanner compares strings and detector verdicts: it has no way to know that forty failures spread across four probe classes all trace to one permissive clause in one system prompt. **What the run costs.** A mid-sized sweep is 2,500-10,000 calls to the target; if the detector is itself a model, roughly double that in grader calls. On a small hosted model that is single-digit to low-tens of dollars; with a long system prompt, retrieval context or an agent loop it is an order of magnitude more, and wall-clock runs 20-90 minutes at the concurrency a rate limit allows. The dominant cost is human and it is set by the hit count: 400 hits read at about two minutes of transcript each is roughly thirteen engineer-hours of triage, every cycle. That is the line item the metric quietly commits the team to. **Where the number misleads.** Three separate failures, and they do not cancel. - *It is configuration-elastic.* It rises with catalogue breadth, samples per prompt and rerun frequency. A quarter-on-quarter comparison across any of those edits is measuring the instrument, and it will read as the product getting worse. - *The target is stochastic.* The same prompt fails on one sample and passes on the next, so a decoding-temperature change or a silent provider model update moves the count with no weakness created or removed. - *The detector errs in both directions.* Refusal-phrase detectors mark a hedged, harmless non-answer as a failure, and mark a fluent, on-topic harmful completion that never used a refusal phrase as a pass. So the hit count is neither an upper bound (false positives inflate it) nor a lower bound (false negatives deflate it) on the number of real issues. A scanner that also reports a relative score — your model against a bag of calibration models — is describing a peer distribution, not your risk appetite; being average among peers is not a control. **What to publish instead.** Three separately named numbers, never merged into one word. *Attempts run* — volume and spend, which explains the query-budget line. *Confirmed issues* — human-triaged, deduplicated to one behaviour with one cause and one owner. *Precision of the automated stage* — confirmed issues over hits triaged, which is what tells you whether the detector is worth its triage hours. Stamp all three with a configuration fingerprint: probe list, samples per prompt, target model version, detector version. **What to check before believing any of it.** Draw a stratified sample of about thirty hits and thirty passes and read the raw transcripts. The hits give you an observed precision; the passes are the only place a false negative can be caught, and skipping them is how programmes discover a year later that their detector never fired on the failure mode that mattered. Until that sample exists, the count has an unknown error rate, and a number with an unknown error rate cannot be defended the first time someone in a funding review pushes back on it.

  • Two hits came from different probe classes but the same weak refusal boundary. One issue or two?
    One, if the fix is the same fix. Deduplicate by cause and owner, and note that two probe classes reached it — that is useful evidence the behaviour is easy to reach, and it belongs in the item, not as a second item.
  • How would you report a behaviour that reproduces on only three of ten samples of the same prompt?
    As one issue with a reproduction rate attached. Probabilistic reachability is part of the finding, not a reason to split it or to drop it; the rate is what the fix has to move.

saying these in an interview costs you the question

  • Treating the scanner's summary count as the finding count in the executive report.
  • Comparing this quarter's count to last quarter's after changing which probe classes were enabled.
  • Never reading raw transcripts, so the detector's false-positive and false-negative rates are unknown.
  • Deduplicating by prompt string rather than by underlying behaviour.

context

open as a page

Your AI red-team findings were filed against a hosted chat application whose endpoint URL and request contract never change. Which changes underneath that unchanged endpoint can silently invalidate those findings, and why does nothing alert you?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Three: the provider swaps the model behind the same endpoint alias, the safety classifier in front of it is upgraded, or the product team edits the system prompt. None changes the URL or the contract, so no deploy alert or failing test fires. Your filed pass and fail results simply stop describing the live system.

open as a page

If an AI red-team programme's headline number is confirmed findings per quarter, which part of its automated testing gets more investment over time, and why is that the wrong part?

level: middleimportance: must knowfreq 55%

basics

~20 s

Whatever produces findings cheapest wins the budget. That is usually a high-yield, low-severity probe family against an easy surface. The hard work — a slow multi-turn attack on a tool-using agent — yields few items and looks unproductive. The programme drifts toward whichever instrument is noisiest, not toward where the risk is.

open as a page

What must you record about an AI red-team run so that repeating it months later is a comparison against the filed result rather than a fresh experiment?

level: middleimportance: must knowfreq 58%

basics

~20 s

Record the whole configuration, not the numbers: the concrete model identifier and any version metadata returned, decoding parameters and seeds, a hash of the system prompt and context, the guardrail versions and thresholds, the exact attack prompts and harness version, the deciding component's version, trials per prompt, and the raw transcripts.

open as a page

A red-team engagement on a customer-facing chat assistant closes with a list of confirmed findings, and your release gate offers three dispositions: block the release, ship behind a compensating control, or ship with a named acceptance. How do you decide which disposition each finding gets?

level: middleimportance: must knowfreq 68%

basics

~20 s

Decide by finding class and consequence, not by a tool's score. Classes whose harm is irreversible or legally exposed block. A finding ships behind a compensating control only if a control already deployed demonstrably stops it. Everything else ships with a named owner accepting it in writing, with a review date.

open as a page

Nothing on a typical AI red-team dashboard says what was never tested. How would you put the untested part on the sheet, and why must any "coverage" percentage you publish state its denominator?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Publish an explicit not-tested list beside every result: surfaces, languages, modalities and attack classes nobody ran, with a reason for each. Coverage needs a denominator because the word means at least three things — probes run out of those available, attack surface reached, and behaviours exercised from a benchmark. A number without one is unreadable.

open as a page

A red-team report on an AI feature marks a finding as 'accepted, signed off' at the release gate. What does that disposition actually mean, and who has to be the signer for it to carry weight?

level: juniorimportance: should knowfreq 46%

basics

~20 s

It means the product ships with the finding still present and a named person has taken responsibility for that choice. The signer must own the product's risk and have authority to stop the launch, not the red teamer who found it or the engineer who would fix it. It needs a date and a reopening trigger.

open as a page

Nobody tells you when a model provider swaps the model served behind an unchanged hosted endpoint. What can a red-team programme actually do to detect that a swap has happened?

level: middleimportance: should knowfreq 50%

basics

~20 s

Ask for a pinned dated model identifier instead of a floating alias, and record whatever version field the response carries, diffing it per run. Where neither exists, run a small fixed canary prompt set on a schedule and watch its behavioural fingerprint: refusal wording, output formatting, length and latency. Drift there means rerun.

open as a page

An AI red-team programme's dashboard leads with mean time to remediate, but the fixes are made by application teams and by a model provider the programme does not control. What breaks about that metric, and what would you put on the dashboard instead?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It scores teams you do not manage, so it measures their backlog, not your work. Worse, many AI findings have no fix you own: the behaviour lives in a hosted model or a third-party guardrail. Report what you control — time to triage, time to notify, and the age and acceptance status of open items.

open as a page

After the product team edits the system prompt, you rerun your red-team suite against the same sampling LLM endpoint. One attack prompt now succeeds on 3 of 20 attempts where the filed run recorded 0 of 20. How do you establish whether this is a real regression before you file it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Treat both numbers as estimates. Zero of twenty never meant zero; its upper confidence bound is well over ten percent, so 3 of 20 is barely separable. Raise the trial count, pin decoding, re-decide the old transcripts with the current judge, and if possible run the old configuration alongside the new one.

open as a page

At a release gate, a product team asks to ship a confirmed red-team finding behind an output classifier as a compensating control rather than blocking. What evidence do you require before accepting that control, and what goes into the gate record?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Require the original finding retested end to end with the control enabled in the exact shipping configuration, not the control's own published coverage. Ask what happens when the control is unavailable, degraded or times out, and who notices. Record the control, the retest evidence, residual risk, an owner and an expiry date.

open as a page

You get three numbers on the slide that decides whether an AI red-team programme is funded next year. Which three do you choose, and how do you defend each against the objection that the team producing it can move it?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pick numbers the team cannot inflate alone. One: issues whose cause was new to the risk register, severity-banded. Two: named untested surfaces with reason codes, especially budget-blocked ones. Three: count and maximum age of open items with no owner and no acceptance. Defend each by naming who verifies it outside the team.

open as a page

The team behind an LLM product edits its system prompt most weeks, the safety-classifier service upgrades on its own schedule, and the provider can swap the model behind the endpoint at any time. A full red-team suite rerun costs real metered spend and analyst time. How would you design the assurance cadence?

level: principalimportance: should knowfreq 33%

basics

~20 s

Make the trigger mechanical rather than calendar-based: subscribe to the config and deploy pipeline, and tier the suite. A tiny high-severity tier runs on every prompt or guard change, a mid tier per release train, the full suite on a model swap or quarterly. Any tier hit escalates to the next.

open as a page

You own the release gate for AI features across several product teams. Every additional finding class the gate blocks on buys real release delay, and no class ever reaches zero occurrences. How do you decide how many classes the gate blocks on?

level: principalimportance: should knowfreq 40%

basics

~20 s

Block only on classes where one instance is unacceptable and the harm cannot be undone. Size that set to what fix teams can actually clear inside a release cycle, or the gate gets routed around. Everything else takes a control or a named acceptance. Revisit the set from what escaped, not from opinion.

open as a page