Your team's standing PyRIT suite runs every seed prompt through a matrix of prompt-converter variants. How do you decide how many variants the suite carries, and what does a per-converter success number actually measure when you publish it?
answer
- basis of transform classes, not accretion
- unique-cause promotion, rotation for the rest
- denominator you invented
- dedupe by behaviour, not by variant
- matrix drifts toward the judge's blind spots
basics
~20 sSize the matrix by distinct transform classes, not by count: variants multiply target and scorer calls linearly and mostly re-test the same input path. A per-converter number measures which surface form got past the input handling and the judge — not distinct model weaknesses, and not attack-surface coverage.
solid answer
~50 sEvery variant you add multiplies the whole run — target calls, scorer calls, triage time — while usually probing the same input-handling path with a cosmetic difference. So the suite should carry a **basis**: one representative per genuinely distinct class of transform, chosen so that dropping any one loses information. The promotion rule is empirical: a variant earns its place if it has ever been the sole cause of a verdict. One that never has goes into a rotating pool rather than every run. On reporting: 'twelve of twelve converters run' is a denominator you invented — not the share of attack surface reached, not the share of known techniques covered. And when several variants trip on the same underlying behaviour, that is **one** report item with several reproductions; otherwise the matrix inflates the report in proportion to how many variants you happened to implement.
go deeper
Understand that more variants cost more calls and more time, and that they mostly re-test the same thing.
Explain that a per-converter number reflects the input path and the judge as well as the model, so it is not a model-weakness ranking.
Bring promotion and retirement rules for variants and insist on deduplicating hits by underlying behaviour before anything is reported.
Own the whole frame: a documented basis, an explicit denominator in every published number, report units defined by behaviour, and an active guard against the suite optimising toward the judge's blind spots.
Three separate judgments hide in this question, and a principal-level answer keeps them apart: how big the matrix should be, what a per-converter number is a statement about, and what unit leaves the team in a report. ### 1. Sizing — cost is linear, triage is the binding constraint Every variant multiplies the run. Twelve variants over 200 seeds is 2,400 target calls plus 2,400 scorer calls per pass, before any variant that is itself model-backed adds its own call per prompt. Run that nightly and the compute is annoying; the real cost is human. Someone has to look at the hits, and triage time scales with hits, not with configuration lines. That is the scarce resource, and a matrix grown by accretion — every interesting transform anyone read about, kept forever — spends it on re-confirming the same input path in twelve cosmetic dresses. Replace accretion with an explicit **basis**: one representative per genuinely distinct class of transform, documented with the reason it is in, chosen so that dropping any single member loses information you would otherwise not have. Pair it with rules rather than opinions: - **Promotion:** a variant earns a permanent slot only if it has been the *sole* cause of at least one verdict within a defined window. - **Retirement:** a variant that has not been a sole cause in that window, or whose hits are overwhelmingly judged artefacts, moves to a rotating pool rather than being deleted. Novelty still enters the suite; not every run pays for all of it. ### 2. What a per-converter number actually measures It is a statement about a **pipeline**, not about a model. Read in full: *this surface form was accepted by the input path, reached the model, produced a reply, and that reply read as compliance to our judge.* It bundles input handling, the model, and the judge's competence on that specific form. Three consequences you should say aloud when publishing it: - It is not a ranking of model weaknesses. Comparing two converters' numbers compares their effect on the whole pipeline, judge included. - It has to be published with its denominator spelled out — prompts attempted, how many were judged, and by what. - It is not "coverage". In this tree that word means at least three different things: converters run out of converters implemented, attack surface reached, or behaviours exercised. Only the first is what a converter matrix can give you, and it is a denominator your own team invented. ### 3. The report unit is not the tool's unit The tool's unit is a hit per prompt per variant. The report's unit must be a **distinct behaviour**, with reproductions attached. Deduplicate by behaviour before anything leaves the team; otherwise a wider matrix mechanically produces a scarier report about an entirely unchanged system, and the first person to notice that is the customer. Ten variants tripping the same behaviour is one item with ten reproductions — and the reproductions are genuinely useful, because a behaviour reachable by ten distinct forms is harder to fix with a filter rule than one reachable by a single form. That, not the count, is the insight to sell. ### The strategic risk to name A matrix that is grown by keeping whatever produces hits will drift toward transforms your own judge mis-reads, because those produce hits. From the dashboard this looks like a productive, improving suite; in reality it is an instrument tracking its own blind spots, and the quarter-over-quarter rise is the drift, not the target. Guard against it explicitly: re-label a hand sample each cycle, publish the judge's false-positive rate **per converter** next to the success numbers, and treat a variant whose hits are mostly artefacts as a scorer defect with a named owner rather than a red-team win. ### What I would check before publishing the quarterly number Did the target change at all in the window — model version, system prompt, guard configuration? If not, an increase must be explained by the suite, and the two candidate explanations are a wider matrix and a drifting judge. Recompute the number against the previous quarter's basis as well as the current one, so the trend is not silently a denominator swap. And confirm every published figure carries the sentence that says what its denominator is; a number whose denominator is unstated is the one that will be quoted back at you.
- Ten variants all trip on the same model behaviour. How many findings is that?One, with ten reproductions. The tool's unit is a hit; the report's unit is a distinct behaviour, and conflating them inflates the report in proportion to matrix size.
- How do you decide a variant should be retired?It has not been the unique cause of any verdict over a defined window, or its hits are overwhelmingly judged artefacts. Move it to a rotating pool rather than deleting the capability.
Reporting 'twelve of twelve converters run' as coverage is like saying you tested a building's locks because you tried every key on your own keyring. The number describes your keyring, not the building.
saying these in an interview costs you the question
- Presenting 'N of N converters run' as coverage of the attack surface.
- Growing the matrix indefinitely with no retirement criterion.
- Counting each variant's hit on the same underlying behaviour as a separate finding.
- Never measuring the judge's false-positive rate per converter.
- Treating a rising hit count as progress without checking whether the target changed at all.