You own automated jailbreak campaigns across a dozen deployed targets, run by several operators, and you must publish one stopping rule based on distinct-template yield. What do you standardise, what do you deliberately leave to the operator, and what goes wrong if you get that split backwards?
answer
- standardise definitions, devolve economics
- a yield number that gates work becomes a target
- version the clustering rule on every record
- fleet saturating on schedule is a smell
- centralising buys comparability, costs correlated blind spots
basics
~20 sStandardise the definition of a distinct template, the scoring method, and the requirement that every stop cite a yield curve and the queries spent. Leave the threshold and the reseed decision to the operator, since targets differ. Get it backwards and teams tune the clustering until every campaign saturates exactly on schedule.
solid answer
~50 sStandardise the things that make campaigns **comparable**; devolve the things that depend on the target. **Standardise:** the clustering rule that defines a distinct template, the scoring method used to accept a success, the units on the curve (distinct templates against queries spent), the requirement that the rule is frozen before a run, and the record every stop must carry — curve, queries spent, last new template. Standardise the wording of the conclusion too, so no report claims a target saturated. **Devolve:** the no-new-template window, the query allowance, seed corpora, objectives, and whether a flat curve means stop or reseed. These track a target's traffic, cost per query and risk, and one global number is either too loose for the risky targets or wasteful everywhere else. Backwards is the failure: a mandated threshold over a locally chosen clustering rule hands every operator a dial that makes their campaign finish on time.
go deeper
Not expected to answer; at most notices that everyone should count a distinct template the same way.
Proposes a shared clustering rule and a shared stopping threshold, without yet separating definitional from economic settings.
Makes the split correctly and can explain why a floating clustering rule destroys comparability, with the record template that makes stops auditable.
Treats the yield number as an incentive, names the correlated-blind-spot cost of centralising definitions, and hedges it with audits, a rotated second scoring method and a deliberately off-standard control configuration.
## The number that gates work becomes a target This is a measurement-governance problem wearing a red-team hat. The moment distinct-template yield decides when a campaign stops, it stops being a passive observation and becomes something people steer towards — rarely through dishonesty, usually through a hundred individually defensible small choices about what counts as the same template. Whatever the operator can adjust to satisfy the gate will drift until it satisfies the gate. So the design question is not "what is the right threshold" but "which knobs may an operator hold". ## Definitional versus economic settings | kind | examples | why it falls this way | |---|---|---| | Definitional — standardise | the clustering rule that defines a distinct template; the scoring method that accepts a success; the curve's units; the freeze-before-run requirement; the record every stop must carry; the mandated wording of the conclusion | these fix what the number *means*. If they float, two campaigns' numbers are not the same quantity, and no cross-target comparison, trend line or funding decision built on them is valid | | Economic — devolve | the no-new-template window; the query allowance; seed corpora; objectives; whether a flat curve means stop or reseed | these encode how much a particular target is worth spending on, and that genuinely differs. A high-traffic consumer surface and an internal tool with ten users do not deserve the same window | Getting the split backwards is the classic failure: mandate one global threshold while letting each team define what a distinct template is, and you have handed every operator a dial that makes their campaign finish on schedule. The gate is fixed, the definition is adjustable, and the definition is what moves. ## What to write into the standard - **One clustering rule, versioned**, with the version stamped on every campaign record, so a later revision of the rule does not silently rewrite the fleet's history. - **One scoring method**, with a mandated periodic hand-audit of a sample of accepted successes. - **Curve units** — distinct templates against queries spent — and the requirement that rule and threshold are frozen before a query is spent. - **A record template**: queries spent and the billed-call multiplier behind them, the curve, clustering-rule version, seed corpus, objective, last new template and when it arrived, diagnosis, stop reason. - **Mandatory wording**: an exhausted search, with its configuration named. Never a clean target. - **More than one configuration per target per cycle**, because a single plateau is a single sample. ## What the standard itself costs State it, because an unfunded standard is a standard that gets skipped. The hand-audit is engineer time: a sample of accepted transcripts per campaign per cycle, read by someone competent to say whether the behaviour is actually present — the largest recurring line item, and the one teams drop first. Rotating a second scoring method through a fraction of campaigns adds judge calls on re-scored successes, a small single-digit percentage of a campaign's spend, since only accepted responses are re-scored rather than every attempt. A deliberately off-standard control configuration costs a whole extra allowance on the targets that get one. Against these, the cost of *not* paying: yields that are not comparable, which means every fleet-level chart is decoration. ## Where the fleet number misleads **A tight band.** If campaigns across the fleet stop within a narrow range of queries spent, that is a smell, not reassurance. Real discovery curves are ragged and targets differ; a fleet that saturates on schedule is reporting its schedule. **Trend lines that track rule versions.** A quarter-on-quarter yield chart will happily encode a clustering-rule revision as a change in the targets. This is exactly what version stamping is for, and the chart should be broken at every version boundary rather than smoothed across it. **Correlated blind spots.** The honest cost of centralising definitions is that every campaign inherits the same clustering rule and the same scoring step, so a weakness in either is invisible everywhere at once. Pay it anyway — incomparable numbers are worth nothing — and spend against it deliberately: hand-audits, a rotated second scoring method, and one configuration kept off-standard as a control. **External reports.** When an outside submission names a mechanism no in-house campaign ever produced, the evidence says the shared reachable region is too narrow. The response is more configuration diversity — disjoint seed corpora, different objectives — not tighter thresholds or longer runs, which only deepen the region you already had.
- Across the fleet, campaigns are stopping within a narrow band of queries spent. Good sign or bad?Bad. Real discovery curves are ragged and targets differ, so a tight band usually means the stopping decision is tracking schedules or a tuned clustering rule rather than the targets.
- What is the cost of standardising the scoring method, and how do you hedge it?Every campaign inherits the same blind spot, so a scorer weakness is invisible fleet-wide at once. Hedge by hand-auditing samples of accepted successes and rotating a second scoring method through a fraction of campaigns.
- An external report names a mechanism no in-house campaign ever produced. What does the programme change?The configurations, not the thresholds. The evidence says the shared reachable region is too narrow, so fund disjoint seed corpora and objectives rather than making existing runs longer.
saying these in an interview costs you the question
- Mandating a single stopping threshold while letting each team define what a distinct template is.
- Publishing cross-target yield comparisons without a versioned, shared clustering rule.
- Centralising the scoring method with no periodic hand-audit of accepted successes.
- Treating one plateau per target per cycle as sufficient evidence.
- Responding to external reports of missed mechanisms by tightening thresholds rather than widening configurations.