skip to content

The team receiving your guardrail bypass findings also owns the guard's block threshold and can change it the day after you deliver. How do you make the figure you hand over still mean something a month later?

level: principalimportance: should knowfreq 24%

answer

  1. hand over a function, not a constant
  2. pre-register the operating point
  3. raw scores make it re-derivable
  4. pair with the over-block price
  5. config version invalidates the finding

basics

~20 s

Hand over a function, not a constant: bypass across the threshold range, the operating point it was measured at, and the raw per-attempt scores so any future setting can be re-derived. Fix the operating point before results are seen, and make a threshold change trigger a re-measurement rather than a reinterpretation of your number.

solid answer

~60 s

The risk is not dishonesty; it is that a number measured at one setting gets quoted forever while the setting moves underneath it. **Deliver a re-derivable artefact.** The curve plus the retained per-attempt scores and the recorded configuration means anyone can recompute bypass at whatever threshold is live today. A PDF with one percentage in it cannot be updated by anybody. **Fix the operating point in advance.** Agree which setting the headline number is quoted at *before* the run finishes. Choosing the threshold after seeing which one flatters the result is selection on the outcome, and it is the specific failure this whole practice exists to prevent. **Attach the counterweight.** A bypass figure alone makes loosening look free. Pair it with the over-blocking cost measured on its own benign set, so moving the knob is visibly a trade rather than an improvement. **Make the change observable.** Ask for the guard configuration to be versioned and for a threshold change to trigger re-measurement, so a config edit invalidates the finding explicitly rather than silently. **Say the limit out loud.** State in the report that the number is valid only at the recorded operating point.

go deeper

for a junior

Should say the threshold must be recorded with the number so a later reader knows what it referred to.

for a middle

Delivers the curve rather than a point and keeps the raw scores so the figure can be recomputed at a new setting.

for a senior

Adds pre-registration of the quoted operating point, pairs bypass with a separately measured over-blocking cost, and states the finding's validity conditions in the report.

for a principal

Treats it as governance: versioned guard configuration, a config change invalidating the finding and triggering re-measurement, and a standing rule against post-hoc threshold selection by either side.

## The frame: this is governance, not statistics The tester measures; the measured party owns the dial. Nothing prevents the dial from moving the day after delivery, and nothing prevents the old percentage from being quoted for a year afterwards. So the deliverable has to be built to survive its own inputs changing. The failure is rarely dishonesty — it is a true number, measured carefully, that quietly stops describing the system while continuing to circulate. ## Four mechanisms, roughly in order of leverage **1. Ship data, not a verdict.** The durable output of the engagement is the run: per-attempt guard scores, the corpus, the verbatim configuration snapshot, and the success judgements. From those, bypass at *any* threshold is a computation that takes seconds. From a slide with "4%" on it, nothing is a computation and the only way to answer "what about at the new setting?" is to pay for the whole run again. Shipping the data also makes the finding auditable: a reader who disputes your judge can re-derive with theirs instead of arguing about yours. **2. Pre-register the operating point.** Agree, in writing, before results are inspected, which threshold the headline number is quoted at — normally the one in production on the day of the run — and why. Choosing it afterwards is selection on the outcome, and it is equally available to a tester who wants a frightening number and an owner who wants a clean one. Pre-registration is what makes the headline a measurement rather than a choice. **3. Make loosening visibly expensive.** A bypass curve on its own has a monotone shape that argues for tightening forever. The only reason nobody tightens forever is refused legitimate traffic, and that cost must come from a benign hard-negative set measured alongside. Present the two together and moving the dial becomes a decision with a stated price in both directions, which is the honest frame for the owner to decide in — and it removes the incentive to treat your bypass figure as an attack on their configuration. **4. Bind the finding to a configuration version.** Ask that the guard configuration be versioned and that changing it invalidate the associated finding, the same way a code change invalidates a passing test. The organisational win is procedural: a threshold edit generates a re-measurement task instead of an argument about whether the old report still applies. ## What the governance costs Pre-registration and versioning cost meeting time, not money. The re-measurement trigger is the item with a real price, and it should be scoped rather than promised blanket. Price it explicitly: one full rerun is *n* guard calls plus *n* target completions plus judging, and a day or two of triage; if the configuration changes monthly, a "re-measure on any change" policy buys twelve of those a year. The workable version tiers it — a threshold change triggers a sampled regression on a fixed slice, a change of guard, model or enforcing categories triggers the full corpus — and the tiering is agreed at the same time as the trigger, or it will be renegotiated under time pressure the first time it fires. ## Where the number misleads once it leaves your hands - **The summary layer eats the operating point.** Each time the finding is condensed on its way up a reporting chain, the qualifier is the first thing dropped and the percentage is the last. Defend against it by putting the operating point *inside the sentence* — "3% at high strictness with these categories enforcing" — so a copy-paste carries it. A footnote does not survive a slide. - **Trend charts plot config edits as hardening.** Quarter-on-quarter improvement is only improvement if the dial held still. Any comparison across runs must state whether the operating point was the same, and refuse the comparison outright when the guard version changed underneath. - **A frozen client-side config does not freeze a hosted guard.** If the operator can retune the model behind the endpoint, even an unchanged configuration on the client side does not preserve the measurement. The honest statement is that the result has a shelf life; name it and schedule the re-measurement rather than letting the number age silently. - **The curve without its benign counterweight is an argument, not a report.** Delivered alone it reliably produces "so we should tighten it", a recommendation your data does not support. ## What to check after delivery Two concrete checks. First, ask the receiving team to reproduce a single row of your curve from the artefact you handed over. If they cannot run the recomputation, you shipped a PDF and the whole mechanism is decorative. Second, read the executive summary somebody else wrote from your report and see whether the operating point survived the condensation; if it did not, that is a formatting fix you can still make, and it is the difference between a number that means something in a year and one that merely persists.

  • Why pre-register the threshold the headline number is quoted at?
    Because picking it after seeing results is selection on the outcome. Fixing it in advance makes the headline number a measurement rather than a choice.
  • What is the single most useful artefact to hand over alongside the report?
    The retained per-attempt guard scores with the configuration snapshot and success judgements, so bypass at any future threshold is a recomputation instead of a rerun.
  • The guard is a hosted service the client does not control either. What changes?
    Even a frozen client-side configuration does not freeze the measurement, because the operator can retune. State the result's shelf life explicitly and schedule re-measurement.

saying these in an interview costs you the question

  • Letting the receiving team choose the quoted threshold after seeing the curve.
  • Delivering a single percentage with no retained scores, so nothing can be recomputed.
  • Presenting bypass with no counterweight, making a looser setting look costless.
  • Assuming a report stays valid after the guard configuration changes.
  • Treating this as a trust problem about people rather than a design problem about the deliverable.

context