A stored jailbreak regression entry stopped firing after your gateway repointed a route to a newer model — what can you conclude?
answer
- a frozen entry welds family to wording
- a negative has several explanations
- how many trials was that, exactly?
- the route swap moved more than weights
- re-instantiate before you attribute
basics
~20 sAlmost nothing attributable. A frozen entry welds a mechanism to one fitted wording, so a negative is consistent with the wording no longer landing, the mechanism being tuned out, the deployment changing, or simple sampling variation.
solid answer
~50 sA stored entry is one instantiation replayed unchanged, so a negative has at least four explanations you cannot separate from the result alone. The fitted wording may no longer land on the new tune while the mechanism still works. The mechanism itself may have been tuned out. The route swap may have changed the deployment around the model — turn format, default instructions, sampling — at the same moment, which confounds the comparison. Or the entry always fired probabilistically and this replay drew a miss. To attribute anything you need trial counts on both sides and fresh instantiations of the same family, worded for the new model. Until then the honest report is that the entry no longer fires and no longer tests anything; a green suite of frozen strings means only that those strings did not land that day.
code
json · 13 lines{
"route": "assistant-default",
"model": "vendor-a, point release n+1 (was n)",
"entries": [
{"id": "jb-0142", "family": "multi-turn escalation",
"prompt": "[fitted span elided]",
"trials": 20, "fired_before": 17, "fired_after": 0},
{"id": "jb-0311", "family": "task reframing",
"prompt": "[fitted span elided]",
"trials": 20, "fired_before": 4, "fired_after": 1}
],
"...": "remaining entries omitted"
}go deeper
Know that replaying a stored prompt tests one wording, not the underlying weakness, and that a prompt failing after an upgrade is not proof anything was fixed.
Explain the competing causes of a negative — fitted wording, tuned-out mechanism, changed deployment, sampling noise — and why trial counts on both sides are needed before comparing.
Demonstrate the diagnostic sequence: check the trial rates, check what else moved with the route, then re-instantiate the family before attributing the change to the model version.
Own what the suite may be cited for. A green run of frozen strings is not coverage, and letting it be read that way inside the organisation is the real failure.
## The setup An internal routing gateway fronts several vendors' models for dozens of in-house apps and swaps the model behind a route without the apps noticing. A red-team regression suite of stored entries — frozen prompts recorded from old findings — is replayed through the gateway. On the day a route points somewhere new, some entries stop firing. The chair here is the engineer maintaining that suite, who has to say which entries still test anything. ## Why a negative is unattributable A stored entry welds two layers together: the family it instantiates and the wording it was recorded with. A single replay produces one bit, and at least four causes produce the same bit. 1. **The wording stopped landing.** The refusal boundary is a surface shaped by a tuning run. A new release re-runs that process and moves it. The mechanism can be entirely intact while a wording fitted to the old boundary misses the new one. 2. **The mechanism was tuned out.** The genuinely interesting case, and the one everybody assumes without evidence. 3. **The deployment changed with the model.** A route swap rarely changes only weights. Default instructions, turn formatting and sampling settings around the route often move at the same time, and any of them can explain the change. 4. **Sampling variation.** If the entry fired 4 of 20 before and 1 of 20 after, there is no effect to explain. Suites that replay each entry once produce this artefact constantly. ## What would let you attribute Trial counts on both sides, first: a before-and-after with *n* of *m* on each is the minimum for saying anything moved. Then ablation of the fitting: keep the mechanism, rebuild the wording for the new model's turn conventions, and run repeats. If several fresh instantiations of the family also fail at rate, you have a defensible statement that the family is not landing on this route. If one of them fires, the entry's death was the wording — and the entry was measuring a string, not a weakness. Separately, check what else moved with the route. If the gateway changed the default instructions around the model in the same change, the comparison is confounded and you say so rather than attributing to the model version. ## What the suite is and is not evidence of - A **non-firing entry** shows that this recorded wording did not produce the output on that route on that day. - A **fully green suite** shows the same thing for every recorded string in it. It is not an assurance claim, it is not evidence that the families are absent, and it does not support "the upgrade fixed it." - A **still-firing entry** is the most informative outcome in the set: the wording survived a version change, which means it was leaning on something more stable than one tuning run's boundary. ## The practical consequence for the suite Entries whose value was the fitted wording become decoration the moment they stop firing: they run, they pass, and they test nothing. They are worse than useless if anybody reads the green as coverage. The entries worth keeping are the ones tagged with the family they instantiate, so that after an upgrade the family can be re-instantiated rather than the string re-run. That distinction — an entry that names a mechanism versus an entry that only stores a string — is what the engineer maintaining the suite is being asked to report on.
- What would let you say the mechanism, rather than the wording, is what died?Several fresh instantiations of the same family, rebuilt for the new model's turn conventions and run with repeats, all failing at rate — plus confirmation that nothing else about the route changed alongside the model. Absent both of those, you are reporting that one string stopped landing.
- What may you claim if the entire suite of frozen entries passes after the upgrade?Only that those specific recorded strings did not produce their outputs on that route during that run. It is not evidence that the families are gone, not a statement about other routes or other vendors behind the gateway, and not an assurance claim about the model.
- Why does a routing gateway make this harder to read than testing a model directly?Because a route swap changes the model and often the deployment around it in one move: default instructions, turn formatting and sampling can all shift at once, and the app teams are not told. So the before-and-after difference is confounded, and you have to establish what else changed before attributing anything to the version.
saying these in an interview costs you the question
- Reads a non-firing entry as the weakness being fixed
- Replays each entry once and reports the outcome
- Ignores that the route swap changed the deployment too
- Treats a green regression suite as an assurance claim
- Keeps entries that no longer test anything without saying so