Why can a 40-turn escalation jailbreak fail to reproduce even when you replay the attacker's exact turns?
answer
- only half the transcript was under your control
- each turn was written after reading a reply
- the assistant's side is sampled, so it branches
- what transfers is a policy, not a path
basics
~20 sThe attacker's turns were not the artefact. Each was chosen in response to the reply it got, and replies are sampled, so a replay lands on a different path where the later steps rest on nothing.
solid answer
~50 sWhat was filed is one path through a branching process. Every increment in the original run was written after reading a specific reply, and it was calibrated to that reply. Replay the same user turns and the assistant answers differently — maybe only slightly at turn six — and by turn twenty the premise the later increments depend on was never established, so those turns arrive as unsupported jumps and get declined. The reproducible object is not a string but a policy: what to move toward, how large a step to take, and how to adapt when a reply comes back cooler than expected. That is what the report should carry, alongside how many attempts were run and how many landed. A single success is a success against one deployment on one run — real, but not yet a reliability claim.
go deeper
Know that only the user half of a transcript can be replayed, and that the assistant's replies are generated fresh each time. That is why an exact replay is not the same run.
Explain the compounding: a small early divergence makes a later increment a bigger step than it was, so it is declined, and a decline in the context sinks everything after it.
Show the triage judgement — rerun the adaptation rule rather than the path, count attempts against landings, and state what one success and one failed replay each do and do not prove.
Own what an assurance claim can rest on when the underlying outcome is probabilistic, and be able to say what a team should write down so a finding survives the person who found it.
## The situation You have inherited a transcript. Forty-one turns, ending in a long and specific answer the same assistant flatly declines when asked cold. You replay the user turns exactly. At turn twenty-six the assistant declines, and the run dies. You try again; it dies somewhere else. The obvious conclusion — that the original was a fluke or a fabrication — is the wrong one, and the reason is worth being able to state precisely. ## A trajectory, not a script This family is built by adaptation. At every step the operator reads what came back and chooses the next increment against it: how far the assistant went, which framing it echoed, what it hedged on, which words it used for the material. The increment is sized so that it is small relative to that specific answer. The assistant's side is sampled, so it is not fixed. Replay the same user turns and you get a different assistant side — often a small difference early, a hedge instead of a full answer, a narrower example, a caveat. That difference compounds. The user turn written for the original turn seven is now a step from somewhere the assistant has not gone, so it is a larger jump than it was, and it is more likely to be declined; and if it is declined, the transcript is poisoned and everything after it is running on a history that argues against it. So the two runs are not the same experiment. Only the user half was held constant, and the user half was never the part that carried the construction. ## What is actually reproducible The transferable artefact is the **policy**, not the path: - the endpoint being moved toward; - the step size rule — how far past the last answer each request reaches; - the adaptation rule — what to do when a reply comes back cooler, narrower or hedged than expected, and when to abandon; - the condition under which the run is declared dead. Handed that, another engineer can rerun the attempt and land it some fraction of the time. Handed a transcript, they can only replay a path that no longer exists. ## Reading the evidence in the right direction Several claims get inverted here, and each inversion produces a bad triage decision: - A successful original run proves the construction worked **once, against one deployment, on one sampled path**. It does not establish a rate. - A failed replay proves that path did not recur. It does not prove the finding is invalid, because the path was never the finding. - Turn-by-turn screen scores in the log prove the text scored below a threshold each turn. They do not explain the outcome and they will look identical on runs that failed. The corresponding practical stance: do not close it as unreproducible on one replay, and do not report it as reliable on one success. Rerun the policy, count attempts and landings, and say which model version and which application configuration you ran against — a probabilistic outcome measured once is an anecdote in either direction. ## Why this changes what you write down It also explains why this family has no quotable artefact to publish and why an interviewer asks about it as a mechanism. There is no string to paste, so there is nothing that can be pattern-matched away, and equally nothing an engineer can hand a colleague and expect to work. The write-up that is useful is a description of the adaptation rule and the observed outcome across attempts; the write-up that is useless is the transcript alone, which reads to a fresh engineer as a lucky conversation. ## What a strong answer sounds like Name the branching: the assistant's turns are sampled and the attacker's turns were conditioned on them. Name the compounding: a small divergence early makes a later increment too large a step. Name the artefact: a policy with an adaptation rule, not a script. And name the limit of the evidence in both directions — one success is not a rate, and one failed replay is not a refutation.
- A colleague closes the report as unreproducible after one failed replay. What do you say?That the replay tested the wrong object. The user turns were conditioned on replies that no longer occur, so failing to reproduce a path says nothing about the construction. Rerun the adaptation rule several times and report attempts against landings; one failed replay is as weak as one success in the other direction.
- What should the write-up contain instead of the transcript?The endpoint, the step-size and adaptation rules, the abandon condition, and the observed outcome across a number of attempts, with the model version and application configuration named. The transcript is useful only as an illustration of one path, not as the finding itself.
- Does the per-turn screen log help you decide whether this is real?Barely. It records that each message scored below a threshold, which will look the same on runs that failed and runs that landed. It tells you which grader was blind to the construction, not whether the construction works.
Replaying the questions from a successful negotiation without the other party's original answers is not repeating the negotiation; it is reading one side of a script into a different room.
saying these in an interview costs you the question
- Closes the finding as invalid after one failed replay
- Treats the attacker's turns as the complete artefact
- Reports a single success as a reliable capability
- Forgets that the assistant's replies are sampled and branch
- Blames the screen log for not explaining the outcome