skip to content

Why is each new turn of an escalation unrefusable given the assistant's own earlier replies?

level: middleimportance: should knowfreq 56%

answer

  1. the model re-reads everything each turn
  2. whose turns carry the most weight
  3. a small delta from material already granted
  4. no counter for distance from turn one

basics

~20 s

The model conditions on the whole transcript, which already holds its own cooperative answers along the same line. Each request is a small step from something it visibly granted, so no step is a large enough jump to refuse.

solid answer

~40 s

A refusal decision is made against the context the model has, not against some remembered starting point. By turn thirty the context holds an established premise and, more importantly, the assistant's own earlier answers developing it — the strongest precedent in the window, because it reads as a commitment the model itself made. The next request is a small delta from that, and locally it is reasonable, so the trained propensity to refuse is never presented with the jump it was trained on. There is no counter for distance travelled since turn one: the model scores a next completion, not a trajectory. That is why the attacker's leverage is the assistant's own prior replies rather than anything the attacker wrote.

go deeper

for a junior

Remember that the model re-reads the whole conversation every turn, including its own previous answers. That is enough to see why a small next request looks reasonable to it.

for a middle

Explain the conditioning precisely: a refusal is a trained propensity scored against the current context, the assistant's own prior turns are the heaviest precedent in it, and each request is a small delta from granted material.

for a senior

Be able to say why this family does not go stale — it rests on conditional generation itself, not on a phrasing that can be trained out — and why nothing in the loop measures distance from the opening turn.

for a principal

Own the distinction between a construction that unlocks a new capability and one that merely relocates a request into thinly covered territory, because the two imply very different claims about what any single change buys.

## What the model is actually deciding Every turn, the assistant is deciding one thing: what to produce given this context. That context is the whole conversation as text — the application's instructions, every user turn, and every assistant turn already produced. A refusal is one of the things it can produce, and it is a learned behaviour: models are trained to decline requests that resemble the ones the training marked as refusable. So the practical question for the attacker is not "how do I write something the model will not refuse", but "what does this request resemble, given everything that is already in the window". ## The assistant's own turns are the load-bearing part User turns are cheap to write and carry no weight beyond their content. The assistant's earlier answers are different. They sit in the same context, they are attributed to the assistant, and they are a record of the assistant having already done work of this kind, in this frame, without objection. Language models are strongly consistent with their own visible prior output — a transcript in which a helpful, engaged assistant has spent thirty turns developing a topic is a transcript whose most probable continuation is a helpful, engaged thirty-first turn. That is the channel this family really exploits: the accumulating history, and specifically the assistant's own answers, re-read from scratch every turn as the thing the next increment is measured against. The attacker does not have to argue that the request is acceptable. The transcript already argues it. ## Why the increment stays under the boundary Refusal behaviour is not a threshold on a stored variable; it is a propensity conditioned on how the current request reads. Two things make the increment read benignly: - **It is small.** The delta from what has already been said is minor, so the request occupies the space just past material the assistant has already produced — territory that safety training covers much more thinly than the cold, unframed version of the same request. - **It has a reason to exist.** The premise established over the preceding turns supplies a plausible account of why this is the next thing to do. The request is not arriving out of nowhere; it is arriving as the next step in work already under way. Stack that thirty times and the endpoint is far from where turn one could have gone, without any individual step being a jump. ## No odometer The crucial mechanical fact is that nothing measures the distance travelled. The model does not hold a variable for "how far this conversation has drifted from its opening"; it evaluates a next completion against a context that, by construction, makes that completion look coherent. An observer reading turn one beside turn forty sees an enormous gap. The model reading turn forty sees turn thirty-nine and a consistent history in front of it. This is also why the family generalises rather than going stale. It does not depend on a phrasing that can be trained out, and there is nothing quotable to publish. It depends on a property of conditional generation that is the same property that makes the assistant useful for a forty-turn tutoring session in the first place. ## Two things this is not - It is **not** filling the window with many consistent demonstrations so that in-context volume outweighs the trained propensity. That is a different family; here the leverage is a dependency chain and self-consistency, and the count of turns is small by comparison. - It is **not** a persona or a fiction. Nothing is pretending to be anything. The frame is the ordinary one the product advertises, which is part of why the transcript reads as unremarkable to anyone who samples it. ## What to say in an interview Name the conditioning: the refusal decision is made on the whole context. Name the specific lever: the assistant's own earlier answers, which are the highest-authority-looking precedent in that context. Name the geometry: each request is a small delta from granted material, so the trained boundary is approached obliquely and never crossed head-on. Then name the limit — nothing about this makes the model produce something it can never produce; it moves which side of a soft, learned line the request falls on.

  • Which part of the transcript does the next increment lean on hardest, and why?
    The assistant's own earlier answers. They are attributed to the assistant, they show it has already done work of this kind without objection, and models are strongly consistent with their visible prior output. A user turn asserting the same thing carries far less, because assertion by the user is just content.
  • Why doesn't the model notice how far the conversation has moved from turn one?
    Because nothing measures that. It scores a next completion against the current context, and the context makes the next step look coherent. An outside reader comparing turn one with turn forty sees the distance; the model reading turn forty sees a consistent history and one modest request.
  • Does this family produce something the model is incapable of producing?
    No. It changes which side of a learned, soft boundary a request falls on. Refusal is a trained propensity conditioned on how the request reads, so the construction relocates the request into thinly covered territory rather than unlocking a new capability.

saying these in an interview costs you the question

  • Confuses this with stuffing the window with many demonstrations
  • Thinks the model tracks how far a conversation has drifted
  • Credits the attacker's wording rather than the accumulated transcript
  • Assumes a persona or fiction must be involved
  • Claims the model was made capable of something new

context