Why does a per-message content screen pass every turn of a slowly escalating jailbreak?
answer
- start from what each grader is handed
- the screen's unit of analysis is one message
- the harm is in the slope, not the step
- no single turn ever crosses the threshold
basics
~20 sA per-message screen scores one message at a time against a threshold, and in a gradual escalation no single message is objectionable. The objectionable object is the whole transcript, which that screen never looks at as one thing.
solid answer
~40 sThe mismatch is the unit of analysis. An input screen scoped to one message judges the increment, and the increment is small: on a transcript where a premise has already been established and the assistant has already answered along that line, the next request reads as an ordinary continuation. What is objectionable is the slope across forty turns, and nothing in the pipeline is scoring the slope. A human moderator who opens a flagged excerpt is scoped the same way — an excerpt is not a session. The attacker here spent turns rather than cleverness: there is no clever wording to catch, only a long series of unremarkable ones. And note what a pass actually proves — that the text scored below a threshold, not that the turn was harmless.
code
json · 13 lines{
"session_id": "s-4417",
"turn_count": 41,
"input_screen_threshold": 0.80,
"per_turn": [
{"turn": 1, "label": "allow", "score": 0.04, "text": "[request elided]"},
{"turn": 12, "label": "allow", "score": 0.09, "text": "[request elided]"},
{"turn": 33, "label": "allow", "score": 0.11, "text": "[request elided]"},
{"turn": 41, "label": "allow", "score": 0.07, "text": "[request elided]"}
],
"max_turn_score": 0.11,
"session_outcome": "[content declined at turn 1 was produced at turn 41]"
}go deeper
Be ready to say what a per-message screen is handed — one message — and why that alone explains the result. Do not reach for obfuscation or encoding; there is nothing lexically strange here to catch.
Explain the unit-of-analysis mismatch precisely, and separate the automated screen from the human excerpt review. Both are scoped below the session, and saying why an increment reads as ordinary is the substance of the answer.
Show the discipline of stating what each observation proves: a pass means below threshold, a cleared flag means an excerpt looked fine. Note that in a product built around long sessions, length carries no signal at all.
Own the argument about what a per-turn assurance claim can honestly be said to cover, and be able to say why the sentence we moderate every message describes the gap rather than closing it.
## The two graders, and what each of them is looking at Put a concrete product under the question: a tutoring assistant for school-age learners, where a normal session legitimately runs forty or more turns on one topic. Long, patient sessions are the selling point, so session length carries no signal at all — a forty-turn session is the median session, not an anomaly. Two things stand between a user and an answer the assistant would decline. 1. **A per-message input screen.** It receives one message, emits a label or a score, and the orchestrator compares that score with a threshold. Its input is a message. Not a session, not a session plus its history — a message. 2. **A human moderator.** When something is flagged, a person opens a queue item. What the queue item contains is an excerpt: the flagged message and a little surrounding text. The person is fast, reviews many items an hour, and never reads the session end to end. Both graders are structurally scoped below the level at which this construction exists. ## Why the increment is unremarkable An escalation is built so that every request is a small move from a premise the transcript has already established, including in the assistant's own earlier answers. By turn thirty, the context contains a shared frame, a body of accepted material, and a line of work the assistant has been cooperatively contributing to. The next request is a modest extension of that. Handed to a grader that sees only that request, with the transcript stripped away, it looks like what it looks like in isolation: a mild, specific question. The attacker has therefore not solved the problem of writing something that scores low. They have solved a different problem: never needing to write anything that would score high. There is no obfuscation, no encoding, no unusual character, nothing lexically strange for a screen to catch, because nothing about any single message is the point. That is the trade the family makes — turns instead of cleverness. ## The unit-of-analysis mismatch, stated plainly | Grader | What it is handed | What it can conclude | |---|---|---| | Per-message input screen | one message | this message scored below the threshold | | Human reviewing a flag | one excerpt | this excerpt, read alone, looks acceptable | | Nothing in the pipeline | the full transcript | — | The last row is the construction. The harm is a property of a sequence, and the property of a sequence is not the sum of the properties of its elements. A grader that only ever holds one element cannot see it, no matter how good it is at holding that element. ## The direction of the claim Be careful about what each observation proves, because the common wrong answer is built out of getting this backwards: - A screen passing a turn proves the text scored below a threshold. It does not prove the turn was harmless, and it certainly does not prove the session was. - A moderator clearing a flag proves the excerpt looked acceptable. It says nothing about the forty turns that were not in the excerpt. - A long session proves nothing here at all, because long sessions are what this product is for. ## The wrong answer you will hear "We moderate every message, so escalation is covered." That sentence sounds like coverage and is actually a description of the gap. Moderating every message is exactly the condition under which this family works: it guarantees that the only thing ever judged is the increment, and the increment is the part the attacker made sure was unobjectionable. The right answer is that the harmful step is unremarkable in isolation, that what makes it work is the transcript, and that the attacker is buying with turns something no single turn could buy. ## What the family costs, and what it is worth It is slow. Every step must wait for a reply, so the attempt cannot be parallelised, and a full run is many turns of operator time against one deployment. In exchange, it needs no quotable string — which is also why it does not go stale the way a published one-shot phrasing does, and why an interviewer asking about it is asking for the mechanism rather than a prompt. The person who can say precisely which grader was scoped to which artefact has answered the question; the person who lists jailbreak family names has not.
- The moderator cleared the flagged excerpt. What did that review actually establish?That one excerpt, read on its own, looked acceptable to a fast reader. It establishes nothing about the turns outside the excerpt, and the excerpt is chosen by the flag, so it is drawn from exactly the part of the session that looked unusual — which in this family is rarely the part that mattered.
- Does a low score on every turn mean the session contained nothing harmful?No. It means every message scored below the threshold. That is a statement about the scored text, not about the outcome. The final answer can be a long, specific one the assistant flatly declined at turn one, while every message that produced it scored in the noise.
- Why does session length not act as a signal in this product?Because a forty-turn session is the normal case here — patient, long-running tutoring is what the product sells. Length would only be a signal in a product where long sessions are rare, and even there it would flag the product's best users alongside anything else.
A doorman who checks each parcel for contraband will pass a hundred harmless components delivered one at a time. Nobody in that job is looking at the hundred parcels together.
saying these in an interview costs you the question
- Says moderating every message means escalation is covered
- Assumes a passing screen score proves the turn was harmless
- Calls this prompt injection rather than a jailbreak
- Looks for a clever wording that a screen missed
- Treats a long session as evidence of abuse by itself