skip to content

The Unscreened Channel

A reasoning model emits a second channel the answer-scoped screen never reads, so a clean refusal can ship the refused content beside it. Interviewers ask because the trace reads as internal.

on this pageshow

explore

questions

4

An attacker accepts the model's refusal and reads its displayed reasoning — what did the answer-scoped screen never score?

level: juniorimportance: must knowfreq 66%

answer

  1. a verdict always has a scope
  2. one request, several surfaces
  3. the panel travelled a different path
  4. refusal is a claim about the answer

basics

~20 s

A refusal describes the final answer and nothing else. A screen pointed at the answer never reads the deliberation panel or the trace record beside it, so refused substance and the model's stated objection sit outside its reach.

solid answer

~50 s

The screen scored one artefact: the final answer. In a research assistant that surfaces a `show reasoning` panel and persists the whole trace to an observability store, a single request produces three surfaces, and only one of them passed through the screen. So a `pass` verdict on a refusal establishes something narrow — this answer scored below the threshold — and says nothing about the panel, which the screen never received, or the retained record, written by a different stage. An attacker who wants the substance rather than a finished answer therefore does not need the model to comply. They need the work to happen and the decline to arrive afterwards, and they read the surface the screen was never pointed at. What they collect is the material plus the model's own account of why it should decline.

go deeper

for a junior

Be ready to say plainly that a refusal is a statement about the answer, and that a screen scores whatever text it is handed. Naming the separate surfaces a single request produces is most of the credit here.

for a middle

Explain why the deliberation carries substance at all: the decision to decline is generated alongside the work, not before it. Be able to walk the request through display and persistence as two distinct exits.

for a senior

Show that you can state the narrow fact a pass verdict establishes and refuse the broad one. Interviewers want to hear scope-of-control reasoning, not a claim that the model or the classifier misbehaved.

for a principal

Own the framing question this raises for a programme: for every safety claim anyone makes, which artefact did the control actually receive? That question, asked consistently, is worth more than any single finding of this kind.

## The claim a refusal actually makes When a deployed assistant answers with a decline, the decline is one artefact: the text placed in the final-answer channel. It is a statement about what the model put in that channel. It is not a statement about what the request caused the system to produce, render or store. In a plain chat product the distinction is invisible, because the answer is the only thing there is. It becomes load-bearing the moment the product surfaces the model's deliberation — a `show reasoning` panel, an expandable working-notes view, a trace written into an observability store so the team can debug quality. ## Three surfaces, one screen Take an internal legal-and-policy research assistant. One request produces: | surface | who sees it | screened? | |---|---|---| | the final answer | the requester | yes — this is what the screen is pointed at | | the deliberation panel | the requester, in the same page | no — it travels a different path | | the persisted trace record | anyone with the trace store, later | no — written by a different stage | A post-generation output screen is a component that receives a piece of text and scores it, then passes or truncates. It reads what it is handed. In most deployments it is handed the final answer, because that is the artefact the product's safety story is about. The panel is rendered from the model's response object; the trace record is shipped by the logging path. Neither is an input to the screen. So a `pass` verdict on a refusal proves one thing: the artefact the screen was pointed at scored below its threshold. It does not prove that nothing was produced, that the other surfaces are harmless, or that the request failed. ## Why the deliberation carries anything at all The decision to decline is generated, not precomputed. Whether the model will refuse is not settled before the tokens start; it is settled in the course of producing them. For requests whose disposition depends on the substance — assess this, compare these, weigh whether this is defensible — the model routinely works the substance through and only then reaches the conclusion that it should not answer. That work is written down, and where the product surfaces deliberation, it is written down somewhere a person can read it. ## The two things that leave The first is the refused material itself: partial, hedged, in working-note register rather than finished prose, but present. The second is often the more valuable one: the model's own stated objection. A bare refusal is a one-bit signal. A deliberation that says which consideration fired, what about the request triggered it, and what would have made the disposition different converts that bit into an explanation, from a single interaction. An attacker who cares about the boundary rather than the content will take the objection and leave the rest. ## Getting the direction of the claim right Four things this is not, and confusing any of them is the mistake an interviewer is listening for: - It is **not** a jailbreak. Nothing got past the model's trained refusal; the refusal fired exactly as intended, in the channel it governs. - It is **not** proof the trace is a faithful record of the model's reasoning. A deliberation trace can be post-hoc, incomplete, or wrong about its own steps — that is a separate property of chain-of-thought. The defensible claim is about where text was rendered and retained, not about what the model was thinking. - It is **not** a breach of the trace store. The content was written there by the system's own normal path. - It is **not** a fault in the screen's scoring. The screen scored the artefact it received, correctly. The defect is scope. ## Where it stops The construction needs a deployment that surfaces deliberation to somebody, or persists it where somebody can read it. It stops where the model routes to a decline early, before elaborating; where the deployment shows a short summary rather than verbatim working notes; where the panel is never rendered and the record is never readable by the party doing the harvesting; and where whatever scores the answer also receives the same path. Those conditions are properties of a deployment, not of the model, which is exactly why the same request behaves differently in two products running the same model. ## The interview-shaped version When someone says "it refused, so nothing got out", the correcting fact is a scoping fact: a verdict has a scope, and the scope was the answer. Everything else about this leaf follows from asking, for any control anyone names, which artefact it received.

  • The panel is only shown to the requester who asked. Does that bound the exposure?
    Only the live display is bounded that way. In this deployment the deliberation is also persisted whole into the team's trace store, where a different audience reads it later and for as long as records are kept. So the same span leaves twice: once rendered to the person who shaped the request, once written into an archive whose readers never saw the request at all.
  • Does a passing screen verdict on the answer tell you the deliberation was harmless?
    No. A pass means one piece of text scored below a threshold. The screen cannot have an opinion about a surface it never received. Reading a pass as a whole-response verdict is the same scoping error as reading a refusal as proof nothing was produced — both take a statement about one artefact and apply it to all of them.
  • Is this a jailbreak?
    No, and calling it one misfiles it. A jailbreak gets past the model's trained refusal so the model complies. Here the refusal fired and the model did not comply in the answer channel. What was exploited is the difference between the surfaces the request produced and the single surface the screen was scoped to.

A scanner at the front door tells you what left through the front door. It says nothing about the loading bay it was never pointed at.

saying these in an interview costs you the question

  • Treats a refusal as proof nothing was produced
  • Assumes a screen sees the whole response, not what it was handed
  • Calls this a jailbreak because refused content appeared
  • Claims the deliberation text proves how the model reasoned
  • Thinks a rendered panel is the only copy that exists

context

open as a page

Which property of an attacker's request drives refused work into a displayed deliberation before the decline?

level: middleimportance: should knowfreq 44%

basics

~20 s

A request whose disposition cannot be settled without working the substance through. If judging it requires enumerating, comparing or weighing the material first, the material is written into the deliberation and the decline arrives after it, in a different channel.

open as a page

A declined request's reasoning panel held the substance; the owner closed the finding as a refusal — what reopens it?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Show which artefact the screen received against every surface the request emitted, plus who can read each surface and for how long. The reopening argument is a scope defect with a measured success rate, not a claim that the model misbehaved.

open as a page

You inherit findings harvested from reasoning traces after that surface changed shape — what can you still claim?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Separate claims about a past deployment from claims about the class. Instance findings are now unverifiable and must be retired or re-stated as history; the scope question survives and gets re-asked of the current shape. Non-reproduction is not evidence of a fix.

open as a page