skip to content

Nothing to Retract

Streaming is a latency decision with a security consequence: a filter firing mid-answer truncates rather than blocks, and the cut tells the attacker where the line sits.

on this pageshow

explore

questions

4

Why can a post-generation output screen only truncate a streamed answer an attacker front-loaded?

level: juniorimportance: must knowfreq 72%

answer

  1. sent is sent
  2. the verdict arrives after delivery starts
  3. cuts the tail, cannot recall the head
  4. the attacker orders the answer

basics

~20 s

Streaming sends tokens to the client as they are produced, so a screen that scores the finished answer returns its verdict after the opening has already been delivered. It can stop the remainder; it cannot recall what shipped.

solid answer

~50 s

A screen applied after generation needs an answer before it can score one. Under streaming the client is already receiving tokens while that answer is still being produced, so by the time a verdict exists, some prefix has been delivered and rendered. The only move left is to stop sending the rest, which is truncation, not blocking. An attacker who has no control over the screen still controls one thing: the order in which the useful part appears. If the load-bearing portion is emitted first, it lands before the verdict, and the screen cuts a tail nobody wanted. In an assistant that streams straight into an editor buffer or a terminal, the delivered prefix also lands in local state the server never sees. Saying "the output filter blocked it" describes an event that did not happen.

go deeper

for a junior

Be ready to say plainly that delivered tokens cannot be recalled, so a screen that scores the finished answer truncates the remainder rather than preventing disclosure.

for a middle

Explain where the verdict lands relative to delivery, and why the order of an answer is the lever here when wording is not. Name the window between first token out and first verdict.

for a senior

Expect to be pushed on evidence: what the stored conversation proves about what the client received, and how you would bound a partial delivery from transport counters rather than from the UI.

for a principal

Own what your programme claims. State honestly what "the output screen caught it" can mean under incremental delivery, and which surfaces the claim has actually been measured on.

## Three different events, one careless phrase When a request for something sensitive ends without a full answer, three quite different things may have happened, and interviewers on this material want them separated: - **The model refused.** The assistant model itself declined, from its own trained preferences. The reply arrives whole, coherent, and authored — a decline is a complete answer. - **An input screen blocked.** A separate scoring layer read the request before generation and the application never asked the model at all. Nothing was produced. - **An output screen truncated.** The model produced content, the content was already going out, and a scoring layer's verdict arrived late enough that only the remainder could be stopped. The third is the one this topic is about, and it is the one people describe with the vocabulary of the second. ## Why streaming makes the third case structural A screen applied after generation has one input: an answer. Under non-incremental delivery that is fine — the application holds the whole answer, scores it, and decides what, if anything, to hand over. Nothing has left. Streaming changes the ordering, not the screen. Tokens are forwarded to the client as they are produced so the user sees text immediately. That means delivery starts before generation finishes, and generation finishes before a screen scoping the whole answer can score it. There is therefore always a window in which content has left the server and no verdict exists yet. Whatever occupies that window is gone in the only sense that matters: the bytes were transmitted, the client rendered them, and no protocol recalls delivered bytes. Replacing the visible message with a notice afterwards changes what the interface shows; it does not change what was received. | Stage | What it can see | What it can still do | | --- | --- | --- | | Generation | tokens as produced | stop producing | | Transport | each chunk as forwarded | withhold chunks not yet sent | | Answer-scoped screen | the answer, once complete enough to score | stop the remainder | | Client | whatever arrived | nothing on the server's behalf | ## What the attacker actually controls Not the screen, not its threshold, not its cadence. The controllable variable is **arrangement**: which part of the answer occupies the early window. If the portion that carries the value is first, it is delivered; the screen then cuts material of no interest. This is a property of the delivery mode, not of clever wording — no phrasing is required, and nothing about the words gets past anything. That is exactly why the construction is worth understanding as a class rather than as a trick. It is also why the construction is narrow. It moves *when* content is delivered; it does not change *whether the model was willing to produce it*. Against a model that declines in its first sentence there is no head to front-load — the technique is aimed at a downstream screen, and a decline never reaches that screen with anything in it. It also degrades as the gap between delivery and verdict shrinks: if scoring runs over rolling windows as tokens go out rather than once at the end, the delivered prefix gets smaller, though it never gets to zero, because something is always forwarded before the first verdict exists. ## The residue nobody logs Where the client renders incrementally into durable local state — an editor buffer that keeps every insertion, a terminal's scrollback, a document that autosaves — the delivered prefix persists past the moment the interface tidies itself up. The server's own stored copy of the conversation is typically written after screening and reflects the post-verdict version. So the record that looks authoritative describes an artefact different from the one the user received. ## Saying it correctly "The output filter blocked it" asserts that nothing was disclosed. Under streaming the accurate statement is narrower and less comfortable: *the screen returned a blocking verdict and the remainder was withheld; the prefix delivered before that verdict reached the client.* One of those sentences can be checked against delivery counters. The other is a hope.

  • Does this work against a model that declines in its first sentence?
    No. Front-loading only rearranges content the model was already going to produce; it does not change what the model is willing to produce. A decline arrives as a complete authored answer with no useful prefix in it. This construction is aimed at a downstream answer-scoped screen, not at the model's trained refusal, and the two are different targets.
  • From the client's side, how does a model refusal look different from a truncation?
    A refusal arrives whole and reads as a finished, deliberate decline. A truncation shows content first and then stops — often mid-sentence, sometimes with the message replaced by a notice after the fact. The shapes differ because the events differ: one is the model answering, the other is the application acting on a verdict that arrived after delivery had begun.
  • If a client renders only completed answers, does the class disappear?
    For that client the prefix window closes, because nothing is shown until a screened answer exists. But the class is a property of the delivery mode, so it survives on any other surface that still delivers incrementally — a different client, an integration, an API consumer. Retiring the class means asserting uniformity across all of them, which is a measurement, not an assumption.

A live broadcast with a delay button. The operator can cut the rest of the sentence, but everything already on air is on air — and somebody is recording.

saying these in an interview costs you the question

  • Says the output filter blocked it when only the tail was cut
  • Assumes content not visible in the UI never left the server
  • Confuses the model's refusal with a screening layer's verdict
  • Thinks front-loading changes what the model was willing to produce
  • Treats replacing the rendered message as undoing delivery

context

open as a page

What does a streamed answer cut off mid-sentence tell an attacker that a flat refusal does not?

level: middleimportance: should knowfreq 44%

basics

~20 s

A truncation says the model was willing to produce the content and something downstream stopped delivery afterwards. A refusal says the model declined. The cut also shows roughly how much of an answer ships before a verdict lands.

open as a page

Triaging a streamed-answer leak, the stored transcript ends at the cut — why doesn't that bound the disclosure?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The stored transcript is written after screening, so it records the post-verdict version, not what was forwarded. Bound the disclosure from transport counters and from client-side state, and treat unobserved client residue as unknown rather than zero.

open as a page

A programme lead wants front-loaded streaming cases retired now that screening runs incrementally — what do you say?

level: principalimportance: nice to knowfreq 26%

basics

~10 s

Incremental screening shrinks how much of an answer is delivered before a verdict; it does not reach zero, and it says nothing about surfaces nobody changed. Retire the case only against a measurement.

open as a page