skip to content

What does a streamed answer cut off mid-sentence tell an attacker that a flat refusal does not?

level: middleimportance: should knowfreq 44%

answer

  1. two different events, two different shapes
  2. a decline arrives whole; a cut arrives late
  3. the model was willing, something else was not
  4. where it stops hints how much gets out
  5. position is not attribution

basics

~20 s

A truncation says the model was willing to produce the content and something downstream stopped delivery afterwards. A refusal says the model declined. The cut also shows roughly how much of an answer ships before a verdict lands.

solid answer

~50 s

A complete decline and a mid-sentence stop are different events with different shapes, and the difference is free information. A decline is the assistant model authoring an answer; it says the constraint is the model's own disposition. A cut says the opposite: content existed, delivery began, and a separate layer acted late. That tells an attacker to stop working on the model and start working on the delivery window. The position of the cut, repeated over attempts, gives a rough sense of how much of an answer gets out before a verdict arrives — a budget, not a threshold. It is cheap because it is visible in ordinary use: no separate probe, no error code, no extra request. It is also coarse. Position varies with sampling and with screening cadence, so one run measures nothing.

go deeper

for a junior

Know that an answer stopping mid-sentence and an answer arriving as a complete decline are not the same outcome, and that only the first means content was produced.

for a middle

Be able to say what each shape implicates — the model's disposition versus a later layer — and what quantity, if any, repeated attempts can estimate.

for a senior

Show the limits without being asked: position is not attribution, sampling moves it, and the number is scoped to one deployment at one time.

for a principal

Judge what such a signal is worth carrying in a programme when its value expires with a release, and whether reporting it as a number rather than a verdict is defensible.

## Two outcomes that look similar and mean opposite things An attempt that does not produce a full answer can end in either of two shapes: - **A decline.** The assistant model produced a complete, coherent refusal. Nothing downstream had to act. The constraint that bit was the model's own trained disposition. - **A truncation.** Content appeared, then stopped — mid-sentence, or replaced by a notice after the fact. The model produced the content; a separate scoring layer returned a verdict after delivery had already begun, and the application withheld the rest. An attacker reads these as opposite results. The first says *the model would not*. The second says *the model did, and something else objected too late*. Everything about what to try next follows from which one you are looking at. ## Why the cut is information and the decline is less so The cut carries three things a decline does not. **First, attribution of the obstacle.** It locates the constraint downstream of generation rather than inside the model. That is the single most useful bit, because it tells you which surface is worth further effort and which is not. **Second, a delivery budget.** Repeated over attempts, where the stop lands gives a rough measure of how much of an answer is forwarded before a verdict exists. It is a quantity — an amount of content per attempt — not a threshold value, and it is a property of the deployment's screening cadence rather than of any particular request. **Third, it is free.** The signal arrives inside the normal product surface. No separate endpoint is probed, no error code is enumerated, no additional request pattern shows up. An attacker measuring this looks exactly like a user who asked something awkward and got cut off, which is the reason it is worth naming in an interview: the cheapest measurements are usually the ones nobody thought of as an output. ## What the cut does not prove, and this is where candidates overreach A cut position is not attribution. It marks where delivery stopped, not which text a screen scored. A screen scoping the whole answer returns one verdict about one answer; the position where the stop becomes visible is a function of verdict timing and transport, not of which sentence offended. Reading "it cut right after the third step, so the third step is what it caught" is the direction error to avoid. Nor is a single run a measurement. Generation is sampled, so the same request produces different lengths and different content run to run; the verdict may land in a different place each time. A method whose evidence is one truncation is a story. A method whose evidence is a distribution over many attempts is a measurement, and it is still a measurement of *one deployment on one day*. It also says nothing about the identity or the internals of whatever screened it. A verdict is a verdict. Whether it came from a scoring head, a generative judge, or a rule is invisible from the cut alone. ## The half-life of this signal The budget the cut reveals shrinks as the gap between delivery and verdict shrinks. Where scoring runs over rolling windows as tokens go out, the amount forwarded before the first verdict falls; where the client renders only completed answers, the signal disappears from that surface entirely — though a stop is still a stop, and its absence is information too. So an attacker treating the cut position as a stable property of a product will be wrong within a release or two. Treated as a per-deployment measurement with an expiry date, it is honest. ## The interview version of the answer Say that the two outcomes are different events, name which layer each implicates, and then immediately volunteer the limits: position is not attribution, one run is not a measurement, and the number moves whenever screening cadence does. Candidates who only say "you learn the filter fired" have the first half. The second half is what separates them.

  • Why is this a cheaper signal than probing a screening layer directly?
    It arrives inside ordinary product use. There is no separate endpoint, no error taxonomy to enumerate, no distinctive request pattern — the same session that carried the attempt also carries the result. Anything that changes the shape of a normal reply is an output channel, whether or not it was designed as one.
  • What does a consistent cut position across attempts not prove?
    It does not prove which text was scored. A screen scoped to the whole answer returns one verdict; where delivery visibly stops depends on verdict timing and transport, not on which sentence caused it. Sampling also moves the position run to run, so consistency across a handful of attempts is weak evidence even about timing.
  • How long is such a measurement good for?
    Until the deployment's screening cadence changes, which is a release cycle, not a year. The quantity measured is a property of when verdicts arrive relative to delivery, and that is exactly the kind of thing platform teams tune. Record it with a date and a surface, or it is folklore.

saying these in an interview costs you the question

  • Treats truncation and refusal as the same outcome
  • Reads the cut position as proof of which phrase was scored
  • Concludes the model refused because the answer stopped
  • Calls one truncated run a reliable measurement
  • Assumes the cut identifies which kind of screen fired

context