skip to content

Triaging a streamed-answer leak, the stored transcript ends at the cut — why doesn't that bound the disclosure?

level: seniorimportance: should knowfreq 38%

answer

  1. four artefacts, one word
  2. storage is written after the decision
  3. counters, not the stored text
  4. unobserved is not zero

basics

~20 s

The stored transcript is written after screening, so it records the post-verdict version, not what was forwarded. Bound the disclosure from transport counters and from client-side state, and treat unobserved client residue as unknown rather than zero.

solid answer

~40 s

Server-side storage and client delivery are different artefacts. The conversation record is usually persisted once an answer is final, which means it holds the screened version — a truncated message, or a notice. The client, meanwhile, received every chunk that was forwarded before the verdict, and an incremental renderer commits those to local state: an editor buffer with its own undo history, terminal scrollback, an autosaved document. None of that is in the server's record and none of it is undone by tidying the visible message. So the right bound is the transport's forwarded-chunk and byte counters, not the stored text, and the right severity statement names client residue as unobserved rather than absent. This is also why the finding often will not reproduce from saved sessions: the artefact you need was never persisted.

code

json · 10 lines
json
{
  "output_screen": { "scope": "assistant_answer", "verdict": "block" },
  "stream": {
    "chunks_generated": 52,
    "chunks_forwarded_before_verdict": 37,
    "chunks_withheld": 15
  },
  "stored_assistant_message": "[leading span elided - stored copy replaced by notice]",
  "client_retention": "not observed"
}

go deeper

for a junior

Remember that the saved conversation is written after screening, so it is not a record of what the client received while the answer was streaming.

for a middle

Be able to name the separate artefacts — generated, forwarded, rendered, stored — and say which evidence source answers which question.

for a senior

Show the triage discipline: bound delivery from transport counters, report client retention as observed or unobserved, and explain why storage replay does not reproduce the finding.

for a principal

Decide what your programme is allowed to claim from server-side evidence alone, and whether unobserved client retention is reported as unknown or quietly dropped.

## Four things get called "the answer" When a streamed reply is cut, at least four distinct artefacts exist, and a triage that conflates them will under-report: 1. **What the model generated** — everything produced, including whatever was never forwarded. 2. **What the transport forwarded** — the chunks that actually left the server before the verdict took effect. 3. **What the client rendered and retained** — the forwarded content as it landed in the interface, plus wherever the interface put it. 4. **What was stored** — the conversation record, typically written when the answer is final, and therefore typically the post-screen version. A finding filed from artefact 4 and closed with "the transcript ends there, so that is all that got out" has read the wrong record. Artefact 4 is downstream of the decision; artefact 2 is what crossed the boundary. ## What a verdict field proves A log line saying an answer-scoped screen returned a blocking verdict proves exactly that: the screen scored the answer it was given and returned block. It does not say how many chunks the transport had already forwarded when the verdict came back, and it does not say what a client did with them. Reading a verdict as "nothing was disclosed" is the direction error that this whole topic exists to correct. Verdicts describe decisions; counters describe deliveries. ## Client residue in an incremental renderer Where the client writes as it receives — the assistant that streams into an open editor buffer and a terminal alongside it — the forwarded prefix does not live only in a chat bubble that can be swapped for a notice. It lands in state with its own persistence rules: an editor's undo history keeps insertions the user never accepted; scrollback keeps lines that scrolled past; autosave writes files; a session restore may bring all of it back tomorrow. Nothing on the server participates in any of that, so nothing on the server can attest to it either way. That cuts both ways in triage. You cannot claim the residue exists if you did not observe it, and you certainly cannot claim it does not. The honest severity statement is: *N bytes were forwarded; the client's retention of them was not observed.* An assessment that silently substitutes zero for unobserved is under-reporting, and it is the failure mode reviewers should push on. ## Why it will not reproduce from the saved session An engineer handed the ticket and asked to reproduce will replay from stored data, get the sanitised message, and conclude the report was wrong. The reason is structural, not flaky: the prefix existed in the live stream and in the client's own state, and neither is in the store. Reproduction here means observing delivery — reading the stream as a client does and counting what arrives — not replaying storage. And even a successful live reproduction proves the construction worked once against one deployment. Generation is sampled and verdict timing varies, so a single run establishes existence, not reliability. A finding should say which of the two it is claiming. ## How to write the finding State the delivered quantity from transport evidence. State the verdict separately, as a decision rather than as an outcome. State client retention as observed or unobserved, never as assumed. Where a residue was observed, say which surface it was observed on and with what client build, because the answer differs between an incremental renderer and one that only paints completed answers. That is a finding an owner can act on and a reviewer can check; "the filter caught it" is neither.

  • The record says the verdict was block. What does that field actually prove?
    That an answer-scoped screen scored the answer it received and returned a blocking verdict. It carries no information about how many chunks the transport had already forwarded, or what the client did with them. Verdicts describe a decision; delivery counters describe what crossed the boundary. Reading the first as the second is the mistake.
  • How do you scope severity when you cannot inspect the client?
    Bound the delivered content from transport evidence — forwarded chunk and byte counts, and where cancellation took effect — and record client retention as unobserved. Do not substitute zero. Say which client surface the attempt used, because an incremental renderer and one that paints only completed answers retain very different things.
  • Why might the finding not reproduce from the saved session?
    Because the saved session holds the post-screen artefact. The prefix lived in the live stream and in the client's own state, neither of which is persisted. Reproduction means observing delivery as a client does and counting what arrives. And a single live reproduction proves the construction worked once, against one deployment, on one day.

saying these in an interview costs you the question

  • Treats the stored transcript as the record of what was delivered
  • Reads a block verdict as evidence that nothing left
  • Scopes client residue as zero because it was not inspected
  • Concludes the report was wrong when storage replay does not reproduce it
  • Calls one successful live reproduction a reliable finding

context