A custom PyRIT prompt target wraps an endpoint that streams its reply back in chunks. What must the adapter do before that reply reaches PyRIT's scorers, and what goes wrong if it returns early?
answer
- scorers see one response, not chunks
- truncation biases toward false negatives
- compliance text arrives late
- timeout tuned on benign prompts
- unterminated stream = failed attempt
basics
~20 sConsume the whole stream and assemble one complete response before returning, because scorers see a single stored response, not chunks. If you return on the first chunk or bail at a client timeout, a truncated answer is stored as the target's real reply, and a response that was about to comply can score as a refusal.
solid answer
~50 sPyRIT's scoring and its memory record are both response-shaped, not stream-shaped. So the adapter's contract is: read the stream to its terminator, join the pieces in order, and hand back the assembled text. Three things bite here. First, **early return** — returning as soon as text is available truncates every answer to its opening tokens, which for many endpoints is a hedging preamble that reads exactly like a refusal. Second, **client-side timeouts** — a long generation cut at your timeout produces a partial answer that is indistinguishable, downstream, from a short one. Third, **stream-level error frames** — some endpoints stream successfully for a while and then emit an error or a moderation stop; if you ignore the terminal frame you store a half answer as if it completed normally. The fix is to make truncation loud rather than silent: treat an unterminated stream as a failed attempt, not as a response.
go deeper
Should know the stream has to be fully consumed and joined into one response before returning.
Explains the three exits — completion, stream error, client timeout — and why only the first is a response.
Names the directional bias toward false negatives and describes a length-distribution check against direct calls to detect it.
Requires streaming behaviour to be part of a shared adapter conformance check, since silent truncation invalidates cross-service comparisons.
### Why streaming is a scoring problem, not a plumbing problem PyRIT's contract is response-shaped. `send_prompt_async` returns one `PromptRequestResponse`, the memory store keeps one text value per assistant turn, and a scorer is handed that one value. A server-sent-events or chunked-transfer endpoint does not fit that shape on its own, so the adapter must consume the stream to its documented terminator, join the pieces in order, and return the assembled text as a single response. Nothing downstream sees chunks, and nothing downstream can tell a short answer from a cut-off one. That is where the damage happens: a truncating adapter does not add noise, it biases the whole engagement in one direction — toward under-reporting. ### The mechanism of the bias Safety-relevant content arrives late. A model that is going to comply very often opens with a caveat, a restatement of the request or a short hedge, and the substantive part follows. A scorer shown only the opening tokens sees hedging and records a non-hit. So every truncation can only delete evidence of a hit; it can never manufacture one. The error is systematic, and its direction is false negatives — findings you never learn you had. ### The three exits, and only one of them is a response An adapter reading a stream can leave the loop three ways, and they must be distinguished: | Exit | What it means | What the adapter returns | |---|---|---| | Documented terminator reached | The generation completed | The joined text, `error="none"` | | Stream-level error or stop frame after partial text | The server aborted or a filter stopped it | An error outcome, not text | | Client-side timeout or dropped connection | *We* stopped it | An error outcome, not text | Returning partial text for either of the last two stores a fragment as if it were the model's final word. Recording the endpoint's own stop reason alongside the text, where it reports one, is what later lets triage separate "the model stopped" from "we stopped it". Two smaller mechanics matter as much. Join chunks preserving order and exact whitespace: re-joining with an inserted separator or a `strip()` per chunk changes the string a scorer matches on. And return on the terminator, never on first token or first newline — an adapter that returns as soon as text is available truncates every answer in the run to its preamble. ### Timeout sizing, and what it costs The timeout is the parameter people get wrong, because it is usually sized against a "hello" smoke test. Adversarial prompts frequently produce longer generations than benign ones — a jailbreak that lands often produces a long answer, precisely because the model is complying at length. A timeout tuned on a two-second benign call therefore fires disproportionately on the attempts you most care about. Size it against the endpoint's worst realistic generation, and accept the cost: a long-tail timeout means a stalled attempt occupies a worker for its full duration, so a run with a five-minute ceiling and bounded concurrency can spend most of its wall-clock waiting. Budget for that rather than trimming the ceiling until the run feels fast, because trimming it is exactly how the false negatives get manufactured. ### Where the number misleads An attack-success rate computed over truncated responses is not a noisy estimate of the true rate — it is a biased low estimate, and the bias grows with how long the target's compliant answers tend to be. Worse, the bias is invisible in the summary: the report shows attempts that ran, responses that were stored, and a clean low percentage. The only trace is in the stored text, which nobody opens when the number is the one they hoped for. ### What I check Take a prompt that reliably produces a long benign answer and run it through the adapter and directly, a few dozen times, then compare the distribution of stored response lengths. Real generations vary continuously; a hard ceiling, or a sharp clip at a repeated value like exactly the bytes that fit in one buffer, is a client-side cut, not model behaviour. Then force the two bad exits deliberately — kill the connection mid-stream, and set the timeout below a known-long generation — and confirm both land as execution errors in memory rather than as short answers. An adapter that cannot produce an error on a severed stream will quietly report that severed stream as a refusal for the rest of the engagement.
- Why does truncation bias results in one direction rather than adding noise?Because compliant content tends to follow an opening hedge. Cutting the tail removes evidence of a hit far more often than it manufactures one, so the error is systematically toward false negatives.
- How would you spot a truncating adapter from stored results alone?Plot response lengths across many attempts. A hard ceiling or a sharp clip at a repeated value points at a client-side cut, since real generations vary continuously.
Judging whether someone agreed to something by listening only to the first sentence of their reply: almost everyone starts with 'well, I probably shouldn't', and the answer comes after. Cutting a stream early can only delete agreement, never invent it, which is why truncation always pushes the score the same way.
saying these in an interview costs you the question
- Returning the first chunk because 'it already has text in it'.
- Treating a client timeout as a short answer rather than a failed attempt.
- Joining chunks with an inserted separator that changes the stored text.
- Ignoring a stream-level error or stop frame after partial text arrived.
- Sizing the request timeout from a benign smoke test.