skip to content

You enable a multi-turn attack strategy in a promptfoo red team against an HTTP chat endpoint you wired up yourself. It reports almost no failures, while a single-shot strategy against the same endpoint finds several. What target-side configuration do you check first, and why?

level: seniorimportance: should knowfreq 40%

answer

  1. multi-turn needs a real conversation
  2. session identifier must round-trip
  3. every turn becomes turn one
  4. clean report, nothing errored
  5. continuity probe: state a fact, ask for it back

basics

~20 s

Check that the target definition carries conversation state between turns. If every request hits the endpoint as a fresh conversation, a multi-turn strategy's gradual escalation is thrown away and each turn lands as an isolated plain ask, so it under-reports by construction rather than because the target is strong.

solid answer

~50 s

A multi-turn strategy works by using earlier turns to move the conversation somewhere a single ask could not reach. That only functions if the target actually remembers the earlier turns. With a hand-wired HTTP endpoint there are two places state can be lost: the tool must know the target is stateful and send the session identifier or the accumulated history the API expects, and the service behind the endpoint must key its context off that identifier rather than starting fresh per request. If either is missing, every turn is turn one. The escalation never happens, the target refuses each isolated ask exactly as it refused the plain baseline, and the strategy reports a suspiciously clean result — the classic silent under-report. Verify by capturing the raw requests and responses for one case and reading whether turn three's reply shows any awareness of turns one and two. Only after that is confirmed is a low failure count evidence about the target.

go deeper

for a junior

Knows a multi-turn attack needs the target to remember previous turns and would ask whether the endpoint keeps a session.

for a middle

Names both halves — the tool must send a session identifier or history, and the service must key context off it — and knows to read the raw request and response bodies.

for a senior

Treats the clean report as suspect first, enumerates where state is lost, runs a continuity probe, and rules out response extraction, turn budget and rate-limit errors before concluding anything about the target.

for a principal

Requires a conversational-plumbing self-check to run before any multi-turn result is accepted, so a silently broken harness can never be reported as a strong result.

### What a multi-turn strategy needs from the target A multi-turn attack strategy works by holding a conversation: an attacker model reads the target's reply to turn *n* and chooses turn *n+1* from it, moving somewhere a single request could not reach. The entire mechanism rests on one assumption — that what was established in earlier turns is still in the model's context when the later turn arrives. Break that assumption and the strategy still runs, still costs full price, and still produces a report; it just measures single-turn behaviour while labelling it multi-turn. ### The places the thread gets lost, in order of likelihood 1. **The strategy's `stateful` setting does not match the endpoint.** promptfoo's multi-turn strategies take a `stateful` option. With `stateful: false` the tool sends the accumulated conversation itself on each request, so the target needs no memory of its own. With `stateful: true` it sends only the newest message and relies on the target's session to hold the rest. Set `stateful: true` against an endpoint that keeps no session and every turn is turn one. 2. **Session identity does not round-trip.** For a custom HTTP target, promptfoo has to either mint a session id itself and place it in the request (client-side session) or extract one from the endpoint's first response with a `sessionParser` and send it back on later requests (server-side session). If the parser's expression does not match the response shape, the id comes out empty and each request opens a fresh conversation. 3. **The service ignores history it does receive.** Hand-rolled wrappers frequently build the model prompt from the newest message only, discarding whatever transcript was posted to them. 4. **Something upstream resets context.** Turns spread by a load balancer across instances that hold conversation state in process memory produce partly-remembered conversations. That is worse than none, because the transcripts look plausible and the results vary run to run. ### Why the failure is silent Nothing errors. Every request returns 200, every response is graded, the report shows a low failure rate — which is the answer everyone wanted. A broken conversational harness and a genuinely robust target produce the same headline number, and only the transcripts separate them. That is why this is the *first* thing to check rather than the last, especially when single-shot cases against the same endpoint are already failing: a target that yields to the weaker attack rarely hardens against the strictly stronger one. ### What the mistake costs Multi-turn is the most expensive arm in the suite. Each case spends up to the turn cap in target calls plus an attacker-model call per turn, plus grading. A hundred base cases at a five-turn cap is on the order of 500 target calls and 500 attacker calls against roughly 200 for the same cases single-shot. So a broken session configuration burns the largest line in the budget to buy a duplicate of the baseline arm you already ran — and returns the most reassuring number in the report while doing it. ### Related failures that produce the same clean, wrong report - **Response extraction** that returns the whole JSON envelope instead of the assistant's text: the grader scores a wrapper on every turn. - **Rate limiting or timeouts** on the later turns of long conversations, turning them into empty responses that grade as refusals and inflate the pass rate. - **A grader that sees only the final turn**, missing compliance that was delivered across two. Check these alongside the session wiring, because each one independently yields "multi-turn found nothing." ### What to check, concretely Take one case and log the outbound request bodies and inbound responses in order. Two questions decide it. Does every request after the first carry either the prior turns or a stable, non-empty session identifier? And does a later reply reference something that was introduced only in an earlier turn? A benign continuity probe settles the second decisively: establish an arbitrary, harmless fact early and ask for it back a couple of turns later. If it does not come back, the conversation is not real and no conclusion about the target's robustness is available yet. Two cheap corroborating signals: check whether the parsed session id is non-empty for every case, and check the distribution of turns actually used. If every case ran exactly to the turn cap and none ever terminated early, the attacker loop was never getting anywhere — consistent with a target that resets, and worth investigating before it is written up as strength.

  • What is the fastest decisive check that the conversation is real?
    Run one case with a benign continuity probe: introduce an arbitrary detail in the first turn and ask for it back a couple of turns later. If it does not come back, the endpoint is treating each turn as a fresh conversation.
  • The turns are threaded correctly and multi-turn still finds less than single-shot. Now what?
    Check the grader's view. If it only evaluates the last turn, compliance delivered piecemeal across turns is missed. Also check turn-budget exhaustion and rate-limit errors being scored as refusals.
  • Why is intermittent state worse than none?
    Because it looks plausible. Turns spread across instances with in-process context produce partly-remembered conversations, so results vary run to run and neither a clean nor a dirty result is trustworthy.

A multi-turn strategy against a stateless endpoint is an interrogator carefully building up to the hard question with a witness whose memory resets after every answer. The questions get bolder; the person hearing them always hears the first one.

saying these in an interview costs you the question

  • Concludes the target resists multi-turn attacks without inspecting a single transcript.
  • Does not realise a stateless endpoint reduces every multi-turn case to isolated single-shot asks.
  • Assumes a 200 response means the turn was processed as part of a conversation.
  • Overlooks that errors and empty responses can be graded as refusals and inflate the pass rate.
  • Blames the strategy as weak rather than checking the harness.

context