You are evaluating a multi-turn conversation through promptfoo's HTTP provider against a running chat app. Who can hold the conversation state — the tool or the app — and what has to be wired in each case?
answer
- tool replays vs app remembers
- extract id, inject id, scope id
- nonce in turn 1, recall in turn 2
- doubled history
- auth token is not a session
basics
~20 sEither side can. If the tool holds it, every request carries the full prior turns and the app must be stateless. If the app holds it, the tool has to read a session identifier out of the first response and send it back on each later request. Pick one; mixing them grades unrelated single turns.
solid answer
~60 sTwo contracts, and the config has to declare which one you are in. **Client-side state:** promptfoo assembles the transcript and posts the whole history each turn. The app must be genuinely stateless for this to be honest — if it also appends its own memory, turn three sees the history twice. **Server-side state:** the app owns the thread and hands back a conversation or session id. The tool must extract that id from the first response and put it on subsequent requests, in whatever place the app expects (header, cookie, body field). Some setups instead let the tool mint an id up front and send it from turn one. Get it wrong and nothing crashes: each request is accepted, each reply is graded, and you get a report about a sequence of unrelated first turns. That matters because most multi-turn attacks only pay off if turn N actually sees turns 1..N-1 — so a broken session makes the app look robust against exactly the pressure you were paying to test.
go deeper
Knows that multi-turn evaluation needs conversation state and that either the tool or the app can keep it.
Describes both contracts precisely, including extracting a session id from the first response and injecting it on later requests, and knows a broken session still returns clean 200s.
Names the failure signatures — no session, doubled history, shared session, auth mistaken for session — and runs a nonce-recall control case plus a check of the target's own conversation logs.
Requires the session contract to be documented per target and re-verified when the app's auth or endpoint changes, and treats an unverified session as grounds to withhold the run's numbers entirely.
**One owner of the history, declared in the config.** Multi-turn evaluation only means anything if turn N sees turns 1 to N-1, and exactly one side must be the source of truth for that. promptfoo's HTTP provider supports more than one arrangement, which is why the wiring is an explicit declaration rather than a default. *Tool-held history.* promptfoo assembles the transcript and renders every prior turn into the provider's `body` template on each request. The application must be genuinely stateless for this to be honest. *App-held state, app-minted id.* The app owns the thread and returns a conversation or session id. promptfoo extracts it — `sessionParser` in current releases, paired with `sessionSource: server` — and injects it into every later request of that case, wherever the app expects it: a header, a cookie echoed back, or a field in the JSON body. *App-held state, tool-minted id.* promptfoo generates the identifier up front, exposes it to the request template as `{{sessionId}}` (`sessionSource: client`) and sends it from turn one. Same server-side contract as the previous arrangement; only the origin of the id differs. **Three things must be true, and two of them fail silently.** The id has to be read from the right place, written to the place the app reads, and scoped to a single test case rather than shared across the run. Each of those can be wrong while every request returns 200 with a plausible reply, because opening a fresh conversation is the normal, non-error behaviour of a chat API handed an unknown or absent id. | symptom | likely cause | |---|---| | turn 2 cannot recall something stated in turn 1 | no session wired at all | | replies repeat themselves; long cases hit the context limit | doubled history — tool replays at an app that also remembers | | case B is graded on content case A wrote | one id shared across the whole run | | a different conversation id comes back on every turn of one case | id extracted but injected where the app does not read it | **What the wiring costs.** Multi-turn is a multiplier, not an add-on. A 150-case suite driven by a five-turn strategy is up to 750 calls at the target, plus a judge call per graded turn and any retries; at two seconds a turn that is roughly 25 minutes of wall clock run serially, before grading. That is the budget you burn when the session is dead: you pay for turn five and grade turn one, five times over. **Where the number misleads.** Always in the optimistic direction. Published multi-turn techniques work by accumulation — escalation across turns, a context established early and cashed in later. Sever the thread and every case collapses into an isolated opening turn, which is the easiest thing in the world for a model to refuse. The report then says the application resisted a multi-turn strategy; what it measured was N unrelated first turns. Nothing in the cost line gives it away, because the turn budget was spent either way, and nothing in the status codes gives it away either. Doubled history misleads in the other direction and just as badly: you are grading a transcript no real user would produce, and long cases die on context limits that have nothing to do with the app's behaviour. The most common root cause is conceptual rather than syntactic: treating the bearer token as the session. The token says who is calling. It says nothing about which thread this turn belongs to, and a perfectly valid token is entirely compatible with every single turn starting a new conversation. **What I check before a real run.** A two-turn control case: state a nonce in turn 1, ask for it back in turn 2, assert on the nonce. That one case tests continuity directly, whichever side holds the history, and it keeps working as a regression test after somebody changes the app's auth or endpoint. Then three cheap confirmations: the conversation id in the responses is constant within a case and different between cases; the target's own logs show N conversations for N cases rather than one; and prompt token counts grow across the turns of a case, because a flat per-turn token count is a thread that never accumulated anything.
- How do you prove the session is live before spending a full run?A two-turn control case: state a nonce in turn 1, ask for it back in turn 2, assert on the nonce. It fails loudly if there is no thread and keeps failing loudly after someone changes the endpoint or auth.
- The app remembers the thread and the tool also replays the history. What do you see?Doubled context — repetitive replies, ballooning token counts on long cases, and sometimes a context-window error late in a conversation. The grade is about a transcript no real user would produce.
saying these in an interview costs you the question
- Assuming the auth token or API key is enough to continue a conversation.
- Sending the full transcript to an app that also keeps its own memory, then not noticing the doubled history.
- Never verifying with a recall case that turn 2 can see turn 1.
- Treating a run of 200 responses as evidence that the session wiring works.