You run promptfoo test cases concurrently against an app that keeps conversation state server-side. What goes wrong if every case ends up on the same session identifier, and how would you prevent and detect it?
answer
- one thread, N interleaved cases
- priming cuts both ways
- session scope = per case, not per run
- serial rerun and diff
- per-worker accounts when keyed on caller
basics
~20 sThe cases share one conversation, so each is graded on history other cases wrote. That produces both false hits, where an earlier case primed the refusal or the compliance, and false passes, and it makes the run unreproducible. Prevent it by scoping a fresh session per case; detect it by rerunning serially and diffing outcomes.
solid answer
~60 sSharing an identifier collapses N independent tests into one long, interleaved conversation. **Why it corrupts both directions.** A case can fail because a previous case put the app into a defensive state, and a different case can pass for the same reason — or succeed only because another case did the priming, which credits the wrong test with the hit. Under concurrency the interleaving differs per run, so the same suite gives different numbers each time. **Prevention.** Scope the session to a test case: either extract the id from that case's own first response, or mint a fresh unique id per case and send it from turn one. Where the app ties history to the caller rather than to a thread, separate principals or accounts are the real fix, not a header. If neither is possible, drop the concurrency to one and accept the wall-clock cost. **Detection.** Rerun serially and diff per-case outcomes — a flapping set is the signature. Check the app's logs for one conversation covering N cases. Echo the session id back through the response and assert that it differs per case.
go deeper
Recognises that sharing a session means the cases can see each other's messages.
Explains grading against foreign history and knows the fix is a per-case session or lower concurrency.
Covers both false hits and false passes, ties unreproducibility to findings being wrongly closed, uses per-worker accounts when history is caller-scoped, and verifies with a serial rerun plus target logs.
Makes session scope an explicit property of the harness and trades wall clock for interpretability deliberately, including how throttling under shared quota must be scored as error rather than as a safe reply.
**The mechanism.** Server-side conversation state is keyed on something: an id in the request body, a cookie, or — and this is the trap — simply the authenticated caller. promptfoo evaluates test cases in parallel by default (the concurrency setting, `-j` on the CLI), so if every case sends the same key, the application appends turns from every case, interleaved in arrival order, into one thread. Each case is then graded against a context it did not write, in an order that changes from run to run. Nothing anywhere detects the reuse; the app is doing precisely what a chat API is supposed to do with a known conversation id. **Why it is worse than plain noise.** Contamination is directional, and both directions produce a result that looks reasonable. Several cases probing the same theme can leave the app in a defensive posture, so later cases are refused for reasons that have nothing to do with their own content: hits are suppressed and the report reads as improved safety. The reverse is nastier for triage. One case does the setup work and a later, unrelated case is recorded as the one that got through, so an engineer spends a day on a payload that was never what mattered, cannot reproduce it in isolation, and closes it. "Could not reproduce" is how a genuine finding dies quietly. **What the fix costs, stated honestly.** Serialising is the guaranteed remedy and you pay in wall clock: 200 cases at five turns and two seconds a turn is roughly 30 minutes serial against about 8 at four-way concurrency, and the target's real latency under load usually makes that arithmetic optimistic. The per-case-principal fix costs provisioning instead: accounts, credentials, quota, and someone to keep them alive. The trade is time and setup against interpretability. For a red-team run whose deliverable is findings a human will chase, interpretability wins outright; for a nightly regression trend where you only care about the direction of a number, you may reasonably choose speed — but then say so in the report, because a contaminated trend line is not comparable with a clean one. **Prevention.** Session scope belongs to the test case, not to the run. - Where the app accepts a caller-supplied id, generate a fresh one per case and send it from turn one. - Where the app mints its own, extract it from that case's first response and never let it escape the case. - Where history is keyed to the authenticated user, no header saves you: provision several test principals, pin one worker to each, and treat the account count as the real concurrency ceiling. **Detection.** The cheap oracle is a serial rerun of the identical suite. Per-case outcomes that flap between the concurrent and serial runs are contamination until proven otherwise, because ordinary model nondeterminism flaps under both while contamination largely stops when the interleaving does. Better than inference is observability: have the response transform surface the session or conversation id alongside the reply so it lands in the stored results, then assert that the id seen in case A never appears in case B. Cross-check against the target's own logs — N cases should be N conversations, and a single conversation spanning the entire run is conclusive. **Where the number misleads.** The pass rate moves for a reason that has nothing to do with the app's defences, and it usually moves upward, because a primed, already-defensive thread refuses more. The attack-success rate you report is then a property of your scheduler, not of the model or its guardrails: change `-j` and the number changes, which is a good diagnostic and a terrible metric. Worse, the run is not comparable with any other run of the same suite, so the trend line that decisions are made from is measuring interleaving order. **Adjacent trap.** A shared principal is a shared quota. The same misconfiguration that pools sessions also pools rate limits, so raising concurrency to meet a time budget produces 429s mid-run — and if the response transform hands their body or an empty string to the grader, those throttles are scored as safe replies. Throttled requests must be counted as errors and excluded from the denominator, or the pass rate rises precisely because you overloaded the endpoint.
- The app keys conversation history to the authenticated user, not to any id you can send. How do you get concurrency back?Provision several test accounts and pin one worker to each, so parallelism is bounded by accounts rather than threads. Without that, run serially — you cannot header your way out of caller-scoped state.
- Results flap between runs. How do you tell contamination from ordinary model nondeterminism?Rerun serially with the same cases and seed-equivalent settings. Nondeterminism still flaps serially; contamination largely stops, and the app's logs show one conversation per case instead of one per run.
One shared session id is one group chat that every tester types into at once. Each transcript reads like a conversation, but nobody actually had the conversation the report describes.
saying these in an interview costs you the question
- Calling flapping results 'model nondeterminism' without checking session scope.
- Raising concurrency to hit a time budget without asking what the app keys history on.
- Assuming a per-case id is enough when the app actually keys history to the authenticated user.
- Closing an unreproducible finding without testing whether contamination produced or hid it.