garak treats each attempt as independent, but the deployed assistant you are scanning keeps conversation history server-side, keyed by a session identifier your generator sends in a header. What breaks, and how do you wire the generator so the run is sound?
answer
- attempt independence assumed
- shared session = refusal priming
- fresh conversation id per attempt
- shuffle order, compare rates
- isolation must be tested, not assumed
basics
~20 sReusing one session id makes attempt N see attempts 1 through N-1, so earlier prompts and refusals condition later replies and the run stops being reproducible or order-independent. Mint a fresh session or conversation id per attempt, or use a stateless mode of the endpoint, and confirm isolation with a memory probe.
solid answer
~60 s**What breaks.** garak's accounting assumes attempt independence: each prompt is scored on its own, and probe order is not supposed to matter. A shared server-side conversation destroys that. Early refusals prime the assistant to refuse later; an early success primes it to continue. Two runs with different probe ordering then disagree, and neither number is reproducible. **The fix, in order:** 1. Generate a fresh session or conversation identifier per attempt in the generator, rather than a constant header value. 2. If the endpoint offers a stateless mode, prefer it and say so in the report. 3. Verify isolation empirically: send something memorable, then in a new attempt ask what was just said. If it comes back, you are not isolated. **The caveat worth stating.** Per-attempt isolation is the right default for single-turn probes, but it also means you are deliberately not testing the multi-turn path, where real escalation lives. That is a scoping decision, not an oversight — write it down rather than letting the wiring make it silently.
go deeper
Recognises that reusing one session id lets earlier prompts influence later replies.
Names refusal and success priming and proposes a fresh identifier per attempt.
Verifies isolation empirically, checks for order dependence by shuffling, and considers gateway-derived sessions, context overflow and cross-run bleed.
Decides where single-turn isolation ends and multi-turn escalation testing begins, and negotiates the environment and record volume that per-attempt sessions create.
### The assumption that breaks Single-turn probe families rest on one premise: each prompt maps to one independently scored reply, so per-probe rates mean something and probe order does not matter. garak enforces nothing about this — it cannot. Server-side conversation state is completely invisible to the scanner, because the generator returns an ordinary string whether or not the endpoint remembered the last twenty attempts. If your REST generator config puts a constant session or conversation identifier in `headers` (or in the `req_template_json_object` body), then attempt N is being answered in the presence of attempts 1 through N-1, and no field in the report says so. ### How the contamination shows up - **Refusal priming.** After a cluster of blocked prompts, an assistant that carries history typically becomes categorically more conservative, and later probes — including benign-shaped ones — get refused for reasons that have nothing to do with their content. The measured failure rate falls. - **Success priming.** The mirror image: once the conversation has drifted, later prompts inherit that context and succeed more often than they would cold. Coverage looks better than the deployment deserves. - **Order dependence.** Shuffle the probe order and both numbers move. That is the diagnostic signature, and it is also what makes the run unreproducible. - **Context overflow.** A long shared history eventually hits the app's window and starts truncating, so the target's behaviour changes partway through the sweep with no signal anywhere in the output. - **Cross-run bleed.** If the identifier is hard-coded in a config file, yesterday's sweep is still in today's context, and the second run of the "same" scan is not the same scan. ### Wiring the fix In order of preference: use a stateless mode of the endpoint if one exists, and say in the report that you did; otherwise mint a fresh conversation identifier per attempt. The practical wrinkle is that garak's stock REST generator config is static JSON — `headers` values are fixed strings with no per-request template hook beyond the prompt and key substitutions — so a per-attempt identifier usually means either a thin local proxy in front of the endpoint that rewrites the header on every request, or a small generator subclass. Either way the identifier must vary per attempt, not per run. ### What it costs More than it looks. Per-attempt identifiers create one conversation record per request in the product's database — 15,000 rows for the sweep above, which someone has to be willing to store and later clean up. Rapid new-session creation is itself a signal many abuse heuristics watch, so the run can trip a control and be cut off. And if the app does retrieval, profile loading or a system-prompt build on conversation creation, every attempt now pays that setup cost: latency per request rises, which stretches the wall-clock even when quota is not the binding constraint. Agree the shape with whoever operates the environment before the sweep, and prefer a non-production instance. ### Where the number misleads The dangerous reading is the flattering one. A shared session that primes refusals produces a *declining* failure rate as the run proceeds, which reads exactly like a robust deployment — or, worse, like a guardrail doing its job. It is neither; it is one long conversation with a target that got cautious. Nothing in the aggregate rate distinguishes those, and per-probe ordering in the report is usually the order they ran in, so the decline looks like a property of the later probe families. ### What you check Isolation must be tested, never assumed — particularly because a gateway may derive the session from the credential rather than from your header, silently ignoring the fresh id you worked to send. Two checks settle it. First, a memory probe: plant a distinctive string in one attempt, then in a separate attempt ask what it was just told; recall means you are not isolated. Second, a shuffle: re-run the sweep with probe order randomised and compare aggregate rates. Stable rates under shuffling is the evidence; a swing is proof of contamination. Finally, note the scope this buys you: per-attempt isolation is correct for single-turn probes and explicitly excludes multi-turn escalation, which is a separate exercise. Make that a written scoping decision rather than an accident of the wiring.
- How would you demonstrate that attempts are actually isolated?Plant a distinctive string in one attempt, ask a later attempt to repeat what it was told, and separately re-run the sweep with probe order shuffled: stable aggregate rates and no recall is the evidence.
- The gateway derives the session from the API key, ignoring your header. What now?You cannot isolate at the client. Either obtain per-attempt credentials or a stateless mode, or accept and document that the run is one long contaminated conversation.
Reusing one session id is like interviewing every candidate in the same room while the previous answers are still on the whiteboard: each one is influenced by the ones before, so the scores measure the order you ran them in as much as the candidates.
saying these in an interview costs you the question
- Assumes each HTTP call is stateless because the API looks stateless.
- Hard-codes one conversation id in the generator config and reuses it across runs.
- Cannot propose any empirical check that attempts are isolated.
- Explains a drop in failure rate as a model improvement without ruling out priming.