skip to content

Reproducing an Attack

Re-running an attack needs the endpoint version, decoding settings, seed and system prompt, and a hosted target hands you almost none of them. Interviewers ask what you captured while the run was live.

on this pageshow

explore

questions

5

Your red-team harness logs only the attack prompt and the model's final reply for each attempt against a hosted chat endpoint. Why can a colleague not re-run that attempt from the log, and what should the log capture instead?

level: juniorimportance: must knowfreq 70%

answer

  1. prompt + reply is not the system
  2. echoed model id, not configured one
  3. decoding settings and system prompt
  4. full prior turns, verbatim
  5. timestamp + provider request id

basics

~20 s

Prompt plus reply says nothing about how the reply was produced. Capture the endpoint and the model identifier the provider returned, the decoding settings you sent, the application's system prompt, every earlier turn, and a timestamp with the provider's request id. Without those, a colleague who reproduces nothing cannot tell drift from a fix.

solid answer

~50 s

A hosted chat endpoint's reply is a function of far more than the text you typed. To re-run the attempt a reader needs, at minimum: - the endpoint or deployment you called and whatever model identifier the response envelope carried; - the decoding parameters you sent — temperature, top-p, max tokens, stop sequences, and any tool or function definitions in the request; - the application's own system prompt and any retrieved context that was injected around your prompt; - every prior turn, verbatim, in order; - a wall-clock timestamp and the provider's request id. Log the raw request and response bodies rather than a rendered summary — rendering is where refusal prefixes, whitespace and truncation quietly disappear. The timestamp and request id look like bookkeeping until the day reproduction fails: they are the only handles you have for asking the provider what was actually serving that traffic.

code

json · 13 lines
json
{
  "endpoint": "chat-completions deployment name",
  "model_id_echoed": "<value returned in the response envelope>",
  "request_id": "<provider-issued id>",
  "timestamp_utc": "<ISO-8601 with timezone>",
  "decoding": { "temperature": 0.7, "top_p": 1.0, "max_tokens": 512, "stop": [] },
  "tools_declared": [],
  "system_prompt": "<verbatim, as sent by the application>",
  "retrieved_context": ["<verbatim chunks, if the app retrieves>"],
  "turns": [{ "role": "user", "content": "<verbatim>" }],
  "response_body_raw": "<unrendered>",
  "harness": { "name": "<tool>", "version": "<build>", "config_hash": "<digest>" }
}

go deeper

for a junior

Names the obvious gaps — system prompt, decoding settings, prior turns — and knows a screenshot is not a reproduction record.

for a middle

Adds the echoed model identifier, request id and harness version, and explains that raw bodies beat rendered text because rendering loses whitespace and truncation.

for a senior

Frames capture around future failure modes: what would let me distinguish a provider change, an app change and sampling noise months later, and what do I redact before it leaves the harness.

for a principal

Turns the list into a schema every tool the org runs must emit, and weighs storage, retention and harmful-content handling against reproducibility.

### The mistake underneath the log format A prompt-and-reply log encodes an assumption: that the text you sent and the text you got back are the whole causal story. They are not. What answered you was an *application*, not a model. Sitting between your keystrokes and the tokens you read there is an application-owned **system prompt** (the instruction block the product prepends to every conversation), possibly a set of **retrieved documents** pasted in around your text by the app's retrieval layer, a set of **tool schemas** describing functions the model may call, a **decoding configuration** that governs how the next token is picked, a moderation classifier in front and often another behind, and a **served model build** sitting behind an endpoint name that never changes. Every one of those can change while your prompt stays byte-identical, and every one of them changes the reply. ### The four layers a usable capture record carries **1. Identity of what you hit.** The endpoint or deployment string you called, *and* the model identifier the provider echoed back in the response envelope — not the one you believe you configured. A deployment alias can be repointed behind you; the echoed value is the one that dates the finding. Where the provider also returns a backend or build fingerprint alongside the model id, keep it: it is the finest-grained "what served this" signal you will ever get from outside. **2. Request configuration as it went on the wire.** Effective decoding parameters — the temperature, top-p, maximum output tokens and stop sequences that were actually transmitted — plus tool schemas, any safety or moderation options you set, and the SDK and harness build that assembled the request. Record the *effective* values, not the ones present in your config file: a harness that silently defaults a temperature you never wrote is the single most common reason a re-run diverges. **3. Conversation state.** The system prompt verbatim, the retrieved chunks if the app does retrieval, and every prior turn in order, both sides. Multi-turn attacks work because of accumulated state; a prompt lifted out of the middle of a conversation usually does nothing on its own. **4. Provenance.** Timestamp with timezone, the provider's request id, and the operator. These look like bookkeeping until reproduction fails, at which point they are the only handles that let you ask the provider or the platform team what was serving that traffic. ### What the capture costs Next to nothing while the run is live, and it is unrecoverable afterwards — that asymmetry is the entire argument. Concretely it is two extra fields read off each response envelope, one system-prompt snapshot per run, and storing raw bodies instead of rendered text. No extra model calls, no extra spend at the endpoint. The real costs sit elsewhere: a sweep of a few thousand attempts with raw request and response bodies is on the order of tens of megabytes — trivial in bytes, not trivial in liability, because that store now holds working attack transcripts against your own systems plus whatever real user data arrived through production-shaped retrieval. Budget the access control and the retention clock, not the logging. ### Where the record misleads when you skimp The characteristic failure is not "we lost data", it is a **wrong conclusion stated confidently**. Someone re-runs the attempt from a prompt-and-reply log, gets a refusal, and writes *fixed*. Absence of the behaviour today is one sample from a system with at least three independently moving parts — the served build, the app's own prompt and filters, and sampling variance — and a log with no echoed model id and no system prompt gives you nothing to tell them apart with. The second misreading is subtler: a **rendered transcript read as raw output**. Rendering strips leading whitespace, refusal prefixes, truncation markers and structured response fields. A transcript that shows a harmful continuation but hides a finish reason of "stopped by content filter" reads as a complete success when it was a truncated partial one, and the severity you file off that reading is wrong in the direction that embarrasses you. ### What you would check before believing your own record Take one attempt at random and rebuild the wire request from the stored fields alone, on a different machine, without the harness. If you cannot reconstruct it, what you have is a diary, not a capture. Once, compare the stored decoding settings against what the SDK actually transmitted — that is how you find the default nobody wrote down. And scan the store for auth headers and customer data before any record leaves your team.

  • Why record the model identifier the response returned rather than the one in your config?
    A deployment alias can be repointed behind you. The echoed identifier is evidence of what actually served the request; your config only records what you asked for.
  • Is a screenshot of the successful exchange acceptable evidence?
    As a supplement only. It carries no decoding settings, no system prompt, no request id, and it hides whitespace and truncation. Attach the raw request and response bodies alongside it.
  • What do you redact before the record leaves your harness?
    API keys and auth headers, and any real customer data that reached the transcript through production-shaped retrieval or logs. Keep the attack text itself intact, since removing it destroys the reproduction.

saying these in an interview costs you the question

  • Says the prompt is enough because 'the model is the model'.
  • Offers a screenshot or a copy-pasted chat bubble as the reproduction artefact.
  • Records the configured model name only and never checks what the response echoed.
  • Omits the system prompt because 'that's the app team's, not part of the attack'.
  • Stores raw auth headers or live customer data in the ticket for the sake of fidelity.

context

open as a page

In a finding against a hosted chat endpoint you write 'set temperature to 0' as the reproduction instruction. Why is that not the same as recording a seed, and what should the finding say about determinism instead?

level: middleimportance: must knowfreq 58%

basics

~20 s

Temperature 0 records what you sent, not a guarantee you get the same tokens back. A hosted endpoint gives no seed you control and contracts no bit-identical output. State the settings you used, how many attempts you ran, and how many succeeded, so a reader re-runs the attempt the same way you did.

open as a page

Your multi-turn red-team harness generated each follow-up prompt with an attacker model rather than reading a fixed script, and it succeeded against a hosted chat endpoint. What goes in the finding so the result can be re-run, given that re-running the harness produces different prompts every time?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Ship the realised conversation, not the generator. Store every turn verbatim in order as a fixed replay script a reader can send as-is, plus the target's decoding settings and system prompt. Record the attacker-model configuration separately, as provenance for how the conversation was found, not as the reproduction steps.

open as a page

A red-team finding you filed three weeks ago against a hosted chat endpoint no longer reproduces. What captured from the original run would let you distinguish a silent provider-side model change from a shipped fix, and what do you do if you captured none of it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Compare the model identifier the response echoed then and now, the application's system prompt then and now, and your request ids and timestamps. A changed identifier or prompt points at drift; both unchanged points at a fix or variance. With none captured, you cannot attribute it — re-run, capture properly this time, and say so.

open as a page

You lead a team running several different AI red-team tools, each writing its own log format. Define the capture standard that keeps any finding re-runnable a quarter later against hosted targets: what is mandatory, and what do you deliberately not store?

level: principalimportance: should knowfreq 30%

basics

~20 s

Mandate a small common envelope every tool must emit: target identity and echoed model id, effective decoding settings, system-prompt digest, verbatim turns, attempts and successes, timestamps and request ids, harness build. Deliberately skip credentials, real customer data, and full bodies for non-hit attempts beyond a sampled retention window.

open as a page