skip to content

You lead a team running several different AI red-team tools, each writing its own log format. Define the capture standard that keeps any finding re-runnable a quarter later against hosted targets: what is mandatory, and what do you deliberately not store?

level: principalimportance: should knowfreq 30%

answer

  1. standardise the envelope, adapter per tool
  2. echoed model id, effective settings, prompt digest
  3. attempts and successes, request ids, harness build
  4. no credentials, no real user data, sample the misses
  5. retention clock + adapter as only write path

basics

~20 s

Mandate a small common envelope every tool must emit: target identity and echoed model id, effective decoding settings, system-prompt digest, verbatim turns, attempts and successes, timestamps and request ids, harness build. Deliberately skip credentials, real customer data, and full bodies for non-hit attempts beyond a sampled retention window.

solid answer

~50 s

Standardise the *envelope*, not the tools. Each scanner, eval framework and multi-turn harness keeps its native output; a thin adapter emits a common record per attempt. **Mandatory:** target identity (endpoint plus the model identifier echoed back), effective decoding settings as sent, a digest of the system prompt with the prompt stored once per run, verbatim turns for any attempt that counted, the criterion that decided it counted, attempts and successes per configuration, timestamps and provider request ids, and the harness build and config digest. **Deliberately not stored:** auth headers and keys; real customer data pulled in through production-shaped retrieval; and full raw bodies for non-hit attempts past a short window — keep counts and a sample instead, or the store becomes unaffordable and becomes a long-lived corpus of generated harmful content. The hard part is enforcement: make the adapter the only write path into the findings store, and set retention and access rules before the first quarter of data lands.

go deeper

for a junior

Lists fields worth capturing but does not address enforcement, cost or retention.

for a middle

Proposes a shared schema and per-tool adapters, and knows credentials and customer data stay out.

for a senior

Adds the omissions as deliberate trades — sampled misses, short retention for sweeps — and names the enforcement point.

for a principal

Owns envelope versioning, the adapter as the only write path, retention tied to the fix lifecycle, access control on a store of working attacks, and a periodic drill that re-runs an old finding from the record alone.

### A standard is a record schema plus a write path, not a memo The failure being prevented is specific: a finding filed by one tool, six weeks later, that nobody can re-run because that tool's log kept a different subset of the world than the tool beside it. You do not fix that by asking teams to be thorough. You fix it by defining one record and making it the only way into the findings store. **Standardise the envelope, not the tools.** Each scanner, evaluation framework and multi-turn harness keeps its native output and its native strengths; a thin adapter per tool emits the common per-attempt record. Trying to make heterogeneous tools agree on a log format is a project that never finishes; adapters are small and independently replaceable. ### The mandatory envelope - **Target identity** — endpoint or deployment string, the model identifier echoed in the response, and any provider build fingerprint. - **Effective decoding settings** as transmitted, not as configured. - **System-prompt digest** on every attempt, with the prompt body stored once per run. The digest is a cheap per-attempt change signal: comparing digests across attempts shows a silent app-side prompt change without storing a copy each time. - **Verbatim turns** for any attempt that counted as a hit, plus mid-conversation tool and retrieval results. - **The criterion that decided it counted**, and that criterion's version. - **Attempts and successes per configuration**, so every rate downstream has a denominator. - **Timestamps with timezone and provider request ids.** - **Harness name, build and config digest.** - **Envelope version**, in every record. You will add fields, and a quarter-old record with no version is unreadable at exactly the moment you need it. A field a tool genuinely cannot supply is written as an explicit **null with a reason code**, never omitted. An omitted field is indistinguishable from a field the tool forgot. ### What you deliberately do not store Credentials and auth headers. Real customer data that entered through production-shaped retrieval — which means the standard must say how a run against production-shaped data is handled *before* the run, not after. And full raw bodies for non-hit attempts past a short sampling window: keep counts, the scorer's verdict and a sample. These omissions are trades, each buying a cost or liability reduction against a reproduction benefit you can estimate. ### What it costs Be honest with the team about the two real line items. **Adapters**: a few engineer-days each, plus a maintenance tail every time a tool changes its output shape, and an onboarding tax that shows up later as "we can't use that new tool yet". **Custody**: the store now holds working attack transcripts against your own systems, which makes it the most attractive single artefact in your estate — it needs its own access control, a named owner, and a retention clock tied to the fix lifecycle (hits live while the item is open plus a verification window; sweep detail expires far sooner). Storage itself is the cheap part: hits are a small fraction of attempts, and a quarter of raw bodies across a handful of tools is measured in hundreds of megabytes. Fund the adapters and the custody, or the standard decays into a schema nobody enforces. ### Where the number misleads The metric everyone reaches for is compliance — "98 per cent of findings emitted a conforming record" — and it is close to meaningless. A record can conform and be unusable if half its mandatory fields are nulls. The honest dashboard is the **null-reason report**: which field, which tool, how often. That is a tooling-gap list, and repeated nulls in one field mean either your standard or your tooling has to move. Two related traps. Coverage measured over *tools onboarded* rather than over *findings filed* hides the one un-adapted tool that produces most of your volume. And a green schema-validation job proves the record parses, not that anyone can re-run from it. ### What you would check The falsifiable test is a drill: draw a closed finding from the previous quarter at random, hand the record alone to an engineer who did not run it, and see whether they can rebuild the request and get a yes or no. Publish the pass rate; that number paired with the null-reason report is what tells you the standard works. Failures indict the standard, not the person who tried. Then verify the two things a drill will not catch: attempt to write into the findings store around the adapter — if you can, your enforcement point is decorative — and confirm the retention job actually deletes, because expired-but-present sweep bodies are the usual finding.

  • A tool cannot report the model identifier the endpoint echoed. What does the standard do?
    Accept an explicit null with a reason and surface it in a report. Repeated nulls in the same field are a tooling gap to fix, not a record to quietly accept forever.
  • How do you keep the findings store from becoming a liability?
    Its own access control, a retention clock tied to the fix lifecycle, and no long-lived raw bodies for non-hit attempts. It holds working attacks against your own systems.
  • How would you test that the standard actually works?
    Pick a closed finding from the previous quarter at random and have someone re-run it from the record alone. Failures indict the standard, not the person.

saying these in an interview costs you the question

  • Tries to make every tool emit the same native log instead of adapting to a common envelope.
  • Stores everything forever 'for completeness', including credentials and customer data.
  • Has no versioning on the record, so a quarter-old entry cannot be read after fields change.
  • Treats the findings store as ordinary logs with no access control or retention clock.
  • Defines the standard but leaves each tool free to file without it.

context