What has to be recorded about a drafting run for a generated case to be explainable months later?
answer
- The code says what, the record says why
- Inputs and their revisions, not the chat log
- Same inputs need not give the same output
- One identifier from the case to its record
basics
~20 sRecord the inputs, not the conversation: the request text, every attached artefact with its revision, which build and settings produced the draft, and the untouched output before human edits. Link the merged case to that record.
solid answer
~50 sMonths later the question asked of a generated case is never *what does it do* - the code answers that - but **why does it assert this**. Answering needs the inputs that produced the draft: the request text; each attached artefact with its revision, so you know which criteria and which rulebook were in force; the model build and sampling settings; and the raw output before anyone edited it. Store that as a record with an identifier, and put the identifier in the case's own metadata so the trail is one lookup rather than an archaeology exercise. Note what this does *not* buy. With sampling above zero the same inputs need not reproduce the same output, and the build shifts underneath you, so the goal is **explainability and provenance**, not byte-identical replay. What you can recover is what the drafter knew, which is what anyone actually asks about.
code
json · 14 lines{
"draft_id": "d-2026-03-11-0417",
"request": "Draft a service-level case for the invite-expiry criterion.",
"attachments": [
{ "kind": "criteria", "ref": "requirements/invites@rev-31" },
{ "kind": "sample", "ref": "samples/invite-record@rev-4" },
{ "kind": "rulebook", "ref": "conventions/case-rules@rev-14" },
{ "kind": "exemplar", "ref": "cases/account/rejects_expired_invite@a91f0c2" }
],
"drafter": { "build": "drafting-service-2026.02", "randomness": 0.2 },
"raw_output_ref": "drafts/d-2026-03-11-0417.txt",
"human_edits": "renamed the case; replaced a fixed pause with a condition wait",
"merged_as": "cases/account/rejects_expired_invite"
}go deeper
Be ready to say that a generated case still needs an account of where it came from. Knowing that the request and the artefacts attached to it matter more than the chat log is enough at this level.
Explain the fields worth recording - request text, each attachment with its revision, the build and its settings, the raw output before editing - and why a pinned revision matters when a criteria document keeps changing under the same name.
An interviewer expects the honest limit: sampling and shifting builds make replay unavailable, so aim at explainability instead. Show how the case points at its record, who follows that pointer, and why capture has to be automatic to survive.
Own the policy. What provenance your domain genuinely requires for generated artefacts, how long records live and what they cost to keep, and how to avoid a scheme so heavy that people quietly stop recording anything at all.
## The question an old generated case provokes A year after it was merged, nobody asks what a generated case does - the code says that plainly enough. The question is always some form of **why does it assert this**. A colleague wants to delete it and needs to know whether the assertion encodes a requirement or a guess. A behaviour is about to change and someone must decide whether the case is protecting intent or protecting an accident. A batch of similar drafts turns out to be wrong the same way, and the useful question is what they shared in their input. None of those are answerable from the code, and none are answerable from the person who merged it, who has forgotten or left. They are answerable only from the **inputs that produced the draft**. That is what a drafting record is for: it exists so a machine-drafted case is no more mysterious a year later than a hand-written one whose author wrote a decent commit message. ## What to record Record the inputs and their versions, not the conversation: 1. **The request text**, exactly as sent - the words that framed what was wanted. 2. **Each attached artefact with its revision.** This is the field people skip and the one that matters most: which acceptance-criteria document, at which revision; which redacted sample record; which rulebook; which exemplar cases, at which commit. A reference without a revision points at today's text, not the text that was in force. 3. **The drafter identity and its settings** - which build produced it and how much randomness was allowed. Behaviour changes across builds, and a family of odd drafts often shares one. 4. **The raw output**, before any human edit. Without it you cannot separate what the machine produced from what a person then repaired, which is precisely the distinction a later reader needs. 5. **The human edits**, as a short note or a difference against the raw output. 6. **The identifier of the case as merged**, so the link runs in both directions. ## Provenance is not replay This is the part that is most often misunderstood, and the misunderstanding leads people to abandon recording altogether when the first rerun disagrees. | Question | Does the record answer it? | |---|---| | What was the drafter given, and at which revisions? | Yes - this is its whole purpose | | What did the drafter produce before human editing? | Yes, if the raw output was stored | | Which parts were repaired by a person? | Yes, from the recorded edits | | Will the same request reproduce the same case? | No - sampling and build drift make that unavailable | | Was the resulting case correct? | No - provenance never certifies correctness | Sampling makes generation nondeterministic, and the build moves under you regardless. Byte-identical replay was never on offer, so its absence is not a defect in the record. Aim at **explainability**: a reader can reconstruct what the drafter knew and therefore judge whether the assertion had a basis. That is the question people actually bring. ## Where the record lives, and how the case points at it A record nobody can find is not provenance. Two properties make it usable. First, the record has a stable identifier and the merged case carries that identifier in its own metadata or header comment, so the trail is one lookup from the file in front of you. Second, the record is stored somewhere with the same lifetime as the code; a link into a chat product's history is a link that outlives nothing. The direction matters. Searching a drafting archive for a case is hopeless; following a pointer from a case you are already reading is trivial. Put the pointer on the case. ## What not to record, and why transcripts disappoint The tempting shortcut is to keep the whole conversation. Transcripts are long, awkward to search, and mix the inputs that mattered with dead ends - but the real weakness is subtler. A transcript contains context that was **pasted**, and pasted text carries no provenance of its own. You can read the criteria that were supplied and still have no way to know which revision of the criteria document they came from, or whether the sample record was the checked one. The transcript preserves the words and loses the versions, which is exactly backwards. Keep transcripts if they are cheap, but do not mistake them for a record. What you need is references with revisions. ## Who reaches for it Three groups, and it is worth knowing which one you are building for. Engineers changing or deleting an old case, who need to know whether the assertion has a basis; reviewers investigating a cluster of similar bad drafts, looking for the shared input that caused them - this is the use that pays for the whole scheme, because a bad rulebook rule or a stale criteria revision explains dozens of cases at once; and, in regulated work, anyone required to show how a control was produced. Size the scheme to those three. A record that takes ten minutes to fill in by hand will not be filled in, and an empty archive explains nothing. Capture it automatically from the drafting step or accept that you will not have it.
- Why not simply keep the whole conversation transcript?A transcript preserves the words and loses the versions. Pasted context carries no provenance of its own, so you can read the criteria that were supplied and still not know which revision they came from or whether the sample was the checked one. Record references with revisions; keep the transcript only if it is free.
- The same inputs produce a different draft on rerun. Has the record failed?No, because it was never a replay guarantee. Sampling makes generation nondeterministic and the build moves underneath you. The record says what the drafter was given and what it produced at the time, which is what explains the case. Treat identical replay as unavailable rather than broken.
- Who actually reaches for one of these records?Someone changing or deleting an old case who needs to know whether its assertion has a basis; a reviewer chasing a cluster of similar bad drafts and looking for the shared input, which is the use that pays for the scheme; and, in regulated work, anyone asked to show how a control was produced.
saying these in an interview costs you the question
- Keeps the chat transcript and calls it provenance
- Expects identical output from identical inputs on rerun
- Records the request but not each attachment's revision
- Stores only the edited case, never the raw output
- Leaves no link from the merged case to its record