How would you manage production prompts as versioned artifacts across a team?
answer
- prompts are behaviour-defining source
- which text was live at 14:00?
- evidence attached to every change
- defensive lines need owning tests
- who may click publish is contested
basics
~20 sTreat each prompt as code: a reviewed file under version control, an identifier recorded with every production response, an eval run attached to each change, a named owner, and a rollback that is one deploy away.
solid answer
~50 sPrompts are behaviour-defining source, so they get the same machinery as source. Concretely: each prompt lives in version control as its own diff-able file, not embedded in a string literal or hand-edited in a console; every change goes through review with its eval delta attached; each version has an identifier that the application records alongside every generation, so an incident can be traced to the exact text in force; and each prompt has a named owner who understands why its lines exist. The part teams underestimate is **scar tissue** — defensive lines added after past incidents that nobody dares remove because nobody knows what they protect. The cure is to pair each such line with an eval case, so deletion is safe when the suite stays green and blocked when it does not. Rollback then means re-pointing at a previous version, with the same audit as any deploy.
go deeper
Know that prompts belong in version control like code, with a review before changes go live, rather than being edited directly in a running system.
Explain the mechanics you would put in place: a diff-able file per prompt, a version identifier logged with every response, an eval run attached to each change, and a rollback path.
Show the operational angles — tracing a customer complaint back to the exact prompt version, keeping output-contract changes and parser changes together, keying caches by prompt version, and pairing each defensive line with an eval case so it can eventually be removed.
Own the organizational tradeoff: who may publish a prompt, whether the source of truth is the repository or a registry, how domain experts get a fast loop without bypassing the eval gate, and how you keep prompts from accreting instructions nobody can justify.
## The failure this prevents The common shape: a prompt starts as a string in application code, gets moved into a database so non-engineers can edit it, and within a year nobody can answer three questions — what text was live at 14:00 last Tuesday, why the fourth paragraph exists, and who is allowed to change it. Meanwhile the prompt directly determines customer-visible behaviour, more so than most of the code around it. Managing prompts as versioned artifacts is the discipline that keeps those three questions answerable. ## The artifact A prompt should be a file, in the repository, on its own, so that a diff shows exactly which sentence changed. Templating variables belong in the file too; assembling the final text from four scattered fragments defeats the review. The version identifier can be a content hash, a semantic version, or a monotonic number — what matters is that it is emitted with the artifact and recorded in the request path. Alongside the text, a prompt benefits from a small header or sidecar recording: the owner, the model version it was tuned for, the eval set and score at its last change, and a changelog line per revision explaining *why* — not "improved wording" but "added escrow tie-break after INC-412". The why is what future engineers cannot reconstruct from the diff. ## Change control Every prompt change should be reviewable and gated exactly like code: - **Review by someone other than the author**, because prompts fail in ways that read fine to the person who wrote them. - **An eval run attached to the change**, showing the delta against the current baseline and which items flipped. A prompt change proposed without an eval number is a change proposed without evidence. - **A blocking gate** on the primary metric and on any case derived from a past incident. - **A recorded baseline update** when the change is accepted, so the next comparison starts from a documented point. ## Observability and rollback Every model response logged in production should carry the prompt version and the model version that produced it. Without that pairing, incident review is guesswork: quality complaints arrive days later, by which time the prompt may have changed twice. With it, you filter the affected window by version and immediately know whether the complaint predates or postdates a specific change. Rollback should be a deliberate, fast path — revert to a previous version identifier, redeploy or flip a config value, and record it. Two caveats are worth raising unprompted. First, if downstream code parses a structured output whose shape changed between versions, rolling the prompt back without rolling the parser back reintroduces a different break; version them together or keep the contract stable across prompt versions. Second, if responses are cached, rolling back the prompt does not roll back the cache, so cache keys must include the prompt version. ## Scar tissue and the right to delete Mature prompts accumulate defensive lines added after specific incidents. Each was justified once; collectively they bloat the prompt, and nobody removes them because removal is unbounded risk. This is the strongest practical argument for pairing evaluation with versioning: **every defensive line should arrive with an eval case that fails when the line is removed.** With that pairing, a later engineer can delete the line, run the suite, and either learn immediately that it still matters or remove it with confidence. Without it, the prompt only grows, and the accumulated instructions eventually start interfering with each other. ## Who may edit, and where The hard organizational question is that the people with the best judgment about prompt wording — support leads, clinicians, legal reviewers — are often not the people with repository access. There are two defensible answers and the tradeoff is genuinely contested as of mid-2026. - **Repo-only editing** keeps a single source of truth, real review, and atomic rollback, at the cost of a slow loop for domain experts who must route every wording change through an engineer. - **A prompt registry or management platform** lets domain experts edit and publish directly, with versions, staged rollout and audit built in. The risk is a second source of truth that drifts from the repo, and edits reaching production without the eval gate. The workable middle is a registry that is fed from version control, or one that enforces the same gates — review, eval run, staged rollout, instant rollback — regardless of who clicks publish. What is not defensible is a production prompt that any authenticated user can edit in a text box with no version history and no eval, which is where a surprising number of systems actually are. ## What an interviewer is listening for Not a tool recommendation. They want to hear that you treat prompts as behaviour-defining artifacts with owners, history, evidence attached to changes, traceability from a production response back to the exact text, and a rollback path — and that you can articulate the editing-access tradeoff without pretending it has a clean answer.
- Why should every logged model response record the prompt version that produced it?Because quality complaints arrive after the fact, and without the pairing you cannot tell which text was in force when the bad output was generated. With prompt and model versions on each response, you filter the affected window, confirm whether a specific change caused the issue, and measure the blast radius before deciding whether to roll back. It also lets you compare cohorts during a staged rollout.
- Non-engineers need to change wording weekly. How do you give them that without losing control?Give them an editing surface that enforces the same gates as the repository: versioned drafts, a required eval run against the held-out set, a review step, staged rollout to a fraction of traffic, and one-click rollback. The failure mode to avoid is a second uncontrolled source of truth — either the registry is fed from version control, or version control is regenerated from the registry, but never two independent copies drifting apart.
- What breaks when you roll a prompt back but nothing else?Two things commonly. Downstream parsers tied to an output shape that changed between versions will now fail against the restored format, so the contract should either stay stable across versions or be versioned together with the prompt. And response caches keyed without the prompt version will keep serving output generated by the version you just rolled back, which makes the rollback look ineffective.
saying these in an interview costs you the question
- Keeps production prompts in an editable text box with no history
- Cannot say which prompt text produced a given past response
- Merges prompt changes with no eval evidence attached
- Leaves defensive lines forever because nobody knows what they protect
- Rolls back the prompt while a cache still serves the old outputs