Who owns a few-shot exemplar pool, and when must it be refreshed?
answer
- exemplars are policy encoded as data
- every example needs provenance
- real records ship PII on every call
- domain team owns the judgments
- refresh on events, not just the calendar
basics
~20 sAn exemplar pool is policy encoded as data, so it needs a named owner in the domain team, provenance and PII review on every entry, versioning alongside the prompt, and a refresh triggered by policy changes, label-set changes and drift — not by the calendar alone.
solid answer
~50 sTreat the pool as a governed artifact, not a constant in a source file. Three properties matter. **Provenance**: every exemplar records where it came from, who labeled it, and when — because exemplars drawn from real records carry personal data that would otherwise be shipped in every request, and because an unattributed example cannot be re-adjudicated later. **Ownership**: the domain team that owns the policy owns the pool, since the exemplars encode that policy more concretely than the written instruction does; engineering owns the harness, not the judgments. **Refresh triggers**: a policy change, a new or retired label, a model swap, or measured drift — each of which invalidates specific exemplars. When a policy change invalidated four of twelve exemplars in one pool, the correct response was to replace those four, bump the pool version, and re-run the evaluation suite before shipping — not to leave stale examples in place trusting the instruction to override them.
go deeper
Know that examples live in a file someone owns, that they can go stale when the rules change, and that copying real customer records into a prompt raises a privacy question.
Explain that exemplars are version-controlled alongside the prompt, that any pool change should re-run the evaluation set, and that a policy change means replacing the affected examples rather than adding to them.
Demonstrate operating discipline: provenance metadata that makes re-adjudication possible, redaction of real records, event-driven refresh triggers, and per-class evaluation gating every pool change.
Own the accountability question — who signs off on the judgments the pool encodes, how much process the stakes justify, and how to keep the pool from becoming an unreviewed shadow policy that contradicts the written rule.
## Why this is a governance question at all An exemplar pool looks like configuration and behaves like policy. The written instruction states the rule in the abstract; the exemplars show what the rule *means* on concrete cases, and where the two disagree, the concrete examples usually win. That makes the pool a de-facto policy document — one that is frequently checked into a repository, edited by whoever is closest, and never reviewed by the people who own the underlying rules. Interviewers at senior and lead level ask about it because it separates candidates who have shipped an LLM feature from those who have run one for a year. ## Provenance: know where every example came from Each exemplar should carry, at minimum: its source, the labeler, the label date, and the policy version it was labeled under. This is not bureaucracy — each field has a job. - **Source** tells you whether the example is a real record, a synthesized one, or a hand-written illustration. Real records are the most representative and the most dangerous. - **Labeler and date** let you re-adjudicate. When the rule changes, you need to find every example labeled under the old rule; without a date and a policy version that search is guesswork. - **Policy version** is what makes refresh tractable at all. ## The PII problem is specific and under-appreciated Exemplars drawn from production records are not like training data sitting in a warehouse — they are transmitted in **every single request**, to the provider, for the life of the prompt. A customer's details embedded in a demonstration are sent thousands of times a day. The controls are ordinary once you see the shape of the risk: redact or synthesize identifying fields, review the pool as you would review any data leaving your perimeter, and record consent or legal basis where the domain requires it. Synthesized-but-realistic exemplars are often the right answer, at the cost of some representativeness. ## Ownership: the domain team, with engineering support The judgments in the pool are domain judgments. If a clinical, legal or compliance team owns the policy, they own which cases exemplify it, and their sign-off should gate changes the same way it gates the written rule. Engineering owns the mechanics — the format, the versioning, the evaluation harness, the deployment. The failure mode when this is unclear is an engineer changing an exemplar's label to fix a failing test, quietly amending policy through a code review nobody from the domain team read. ## Refresh: event-driven, with a periodic backstop Calendar-based review alone is too slow for the changes that matter and too noisy for the ones that do not. The triggers worth wiring up: - **Policy change.** The strongest trigger. When a rule changes, some exemplars now demonstrate the *old* rule and actively teach the wrong answer. In one pool, a policy change invalidated four of twelve exemplars — the right response was to replace those four, bump the pool version, and re-run the evaluation suite before shipping. - **Label-set change.** A new class needs demonstrations or it will barely be predicted; a retired class needs its examples removed or it will keep being predicted. - **Input drift.** New formats, new sources, new phrasing mean the pool no longer spans the traffic. Detected by monitoring per-segment quality, not by intuition. - **Model change.** A different model may need a different number and mix of examples; the pool that was tuned for the old one is an assumption to re-test. - **A periodic backstop.** A quarterly review catches the slow erosion none of the above fires on, and is the moment to check that provenance metadata is still accurate. ## Versioning and evaluation go together A pool version should be pinned to a prompt version, and every pool change should re-run the same evaluation suite, reported per class and per segment. Without that, a pool edit is an unmeasured production change to a system's decision-making. It is also what makes rollback meaningful: reverting a bad exemplar swap should be one pinned-version change, not an archaeological exercise in the commit history. ## The contested part How much process a pool deserves is genuinely contested, and the honest answer is that it scales with consequence. A pool behind an internal drafting assistant does not need domain-team sign-off; a pool behind a regulated screening decision does. The signal to watch is whether anyone can answer "why is this example labeled this way?" — when the answer is "it was there when I joined", the pool has drifted out of governance regardless of what the process document says.
- A policy change invalidates four of twelve exemplars. What exactly do you do?Identify them by the policy version recorded on each exemplar rather than by reading through the pool. Replace those four with cases labeled under the new rule and adjudicated by the domain owner, bump the pool version, and re-run the full evaluation suite reported per class before shipping. Leaving them in place is the worst option: concrete demonstrations of the old rule generally beat a revised instruction that contradicts them.
- How do you decide between real production records and synthesized exemplars?Trade representativeness against exposure. Real records capture the messiness of actual input, which is what makes demonstrations effective, but they ship whatever they contain in every request for the prompt's lifetime. Synthesized or heavily redacted examples remove that exposure and cost some fidelity. In regulated or personal-data-heavy domains, synthesize by default and validate that the synthetic cases still reproduce the failure modes the real ones did.
- What is the smallest useful governance for a low-stakes internal prompt?A version-controlled pool file with a named owner, a short provenance note per exemplar, and an evaluation set that runs on change. That is enough to answer why an example is labeled the way it is and to roll back a bad edit. Heavier controls — domain sign-off, formal PII review, scheduled re-adjudication — are earned by consequence, and applying them everywhere mostly teaches teams to route around the process.
saying these in an interview costs you the question
- Treats the exemplar pool as a constant nobody needs to own
- Ships production records as exemplars without redaction review
- Leaves outdated examples in, trusting the instruction to override
- Changes an exemplar's label to make a failing test pass
- Reviews the pool only on a calendar, never on policy change