skip to content

A retrieved-exemplar store keeps teaching a product name retired six months ago — how do you fix it?

level: seniorimportance: should knowfreq 35%

answer

  1. the pool is data, not prompt text
  2. stale examples surface only near their neighbourhood
  3. retire by validity window, do not delete
  4. unverified write-back trains the pool on itself
  5. pin a snapshot so evals stay honest

basics

~20 s

Treat the exemplar pool as a versioned production dataset, not a folder of examples. Give every entry a timestamp and a validity window, filter expired entries out at retrieval, verify anything written back before it becomes teachable, and audit the pool against the current taxonomy on a schedule.

solid answer

~50 s

Once demonstrations are retrieved rather than hardcoded, the pool is a live dependency that silently shapes output — and stale entries are unusually hard to notice, because an obsolete example only surfaces for the queries that land near it. So the regression is localised and invisible in aggregate metrics. The fix has four parts. Give every entry provenance and time metadata — created-at, verified-by, and a validity window — so retirement is a metadata write, not a deletion. Filter on that window at retrieval time and optionally weight by recency, so old cases lose to newer ones at equal similarity. Gate the write-back path: verified outcomes are admissible, raw model output is not, or the pool slowly retrains itself on its own mistakes. Finally, version the pool and pin evaluations to a snapshot, so a quality change can be attributed to a store update rather than blamed on the model.

go deeper

for a junior

Know that retrieved demonstrations come from a store someone has to maintain, and that an out-of-date example keeps teaching the old answer for as long as it stays retrievable.

for a middle

Explain the metadata that makes maintenance possible — created-at, verifier, validity window — and how retrieval-time filtering and recency weighting turn that metadata into behaviour. Say why deletion is the worse tool.

for a senior

Show you would treat the pool as a versioned production dataset: gated write-back, scheduled reconciliation against the live taxonomy, telemetry on the age of retrieved entries, pool version logged per request, and evaluations pinned to a snapshot.

for a principal

Own who is accountable for the pool. Decide whether it is engineering-owned or operations-owned, what the admission bar and review cadence are, and how a bad batch is rolled back — a demonstration store with no owner degrades silently and steers every request.

## Why staleness bites harder here than in a static prompt When demonstrations are hardcoded in a template, a retired product name is in a file someone reads during review, appears in every request, and shows up the moment anyone looks at the prompt. When demonstrations are retrieved, the same stale example is dormant for most traffic and activates only for the queries near it. The result is a regression that is concentrated, intermittent, and absent from an aggregate accuracy number — which is why teams find it through a customer complaint rather than a dashboard. The deeper point an interviewer is testing: adding retrieval to few-shot prompting converts prompt content into *data*, and data needs a lifecycle. Nobody would ship a training set with no timestamps, no provenance and no review path; an exemplar pool is that training set, applied at inference time. ## Metadata is the whole fix Every entry should carry, at minimum: when it was created, who or what verified its label, and the window over which it is considered valid. That last field is what lets you retire the obsolete product without deleting history — you set a validity end date, and the retrieval filter stops returning it. Deletion loses the audit trail and makes it impossible to answer "why did this system say that in March?" With the metadata in place, retrieval becomes a filtered search rather than a raw nearest-neighbour lookup: drop entries outside their validity window, then rank. Recency weighting is a softer version of the same idea — at comparable similarity, prefer the newer case — which handles gradual convention drift where nothing is strictly wrong but older phrasing is worse. ## Curating what goes in Write-back is what makes a pool valuable and what makes it dangerous. The valuable version: when a case is resolved by a human and the outcome is verified, the input and the verified outcome enter the pool. The dangerous version: the system's own output is written back automatically. The second creates a closed loop — the model's mistakes become demonstrations, the demonstrations make the mistake more likely, and the pool converges on its own errors while looking busy and well-maintained. If you must auto-write, mark provenance as unverified, keep those entries out of retrieval until reviewed, and never let an unverified entry become the nearest neighbour for a class of queries. A second admission rule is worth stating: an entry earns its place by being *teachable*, not merely by being correct. Correct-but-idiosyncratic cases teach idiosyncrasy. ## Auditing on a schedule Taxonomies move: queues get merged, categories renamed, products retired, policies rewritten. None of those events touch the pool unless someone makes them. Two mechanisms keep the pool honest. First, a periodic reconciliation against the current label set and product vocabulary — anything referencing a label or entity that no longer exists is flagged for retirement or relabelling. This can be largely automated because it is a set difference, not a judgement call. Second, coverage and staleness telemetry: the age distribution of the entries actually being retrieved, and the share of retrievals hitting entries older than some threshold. A rising median age of retrieved exemplars means the pool is being outrun by the traffic. It is the same shape of signal as a falling similarity score, and it deserves the same alert. ## Versioning and attribution The pool must be versioned, and every request should log the pool version alongside the retrieved ids. Without that, every quality question becomes unanswerable: a change in output could be a model update, a prompt change, or someone adding forty entries to the pool on Tuesday. With it, you can diff versions, roll back a bad batch, and pin an evaluation to a snapshot so your regression suite is measuring the model rather than the drift of its demonstrations. Pinning also protects against the quieter contamination problem: if production inputs are written back and your evaluation inputs came from production, the pool can eventually contain the eval items themselves, turning the evaluation into a lookup. ## The answer in one line Retire the obsolete entries by validity window rather than deletion, filter and recency-weight at retrieval, gate write-back behind human verification, reconcile against the live taxonomy on a schedule, and version the pool so you can attribute and roll back. The product name is a symptom; the missing lifecycle is the defect.

  • Why not simply delete the obsolete entries?
    Deletion destroys the audit trail and makes past behaviour unexplainable — you can no longer show why the system answered as it did last quarter. A validity window achieves the same retrieval outcome, keeps history, allows a mistaken retirement to be reversed with a metadata write, and lets you reconstruct any prior pool version for debugging or evaluation.
  • What goes wrong if you auto-write the model's own accepted outputs back into the pool?
    You close a feedback loop. Errors that slipped through become demonstrations, demonstrations make the same error more likely, and the pool drifts toward its own mistakes while its size and freshness metrics look healthy. Guard it by writing unverified entries in a quarantined state, keeping them out of retrieval until a human confirms them, and sampling live retrievals for label audit.
  • How would you notice pool staleness before a customer does?
    Track the age distribution of the exemplars actually retrieved, not of the pool as a whole — a rising median age means traffic has moved past your examples. Pair it with the share of retrievals whose top similarity sits below your floor, and a scheduled reconciliation that flags entries referencing labels or entities no longer in the live taxonomy.

saying these in an interview costs you the question

  • Treating the exemplar pool as static prompt text rather than data
  • Writing model output back into the pool without verification
  • Deleting stale entries instead of retiring them with metadata
  • Evaluating against a live pool, so results are not reproducible
  • Assuming an aggregate accuracy metric will surface a stale example

context