Where must a deletion request reach in an LLM app beyond the app database?
answer
- one record, many resting places
- the trace store is the big copy
- evals were mined from production
- memory keeps derived facts too
- some copies cannot be deleted at all
basics
~20 sEvery store that received a copy of the text: trace and log stores, eval and golden datasets mined from production traces, agent memory, vector-index payloads, provider-side retention and caches, and any fine-tuning set. Data already trained into weights cannot be deleted at all.
solid answer
~60 sAn LLM application fans one record out into far more stores than a CRUD app does, and a deletion routine written for the transactional database misses most of them. The checklist: **traces and logs** (the largest and longest-lived copy, often in a third-party observability vendor); **eval and golden datasets**, because production traces get sampled into fixtures that then live in a repo or object store indefinitely; **agent memory** — episodic logs and extracted semantic facts that were written from those conversations; **vector-index payloads**, where the source text is usually stored alongside the embedding; **prompt-cache tiers**, provider-side and your own semantic cache; **provider retention buffers**, deleted by their own clock rather than by your request; and **fine-tuning corpora**. Two of these are not deletable in the ordinary sense. Data absorbed into fine-tuned weights cannot be excised — the remedy is retiring or retraining the model. And a version-controlled eval fixture persists in git history after the file is removed. The engineering answer is a data-flow inventory built when the feature ships, plus a stable subject identifier propagated into every store so deletion is a query rather than an archaeology project.
go deeper
Know that a prompt ends up in more than one place — at minimum the application database and the trace or log store — so deleting a record is not the same as deleting the data.
Enumerate the sinks: traces, eval fixtures, agent memory, vector payloads, caches, provider retention, fine-tuning corpora, and the backups behind each. Explain why the trace store is usually the largest copy.
Demonstrate you have designed for this: a subject identifier propagated into every store so deletion is a query, redaction upstream of the fan-out, and a maintained inventory that a new sink cannot bypass silently.
Own the irreversible cases. Decide in advance which data classes may ever enter a fine-tuning corpus or a committed eval fixture, given that the only remedies there are retiring a model or rewriting shared history.
## Why this is harder than in a normal service In a conventional application, a customer's data lives in a handful of tables plus backups, and deletion is a well-understood cascade. An LLM application copies the same text into a set of stores that were built for debugging, evaluation and performance, each added by a different team at a different time, and several of them run outside your infrastructure. A deletion routine that only touches the transactional database is, in this architecture, close to a no-op. The honest way to answer the question is to trace one request end to end and name each place its content came to rest. ## The inventory **Trace and log stores.** Prompts and completions are captured for debugging, cost attribution and quality review. This is usually the single largest copy of the sensitive text in the system, frequently held by a third-party observability vendor, with retention set by whoever configured the integration and read access granted broadly to engineering. If deletion reaches nothing else, it must reach here. **Eval and golden datasets.** Mature teams mine production traces for eval cases — that is the recommended practice, and it is also a data-propagation path. A trace sampled into a golden set stops being an ephemeral record and becomes a fixture: committed to a repository, copied into CI artefacts, snapshotted alongside experiment results. Removing the file from the working tree does not remove it from git history, and re-writing history on a shared repository is a heavy operation. The mitigation is upstream: fixtures are built from redacted text, or synthesised from real traces rather than copied verbatim. **Agent memory.** If the product remembers users across sessions, there are two layers: an episodic log of what happened, and semantic facts extracted from it. A deletion must clear both, and the extracted facts are the harder half — a fact derived from a conversation no longer contains the conversation, so it is only findable if the write path recorded provenance. This is the practical argument for storing, with every memory record, the subject and source it came from. **Vector-index payloads.** Retrieval systems almost always store the source chunk next to the embedding, because you need the text to put in the context. Deleting the row deletes both. If for some reason only the vector remained, note that embeddings are not anonymised data — inversion research shows meaningful reconstruction of source text from embeddings — so they are in scope too. **Cache tiers.** A semantic cache keyed on request similarity holds request and response bodies. Provider-side prompt caches hold prefixes for a TTL. Neither is addressed by a database delete; the first you must purge, the second expires on its own schedule and you can only account for it. **Provider retention.** Whatever the provider's bounded window is, your deletion request does not shorten it. What you can do is know the number, state it to your reviewers, and ensure the content sitting inside that window was tokenized in the first place — which is exactly why the redaction boundary and the deletion story are the same design problem. **Fine-tuning corpora and weights.** The uploaded corpus is a deletable object. The model trained on it is not: there is no reliable mechanism to remove specific records from weights, so the remedies are coarse — retire the model, or retrain from a corrected dataset. Because this copy is uniquely irreversible, fine-tuning inputs deserve the strictest redaction of anything in the pipeline. **Backups and warehouses.** Every store above has backups, and several are replicated into an analytics warehouse. Backup-cascade policy is a normal engineering problem, but it is now a problem in six systems instead of one. ## Making it tractable Deletion is only cheap if it was designed in. Three moves make the difference: 1. **A stable subject identifier propagated everywhere.** Every trace, memory record, cache entry and eval fixture carries the identifier of the person the content concerns. Deletion becomes a query per store instead of a text search. 2. **Redaction upstream of the fan-out.** If the text that reaches traces, evals and caches is already tokenized, deleting the vault entry breaks the link for every downstream copy at once. This does not discharge a deletion obligation on its own, but it collapses the blast radius from "seven stores hold plaintext" to "seven stores hold surrogates and the map is gone." 3. **A written data-flow inventory, maintained.** New sinks appear constantly — a new tracing vendor, a caching layer, an experiment store. The inventory is only useful if adding a sink is a step that updates it. ## The shadow copy The classic incident is not architectural. Someone enables full prompt capture behind a debug flag to chase a bug, the flag stays on for two weeks, and a second complete copy of every client note lands in a log store nobody included in the review. When the deletion request arrives, the inventory is silently wrong. Treat verbose prompt logging as a change that requires the same review as adding a database. ## What interviewers listen for That you enumerate stores rather than gesture at "we delete the data"; that you know evals and agent memory are propagation paths; that you say plainly which copies cannot be deleted and what you do instead; and that you make the whole thing tractable with a propagated subject identifier and upstream redaction.
- An eval fixture built from a real client conversation was committed to git six months ago. What are your options?Removing the file leaves it in history, and rewriting history on a shared repository is disruptive and often incomplete once forks and CI artefacts exist. Practically you rotate the fixture to a redacted or synthesised version, restrict repository access, and record the residual copy honestly. The real fix is preventive: eval fixtures are built from tokenized text or generated from traces rather than copied verbatim.
- Why is agent memory harder to clear than a trace store?Traces are records of the original text, so they match on content or on a request identifier. Memory holds facts extracted from that text — a summarised preference or a derived attribute that shares no substring with the conversation it came from. Unless the write path stored provenance linking each memory record to its subject and source, those records are effectively unfindable, which is why provenance is a memory-write requirement rather than a nicety.
- Does deleting a vector store row remove the person's data from the retrieval system?It removes the payload and the embedding, which is most of it, but check the rest of the pipeline: many systems keep a document-level copy in object storage, a chunk cache, and a search index alongside the vector index. Also note the embedding itself is not anonymised — inversion techniques recover meaningful source text — so leaving vectors behind after deleting payloads is not a valid halfway position.
saying these in an interview costs you the question
- Deletes from the application database and calls it done
- Forgets eval fixtures mined from production traces
- Assumes a deletion request shortens the provider's retention window
- Claims fine-tuned weights can have specific records removed
- Treats embeddings as anonymised and safe to keep