skip to content

PII & Data-Leakage Handling

You will learn how sensitive data leaks through LLM systems — into prompts, logs, provider retention, and cross-tenant retrieval — and the redaction, isolation, and retention controls that stop it. Interviewers ask this because most real LLM incidents are data-leakage incidents, not jailbreaks.

on this pageshow

questions

4

Why is pasting a full patient record into a chat assistant a leak even without model training?

level: juniorimportance: must knowfreq 70%

answer

  1. prompts are not ephemeral
  2. count the copies you make
  3. history is resent every turn
  4. logs, traces, support tooling, caches
  5. send fields, not documents

basics

~20 s

Text placed in a prompt becomes data your own system holds: it sits in conversation history and is resent every turn, and it lands in request logs, traces and support tooling. Training use is a separate, narrower question.

solid answer

~50 s

"The provider does not train on it" answers only one of several exposure paths, and not the common one. Once text is in a prompt it is in the conversation history, so it is re-sent with every subsequent turn of that session; it is in whatever request/response logging, tracing and error capture the application does; it is visible to anyone who can open that thread — support staff, an admin console, a shared link — and it may be picked up later when someone samples production traffic. A whole discharge summary also carries far more identifiers than the task needs. The control is minimization at the input boundary: pull only the fields the task requires, redact before the call, and design the interface so the narrow path is the easy one rather than leaving a free-text box that invites a full paste.

go deeper

for a junior

Be ready to name the places a prompt gets copied: conversation history, logs and traces, support views, caches. Say plainly that not training on the data is a different question from not storing it.

for a middle

Explain minimization concretely — assemble prompts from the specific fields a task needs instead of accepting a pasted document, and keep full prompt bodies out of default logging.

for a senior

Show that you would go looking for the copies: which store holds threads, what the tracing layer captures, what support tooling can view, what sampling feeds evaluation. Then argue for the interface change that removes the paste path.

for a principal

Own the tradeoff between debuggability and exposure. Full prompt capture makes incidents solvable and makes every log store sensitive; argue for a default of metadata-only capture with deliberate, time-boxed exceptions.

## The claim being tested The reassurance people reach for is "the model does not learn from my input." Whether or not that is true for a given deployment, it addresses one exposure path — the model's weights — and leaves untouched every copy your own system makes of the same text. Most real LLM privacy incidents are of the second kind. Nothing exotic happens; the sensitive text simply ends up in more places than the person pasting it imagined. ## The copies your system makes **Conversation history.** A chat feature is stateless underneath: to answer turn five, the application re-sends turns one through four. A discharge summary pasted at the start is transmitted again on every later message and sits in whatever store holds the thread. It also keeps consuming context, which is why teams later add summarisation or truncation — and a summary of a record is still a record. **Logs and traces.** Application logs, an observability/tracing layer, an error report that captures the request body, a queue message, a retry buffer. These are usually built to be broadly readable by engineers, precisely because they exist for debugging. A prompt written into a log line inherits that audience. **Operational surfaces.** Admin consoles, support tooling that lets a human view a user's session to reproduce a complaint, exported conversation transcripts, screenshots pasted into a ticket, a shared thread link. Each is a legitimate feature and each widens the readership of the pasted text. **Downstream reuse.** Sampled production traffic is a normal source of evaluation examples and debugging fixtures. Text that entered as a prompt can end up in a dataset that outlives the session. **Caches.** Repeated prefixes are frequently cached to save cost and latency. A cached prefix is another stored copy with its own lifetime. None of these are misconfigurations; they are the default shape of a production application. Privacy work here is about noticing that a prompt is not ephemeral just because a conversation feels ephemeral. ## Why the whole record is the wrong unit A discharge summary bundles identity (name, address, record number, dates of birth and admission) with clinical content (diagnoses, medications, notes). Almost no task needs all of it. "Draft a follow-up appointment message" needs a diagnosis category, a discharge date and a clinician name; it does not need an address or a full medication history. Pasting the whole document imports every identifier into every copy listed above, and it also makes the text re-identifiable: identifiers such as record numbers are resolvable by anyone inside the organisation with access to the chart system. ## Minimization in practice *Fetch narrowly.* If the application can pull structured fields from the source system, pull the three fields the prompt template needs rather than the document. A template with named slots is both safer and more reliable than a free paste. *Redact before the call.* Run detection and replace identifiers with placeholders. This is lossy and imperfect, so treat it as reducing blast radius, not as a guarantee. *Shorten the retention of the copies you cannot avoid.* Do not log full prompts by default; log a request identifier and enough metadata to debug. If prompt capture is needed for a specific investigation, make it deliberate and time-boxed. *Design the easy path.* This is the part teams miss. If the only affordance is a large text box, users will paste whole documents, because that is the fastest way to get a good answer. If the interface offers "summarise this visit" with the record selected by identifier and assembled server-side, the sensitive text never passes through the user's clipboard or the chat log at all. Behaviour follows the interface; a policy asking clinicians to paste less will lose to a UI that rewards pasting more. *Give users a way to correct mistakes.* Someone will paste something they should not have. A visible delete on the conversation, that actually removes the stored copies, is worth more than a warning banner. ## How to talk about it The strong answer separates the two questions cleanly. Question one: what does the model provider do with the data — a contractual and configuration matter. Question two: what does *your* application do with it — a design matter entirely within your control, and the one that produces most incidents. A candidate who only answers question one has not thought about their own system.

  • The team says they will just scrub the logs. Is that enough?
    It helps, but it only covers one copy. Conversation history, caches, support views, exported transcripts and sampled traces are separate stores with separate lifetimes. Scrubbing logs while the thread store keeps the full text moves the problem rather than solving it. Minimization at the input boundary shrinks every copy at once, which is why it comes first.
  • If the assistant summarises the record before storing it, is the privacy problem solved?
    No. A summary of a patient record is still patient data, and summarisation happens after the full text has already been sent, logged and cached. It can reduce what persists in history going forward, which is worth something, but the exposure at the moment of the call is unchanged. Treat summarisation as a context-budget technique, not a privacy control.
  • Why does the interface design matter more than a written policy here?
    Because the fastest path wins. A free-text box makes pasting an entire document the lowest-effort way to get a useful answer, so people do it regardless of training. Assembling the prompt server-side from selected structured fields removes the temptation entirely and gives you one place to enforce minimization and redaction.

saying these in an interview costs you the question

  • Says training opt-out means the data is not stored anywhere
  • Thinks a prompt is discarded once the model replies
  • Forgets conversation history is resent on every turn
  • Treats logs and traces as private because only engineers read them
  • Assumes summarising the record removes the privacy concern

context

open as a page

Why do off-the-shelf PII detectors miss hospital MRNs, and what does over-redaction cost?

level: middleimportance: must knowfreq 62%

basics

~20 s

Shipped recognizers cover identifiers standardised nationally or industry-wide; a facility-assigned medical record number has no fixed format, so nothing matches it. Misses leak silently, while blanket masking strips the dosages, dates and lab values the answer depended on.

open as a page

In a multi-facility RAG corpus, where must the tenant filter be enforced, and why?

level: seniorimportance: must knowfreq 55%

basics

~20 s

At query time in the retrieval layer, as a pre-filter derived from the caller's verified identity — or as physically separate indexes. Anything above that, including instructions to the model, is advice rather than a boundary, because a retrieved chunk is already disclosed.

open as a page

Which internal content is unsafe to place in a system prompt, and what should hold it instead?

level: middleimportance: should knowfreq 48%

basics

~20 s

Treat anything in the context window as disclosable to that session's user: no credentials, no other users' data, no internal policy you would not publish. Secrets belong in the tool-execution layer, access decisions in code, sensitive policy behind an authorized retrieval call.

open as a page