skip to content

In LLM APIs, what does a training opt-out still leave retained on the provider side?

level: middleimportance: must knowfreq 60%

answer

  1. two different promises, often conflated
  2. training use is not retention
  3. abuse-monitoring window still exists
  4. caching and batch keep their own copies
  5. your own logs usually outlive the provider's

basics

~20 s

A training opt-out only stops your text improving future models. Providers typically still hold prompts and responses for a bounded abuse-monitoring window, and features such as prompt caching, batch jobs and fine-tuning keep their own copies with their own lifetimes.

solid answer

~50 s

"We do not train on your data" and "we do not keep your data" are two different commitments, and teams routinely buy the first while believing they bought the second. As of mid-2026 the common enterprise shape is: **training use** is off by default on paid API tiers; **retention** is a separate, bounded window — providers typically buffer inputs and outputs for a period so they can investigate abuse and policy violations; and **zero-data-retention** is a further arrangement, usually contractual and account-scoped, that suppresses that buffer. Separately, product features create their own copies with their own clocks: a prompt cache stores the cached prefix for a TTL, asynchronous batch jobs hold inputs and results until collected, and any fine-tuning or evaluation you run leaves the training file on the provider side. The engineering answer is to enumerate every retained copy, get the window for each in writing, verify it in the account console, and treat the residual window as the exposure you must redact against.

go deeper

for a junior

Be able to state that a hosted model call transmits your text to a third party, and that whether it is used for training is a separate question from how long it is stored.

for a middle

Distinguish training use, bounded retention and zero-retention arrangements, and name the features that keep their own copies: prompt caching, batch jobs, fine-tuning and eval uploads.

for a senior

Show how you verify the posture — terms, console, subprocessor list — and argue that the application still redacts before the call, because a provider setting is a control you cannot enforce at runtime.

for a principal

Own the vendor decision: which data classes may reach which providers at all, what you require contractually before that is allowed, and what the fallback is when a provider changes terms or a required feature is excluded from the arrangement.

## Three different promises A provider's data-handling posture is not one switch. In practice there are at least three distinct commitments, and confusing them is the most common failure in this area. 1. **Training use** — whether your inputs and outputs may be used to improve the vendor's models. On paid API tiers this is typically off by default as of mid-2026, and often on by default for free consumer surfaces. This is the promise most easily obtained and the one most often mistaken for the others. 2. **Retention** — how long the provider stores the text at all. Even with training use disabled, providers commonly keep inputs and outputs for a bounded period to investigate abuse, safety violations and billing disputes. A limited retention window is not a secret; it is written into the terms, and it is where your data actually sits. 3. **Zero-data-retention** — an additional arrangement, usually enterprise-tier and contractual, that suppresses that buffer so requests are processed and discarded. It is not universally available, it often excludes specific features, and it is the thing you must ask for by name rather than assume. An interviewer wants to hear that you know which of the three you have, per provider, per feature. ## Feature-level copies with their own clocks Even under a strong account-level posture, individual features create storage: - **Prompt / prefix caching.** Caching a long shared prefix is the single highest-leverage cost lever available, and it works by the provider holding that prefix for a time-to-live. If the cached prefix contains client data — retrieved documents, a case file, a portfolio snapshot — that data is resident for the cache lifetime by design. - **Batch / asynchronous APIs.** Inputs and results are stored until the job completes and the output is fetched, and often for a grace period after. - **Fine-tuning and evaluation.** The uploaded training or eval file lives on the provider side as a first-class object with its own lifecycle, entirely separate from inference retention. And once a model has been trained on that file, the data is in the weights: deleting the file does not remove it. - **Gateways and intermediaries.** If requests traverse an AI gateway, an observability vendor or a framework's hosted tracing, each of those is a further processor with its own retention default — frequently longer than the model provider's. ## Subprocessors The provider you contracted with is often not the only party touching the request. Model providers run on cloud infrastructure and may use subprocessors for hosting, safety review or support tooling. The subprocessor list is a published artefact you are expected to read, because it determines who else holds a copy and where. In a regulated setting this list, not the marketing page, is the thing your reviewers will ask for. ## Human review Abuse-monitoring buffers exist so that flagged traffic can be examined — sometimes by an automated classifier, sometimes by a person. "Retained for 30 days" therefore implies "potentially readable by an employee of the provider during those 30 days." For a private bank summarising client meeting notes, that possibility is the whole reason the redaction layer exists: you assume the residual window is real and make sure what sits in it is already tokenized. ## Verifying rather than assuming The defensible version of this control has three parts: - **Contract.** The data-processing terms name the retention window, the training-use position and any zero-retention arrangement, with the feature exclusions written down. - **Configuration.** The account console shows the setting actually applied — organisation-wide, not per-project, and not silently different for a second account someone created for a prototype. - **Enforcement in code.** The application does not depend on the setting being right. It redacts before the call anyway, so that a misconfigured account is a policy problem rather than a disclosure. That last point is what separates a senior answer from a procurement answer. Provider settings are a control you do not operate; the redaction boundary is a control you do. ## The shadow-copy failure The recurring incident in this space is not the provider. It is a debug flag. Somebody turns on full prompt-and-response capture to chase a bug, the flag survives the fix, and for two weeks every client note is written a second time into a store that was never reviewed for this data class, with a retention far longer than the provider's. Your own logging is usually the longest-lived copy of the prompt in the system. ## What to say in an interview Name the three commitments and keep them distinct; list the features that retain independently; mention subprocessors and human review as the reason the window matters; and finish on the point that you redact regardless, because a setting you cannot audit at runtime is not a control you can rely on.

  • Your account has zero-data-retention. Why might a security reviewer still object to sending client notes?
    Because retention is only one of the questions. Where the inference physically happens, which subprocessors are involved, whether the feature set you use is covered by the arrangement, and what your own gateway and trace stores keep are all untouched by it. Zero-retention also does not change the fact that the disclosure occurred — the text was transmitted to and processed by a third party, which is the transfer your reviewer is assessing.
  • A team fine-tuned a model on client transcripts, then deleted the uploaded training file. Is the data gone?
    No. Deleting the file removes one copy; the model trained on it retains what it learned, and there is no reliable way to excise specific records from weights. The remedies are coarse: retire the model, or retrain from a corrected dataset. This is why fine-tuning inputs get the strictest pre-send redaction of anything in the pipeline — a mistake there is not correctable later.
  • How would you actually verify the posture rather than trust a documentation page?
    Three artefacts: the signed data-processing terms naming retention, training use and feature exclusions; a screenshot or API read of the setting as applied at the organisation level, checked for every account and project that can reach production keys; and an inventory of intermediaries — gateways, tracing vendors, frameworks — each with its own retention answer. Re-check on renewal and after any provider terms change.
  • Which single copy of a prompt usually has the longest retention in a typical deployment?
    Your own. Provider abuse buffers are measured in days to weeks, while application logs, trace stores and analytics warehouses are commonly retained for months or years and are readable by far more people. Any conversation about provider retention that does not also set retention on the trace store has optimised the smaller exposure.

saying these in an interview costs you the question

  • Treats a training opt-out as meaning nothing is stored
  • Assumes zero-data-retention is on by default for everyone
  • Forgets that prompt caching and batch jobs retain independently
  • Believes deleting a fine-tuning file removes the data from the model
  • Checks the marketing page instead of the terms and the console

context