skip to content

In OpenRouter, what does the provider.data_collection setting control?

level: seniorimportance: nice to knowfreq 30%

answer

  1. a compliance predicate on routing
  2. policy metadata, not a contract
  3. failover is what it defends against
  4. allow is the default posture
  5. pair it with a hard allowlist

basics

~20 s

It is a routing filter on provider data policy. Left at its default of allow, any upstream is eligible; set to deny, OpenRouter routes only to providers whose published policy says they do not collect or retain your prompts for their own use.

solid answer

~50 s

OpenRouter tags each upstream provider with the data policy it publishes — notably whether prompts and completions may be stored or used for the provider's own purposes. `provider.data_collection` turns that tag into a routing predicate: `"allow"` (the default) leaves every provider eligible, `"deny"` removes the ones whose policy permits collection. It matters precisely because of fallback: without it, a request that normally lands on an acceptable provider can be failed over to a different one under load, and the compliance property you assumed silently changes. Two caveats for a senior answer. First, it is a filter over declared policy, not a technical guarantee or a contract — your DPA and vendor review still do that work. Second, it interacts with `allow_fallbacks`: if a shrunken pool must never be escaped, pair `data_collection: "deny"` with `allow_fallbacks: false` and accept that exhaustion becomes a hard error rather than a quiet reroute.

code

json · 9 lines
json
{
  "model": "vendor-a/model-name",
  "provider": {
    "data_collection": "deny",
    "only": ["approved-provider-one", "approved-provider-two"],
    "allow_fallbacks": false
  },
  "messages": [{ "role": "user", "content": "Classify this customer record." }]
}

go deeper

for a junior

Know that different upstream providers have different data policies and that OpenRouter can be told to avoid providers that may retain your prompts.

for a middle

Explain the two values, that allow is the default, and that the filter is applied per request against provider policy metadata before selection.

for a senior

Make the failover argument — a per-request predicate is what survives an outage-driven reroute — and pair it with an allowlist, disabled fallbacks, and logging of the serving provider.

for a principal

Own the tradeoff between an auditable data-handling posture and the availability you give up, and make the residual single-provider risk an agreed, documented decision.

## The problem a gateway creates Routing through an aggregator means the counterparty that actually sees your prompt is chosen at request time, sometimes differently from one call to the next. That is the point — it is what buys you failover. It is also a compliance hazard: "which company processed this text?" stops being a static fact about your architecture and becomes a per-request outcome you may not be recording. `provider.data_collection` exists to make one axis of that outcome controllable. ## What the field does OpenRouter maintains, per upstream provider, the data-handling policy that provider publishes — in particular whether prompts and outputs may be retained or used by the provider for its own purposes such as training. The request field takes `"allow"` or `"deny"`: - **`"allow"`** (default) — no filtering on this axis; every provider hosting the model is eligible. - **`"deny"`** — providers whose policy permits collection are excluded from the candidate pool before selection. Like the other `provider` fields it is a **routing predicate**, applied per request, over the same candidate set that `order`, `only`, `ignore` and `require_parameters` also constrain. ## Why the default is the dangerous case The reason this shows up in senior interviews is the interaction with fallback. Suppose your normal traffic lands on an upstream your legal team reviewed. During an outage, OpenRouter fails over — that is its job — and the replacement upstream may have a different policy. Nothing errors, latency barely moves, and the property you told your auditors about was true yesterday and false for an hour on Tuesday. A per-request filter is the only mechanism that survives failover, because it is re-evaluated on every attempt rather than being an assumption about steady state. ## The limits of the guarantee Be precise here, because overclaiming is itself a red flag: - It filters on **declared policy metadata**, not on verified behaviour. It is as good as the provider's own statement. - It is not a data-processing agreement, a residency control, or an encryption guarantee. Jurisdiction and contractual terms are separate questions the field does not answer. - It constrains the upstream, not the gateway. Your own OpenRouter account also has privacy and logging settings, and those are configured on the account rather than per request. A complete answer mentions both layers. - It shrinks the pool, so it carries the same availability cost as every other filter. ## Composing it into a real policy For regulated or customer-confidential traffic, the usual shape is layered: 1. `data_collection: "deny"` to exclude collecting providers on every request. 2. `only` (or a curated `order`) to restrict to the specific upstreams your review actually approved — because "does not collect" is a weaker statement than "is on our approved list". 3. `allow_fallbacks: false` so exhausting that list produces an error instead of routing somewhere unapproved. 4. Logging of the serving provider for every request, so the compliance claim is evidenced rather than assumed. Step four is what turns configuration into an auditable control. If you cannot produce, for an arbitrary past request, the name of the company that processed it, you do not really have the control — you have an intention. ## The availability conversation Each layer costs headroom. A model with many upstreams may have only a couple that satisfy a strict policy, and if you also require capability parameters the intersection can be one provider. At that point your effective availability is that provider's availability, and you should say so explicitly when the control is agreed — ideally alongside a model fallback list of *equally approved* alternatives, so resilience is recovered without loosening the policy. ## Common mistakes Treating the flag as a technical guarantee rather than a policy filter; setting it while leaving `allow_fallbacks` at its default and believing the pool is fenced; assuming it covers OpenRouter's own logging; and never recording which provider served, which leaves the whole control unverifiable.

  • Why is setting data_collection to deny not enough on its own for a regulated workload?
    It filters on a single declared property and still leaves fallback free to pick any remaining provider. "Does not collect" is weaker than "is on our approved list", and it says nothing about jurisdiction or contractual terms. The workable control layers an explicit allowlist, allow_fallbacks set to false so exhaustion errors rather than reroutes, and per-request logging of the serving provider so the claim can be evidenced later.
  • How would you prove after the fact which provider handled a given request?
    By logging it at the time. OpenRouter reports the serving upstream on the response alongside the model that answered, so capture both with your request id and retain them as long as your compliance story requires. Reconstructing this later from configuration is not possible, because routing is a per-request outcome — configuration tells you what was permitted, not what happened.
  • What availability cost should you plan for when stacking these filters?
    Each predicate shrinks the candidate pool, and denying collection plus requiring specific parameters plus an explicit allowlist can intersect down to a single upstream. Your effective availability then equals that provider's. Recover resilience without loosening policy by curating a fallback list of equally approved models, and state the residual risk explicitly when the control is agreed rather than discovering it during an incident.

saying these in an interview costs you the question

  • Treats the setting as a technical guarantee rather than a policy filter
  • Sets deny but leaves allow_fallbacks at its default and calls the pool fenced
  • Assumes it also governs OpenRouter's own request logging
  • Never logs which provider served, leaving the control unverifiable
  • Ignores that stacking filters can reduce the pool to one upstream

context