skip to content

How do you decide where LLM inference physically runs for regulated client data?

level: principalimportance: should knowfreq 30%

answer

  1. classify the data before choosing a vendor
  2. location is a property of the data class
  3. every step down costs capability
  4. tokenize before anything crosses
  5. telemetry is a second egress

basics

~20 s

Treat processing location as an architecture constraint set per data class, not a vendor checkbox. Classify what may leave the jurisdiction, check which providers offer in-region inference and which subprocessors are involved, and accept the capability gap that regional or self-hosted options impose.

solid answer

~60 s

Start by classifying data rather than choosing a provider. Split the corpus into what may cross a border at all, what may cross only after tokenization, and what may never leave — for a Zurich private bank, raw client names and portfolio positions typically sit in the last bucket while the surrounding narrative does not. Then map the constraint onto the deployment options, which form a ladder of decreasing capability: a frontier provider's in-region endpoint, a hyperscaler-hosted version of the same model in your region, an open-weight model you run yourself, and no model at all. Each step down costs capability, and the frontier tier usually lands in one region first — so a strict residency rule means committing to a model generation behind. The architecture that survives this is a routing boundary: a classifier decides the data class, a router picks the endpoint, and the tokenization layer sits in front of anything crossing a border. Verify residency against the subprocessor list and the contract, not the console label, and remember that inference location says nothing about where your own traces are stored — that is a second, independently-configured egress.

go deeper

for a junior

Understand that a model call sends text to a specific place, and that some data is not allowed to leave a country or region — so where inference runs is a real design input, not a detail.

for a middle

Explain the options — default endpoint, in-region endpoint, hyperscaler-hosted, self-hosted open weights — and that tokenizing before the call shrinks what actually crosses a border.

for a senior

Design the routing boundary: classify by calling context, map class to endpoint, fail closed, and verify against contracts and subprocessor lists rather than console labels. Catch telemetry as a second egress.

for a principal

Own the capability-versus-locality tradeoff explicitly, quantified on your own evals, and keep the decision reversible — an internal model interface plus fast evals so a provider or region change is a week of work, not a re-platform.

## Reframing the question "Which provider is compliant?" is the wrong opening. The decision that actually holds up is: for each class of data this product handles, where may that class be processed, and what is the best capability available under that constraint? Residency is a property you assign to data and then satisfy with architecture — not a badge you shop for. This matters because the constraint is rarely uniform. A private bank summarising client meeting notes has, in the same document, a client's name and holdings (strictly local), the narrative of a market discussion (less sensitive), and generic product questions (unconstrained). Treating the whole document at the strictest level is the safe default and also the expensive one: it pins your entire product to whatever the most restricted path can do. ## The deployment ladder Options ordered by decreasing capability and increasing control: 1. **Frontier provider, default endpoint.** Best models, best latency, most features — and processing wherever the provider chooses, with a subprocessor chain you inherit. 2. **Frontier provider, in-region processing option.** Increasingly offered as of mid-2026 for enterprise agreements. Usually a subset of regions, sometimes a subset of models, and often lagging the default endpoint by a generation. 3. **The same or a comparable model via a hyperscaler in your region.** You gain a familiar cloud contract and regional footprint; you may lose the newest model versions and some provider-specific features. 4. **Open-weight model, self-hosted in your own environment.** Full control of location and retention, no third-party disclosure at all, and a substantial capability gap plus the operational burden of running inference — capacity, GPUs, upgrades, evaluation of every model swap. 5. **No model for that data class.** A legitimate outcome. Some content should not go through an LLM at all, and saying so is a stronger answer than engineering around it. A principal-level answer names this ladder and is explicit that steps down cost quality, because the interviewer is testing whether you will pretend the constraint is free. ## The architecture that makes it tractable Do not build one pipeline and hope it satisfies the strictest class. Build a routing boundary: - **Classification at the edge.** Each request is labelled with a data class, from the calling context (which tenant, which product surface) far more than from content inspection, because context-based labels are auditable and content classifiers are probabilistic. - **A router that maps class to endpoint.** Class determines the region, provider and model — not a per-call convenience flag that a future developer can override. - **Tokenization in front of every crossing.** Anything leaving the jurisdiction goes through the redaction layer first, so the residency question shrinks to "surrogates crossed a border" rather than "client names crossed a border." This is the single highest-leverage move: it converts a hard constraint on the whole product into a constraint on one small vault. - **Fail-closed defaults.** Unclassified traffic takes the most restrictive path. A misconfiguration should degrade quality, never expand egress. ## Verification The console label is a claim; the artefacts are the contract's data-processing terms, the published subprocessor list, and the retention terms attached to the in-region option. Ask specifically whether abuse-monitoring buffers and any human review also stay in region, because those can be centralised even when inference is not. Ask which features are excluded — batch endpoints, caching tiers and fine-tuning are common carve-outs. ## The second egress everyone forgets Inference location and telemetry location are configured independently. A team can route inference to an in-region endpoint and then ship every prompt and completion to an observability vendor hosted elsewhere, with longer retention and broader access than the model provider. The same applies to error trackers, analytics warehouses and any hosted tracing built into a framework. If the residency requirement is real, it applies to every store the text reaches, and your own telemetry is usually the copy that violates it. ## The tradeoffs to own - **Capability versus locality.** Committing to strict residency generally means a model generation behind. Quantify that with your own evals rather than arguing it abstractly — sometimes the regional option is entirely sufficient for the task, and knowing that ends the debate. - **Latency and availability.** In-region endpoints can have thinner capacity and fewer fallback options during an outage, and your cross-region failover plan may be exactly the thing residency forbids. - **Operational cost of self-hosting.** Running inference yourself trades a vendor risk for a staffing and capacity problem, plus re-evaluation on every model upgrade. - **Portability as insurance.** Providers change terms, regions and model availability. Keeping the model behind an internal interface, and keeping evals that can score a replacement quickly, is what makes a forced migration a week rather than a quarter. ## What interviewers listen for Data classes before vendors; an honest capability ladder; a routing boundary rather than a global setting; verification against contract and subprocessor list; and the observation that telemetry is a second egress path that must satisfy the same rule.

  • Your strictest data class can only run on a self-hosted open-weight model that scores clearly worse on your evals. How do you decide whether to ship?
    Score it against the task, not against the frontier. If the regional or self-hosted model clears the bar the feature actually needs, the gap is irrelevant. If it does not, the choice is narrowing the feature for that class, adding human review to compensate, or declining to build it for that class — and saying so is legitimate. What is not legitimate is shipping the frontier path anyway and treating the constraint as advisory.
  • A team routes inference to an in-region endpoint but sends traces to a hosted observability vendor abroad. What has residency bought them?
    Very little for the data itself. The trace store typically holds the full prompt and completion, retains it far longer than the provider's buffer, and grants broader read access. Residency requirements apply to every store the text reaches, so the telemetry pipeline needs the same regional decision — self-hosted collection, an in-region vendor, or tokenized payloads so that only surrogates leave.
  • How do you keep a residency decision from ossifying into a permanent capability ceiling?
    Put the model behind an internal interface and keep an eval suite that can score a candidate quickly, so re-testing a newly available in-region model is days of work, not a project. Re-check the provider's regional and feature coverage on a schedule, because it moves; and revisit the data classification itself, since better tokenization can move content from the strict bucket to the permissive one and reopen the capability gap.

saying these in an interview costs you the question

  • Treats residency as one vendor checkbox for the whole product
  • Assumes an in-region endpoint covers abuse buffers and every feature
  • Sends traces abroad while calling inference in-region compliant
  • Ignores that regional and self-hosted options usually lag on capability
  • Applies the strictest data class to all traffic without classifying

context