skip to content

Centralizing OPA as one shared decision service — what leaves each caller that never left before?

level: seniorimportance: nice to knowfreq 33%

answer

  1. the rule can only judge what it is told
  2. identifiers are the useful fields
  3. transit, processing, and retention
  4. decision logs outlive the decision
  5. minimize before you mask

basics

~20 s

Every query's input document. A rule can only judge facts it is given, so the input carries account, tenant and requester identifiers, and centralizing sends all of it across a boundary on every call and usually into a decision-log store.

solid answer

~50 s

Every query's input document. The rule needs facts to judge — the target region, the cloud account, which tenant the request belongs to, who asked — so the input carries exactly the identifiers your data-classification people care about. Centralizing means that document crosses a process, a host and usually a network boundary on every call, and then typically lands in a decision-log pipeline that retains it. The decision service becomes a collection point for every team's request data and a correspondingly attractive target. Ask, before centralizing: do these inputs carry residency constraints, who can read the decision logs, and does the rule genuinely need the identifying fields. The mitigations are minimizing what the caller sends, masking fields out of decision logs, running one decision service per data boundary, or keeping evaluation in-process so the input never leaves at all.

code

json · 12 lines
json
{
  "input": {
    "action": "create_instance",
    "region": "eu-central-1",
    "account_id": "acct-4471",
    "tenant": "northwind-health",
    "requested_by": "svc-provisioning",
    "tags": {
      "contains_patient_data": "true"
    }
  }
}

go deeper

for a junior

Know that OPA decides using an input document the caller supplies, and that a shared decision service therefore receives whatever identifiers the caller put in it.

for a middle

Explain what a realistic input carries for an account or region rule, and why the same document usually ends up in a decision-log pipeline as well as in the evaluation.

for a senior

Show that you would inspect a real input document against data classification before choosing a topology, and that you can rank minimizing, masking, partitioning and relocating the engine.

for a principal

Own the organizational consequence: a central decision service concentrates every team's request data under one operator, which is a governance and residency commitment, not just a deployment.

## The input is the part nobody budgets for Every discussion of OPA topology starts with latency and blast radius. The axis that decides the question in a regulated environment is a different one: **a shared decision service receives, on every single query, the raw `input` document that the caller assembled.** Embedding never emits it. A sidecar emits it only onto the pod's loopback. Centralizing emits it across a boundary, thousands of times a second, to a component operated by someone else. ### Why the input is always sensitive-ish Rego evaluates against facts. A rule cannot judge what it is not told, so the input contains precisely the fields the decision turns on — and the useful decisions turn on identity. A placement rule for a provisioning API needs the region and the cloud account. To make an exception for one business unit it needs the tenant. To attribute the decision it needs the requesting principal. To reason about sensitivity it may be handed the resource's own classification tags. Each of those is a business identifier, and together they describe who is doing what, where, for whom — which is close to the definition of the data an organization has rules about. ### Three boundaries, not one When you centralize, the input crosses three things at once, and they have different owners: 1. **A trust boundary.** The decision service is now inside the data flow of every calling team, including teams whose data the operating team was never cleared for. In a multi-tenant estate that means one platform team holds a live feed of every tenant's activity. 2. **A network and possibly a geographic boundary.** If the decision service is one regional deployment and callers are global, inputs from a residency-constrained tenant are now transiting and being processed outside their region. That is a contract question, not an engineering preference, and it is usually discovered late. 3. **A retention boundary.** OPA can ship decision logs — the input, the result, timing and metadata — to a remote endpoint. That is a genuinely valuable capability: it is how you prove after the fact what was decided and on what basis. It also means the inputs are no longer transient. They are now a dataset with a lifetime, an access-control list and a backup. The last point is the one that surprises people. Centralization does not just move the input; it usually *creates a store of inputs*, because the same central position that makes the service convenient to call also makes it the obvious place to collect evidence. ### What to do about it - **Minimize the input.** Send the fields the rule reads and no more. If the rule only needs to know that the tenant is in a regulated class, the caller can send `tenant_class: "regulated"` rather than the tenant's name. This is the highest-leverage fix because it removes the data instead of protecting it. - **Mask what is logged.** OPA supports masking fields out of decision logs with a policy, so you can keep the decision record while dropping or redacting the identifying fields inside it. Note the limitation honestly: masking protects the *log*, not the transit and processing that already happened. - **Run one decision service per boundary.** Per region, per tenant class, per business unit. You give up some of the operational simplicity that made centralizing attractive, and you keep the freshness benefits. - **Move the evaluation to the data.** A sidecar keeps the input on the pod; embedding keeps it in the process. If the input is the problem, relocating the engine is the only fix that eliminates rather than mitigates it. ### The sidecar is not automatically clean A candidate who answers 'use a sidecar' has solved the transit problem and not the retention one. The sidecar's decision logs still ship somewhere central — that is the point of them — so the same masking and access questions apply to the log pipeline even when the query itself never left the pod. What the sidecar genuinely buys is that the *synchronous* path carries no cross-boundary transfer, so an outage or a compromise of the central component does not expose live traffic. ### How to raise it The practical move is to ask, before a topology is chosen, for one real example of an input document from each candidate caller, and to walk it past whoever owns data classification. It is a five-minute exercise that reliably changes the answer, and it is far cheaper than discovering after rollout that a tenant's contract forbids the flow you just built.

  • Does moving to a sidecar remove the exposure entirely?
    It removes the synchronous cross-boundary transfer — the query stays on the pod's loopback — but not the retention problem. The sidecar's decision logs still ship to a central endpoint, so the same masking and access-control questions apply to that pipeline. Only the live path is cleaned up; the evidence trail still needs handling.
  • The rule needs the tenant to grant one business unit an exception. How do you avoid sending the tenant name?
    Push the classification to the caller: send a derived attribute such as a tenant class or an exception flag that the rule actually branches on, rather than the identifier it would otherwise look up. You lose some ability to write new rules without changing callers, and you gain an input document that carries no customer identity.
  • Why is the decision-log store often the bigger risk than the query path?
    Because the query is transient and the log is not. A decision-log pipeline turns a stream of inputs into a retained dataset with backups, an access list and a lifetime, frequently in a different system with different controls. Anyone reviewing this should ask who can query that store before asking about TLS on the decision call.

saying these in an interview costs you the question

  • Assumes the input is just a policy detail, not data
  • Overlooks that decision logs retain the input
  • Says TLS on the call solves the exposure
  • Treats a sidecar as removing every confidentiality concern
  • Never asks whether the rule needs the identifying fields

context