skip to content

What is OPA's data document, and what are the ways facts get loaded into it?

level: juniorimportance: must knowfreq 72%

answer

  1. two documents a rule can read
  2. startup files, polling, pushing
  3. bundle activation versus a Data API write
  4. pushed facts die with the process

basics

~20 s

OPA's data document is the in-memory JSON tree of standing facts policies read as data.something, separate from the input being decided. Facts arrive as startup files, inside a bundle OPA polls for, or as pushes to its Data API.

solid answer

~50 s

A Rego policy is evaluated against two documents: `input`, the thing being decided right now, and `data`, a JSON tree of standing facts OPA holds in memory. A rule that only lets a workload request an approved instance type reads the requested type from `input` and the approved list from `data.approved.instance_types` — the list is a fact about the organisation, so it cannot ride in on the request. Three loading paths: files and directories handed to `opa run` at startup; a **bundle**, a tarball of policy plus data that OPA periodically fetches from a bundle service and activates as a unit; and a **push** to the Data API, `PUT /v1/data/approved/instance_types`, from whatever job owns the source of truth. Bundles suit facts curated and reviewed alongside policy; pushes suit facts that change faster than you want to rebuild bundles. Pushed data lives only in that process's memory.

go deeper

for a junior

Be ready to say what input and data each hold, and to name the loading paths: files at startup, a polled bundle, a push to the Data API. Knowing that data lives in memory is enough at this level.

for a middle

An interviewer expects you to explain the mechanics: a bundle is fetched and activated as a unit under declared roots, while the Data API is written to per instance and holds nothing across a restart.

for a senior

Show you have operated this. Talk about which facts you put in the bundle versus push, how every replica converges, and what a new pod holds before its first fetch completes.

for a principal

Own the ownership question: which team owns each subtree of the data document, how fact changes get reviewed, and why letting two producers write one path eventually produces decisions nobody can explain.

### The two documents a rule can read Every OPA decision is a query evaluated against two things. `input` is the object under decision — an admission request, a Terraform plan JSON, a CI job payload. `data` is everything else OPA knows: a JSON tree, held in the process's memory, that any rule can reference by path as `data.<something>`. Their lifecycles differ completely. `input` arrives with the query and is gone when the answer returns; `data` is loaded once and then serves every decision until something changes it. A realistic guardrail needs both. "A workload may only request an instance type on the organisation's approved list" reads the requested type out of `input` and the list out of `data`: ```rego package platform.workloads deny contains msg if { not data.approved.instance_types[input.instance_type] msg := sprintf("instance type %v is not on the approved list", [input.instance_type]) } ``` The approved list is a fact about the organisation rather than about the change, so it has to already be inside the engine when the request arrives. ### Path one: files at startup `opa run` accepts files and directories. JSON and YAML files are loaded into `data` at the path implied by their location on disk, so `approved/instance_types.json` lands at `data.approved.instance_types`. This is the simplest thing that works and is fine for a fact that only changes when you redeploy — but it makes every fact update a restart, and it makes the facts invisible to anyone reading the deployment rather than the image. ### Path two: bundles OPA periodically fetches a bundle — a gzipped tarball carrying policy modules and data files — from a bundle service, and activates the whole thing at once. This is the dominant model, because the facts a policy reads are versioned, reviewed and shipped exactly like the policy that reads them: one artifact, one revision, and any new replica becomes correct as soon as it has completed its first fetch. A bundle declares `roots` in its manifest; everything under those paths is owned by that bundle and is erased and rewritten on each activation. ### Path three: the Data API A running OPA exposes its store over HTTP. `PUT /v1/data/approved/instance_types` writes the document at that path; `PATCH` applies JSON Patch operations to it; `GET` reads back exactly what the engine currently holds, which is the single most useful diagnostic on this whole subject. Here the direction of travel is inverted: instead of OPA pulling on a schedule, whatever owns the source of truth pushes when it changes. That suits a fact that moves many times an hour, or one produced by a system that has no place in a bundle build. ### The three things people get wrong **Pushed data is memory only.** OPA's default store is in-process and in-memory. Restart the pod and everything pushed over the Data API is gone; what comes back is only what loads at boot — startup files and the first bundle fetch. A push pipeline therefore has to be re-drivable on demand, not fire-and-forget. **Replicas do not share a store.** Each OPA holds its own copy. There is no replication and no shared cache between them, so a push must reach every instance. In a sidecar deployment that is one write per pod, which is exactly why bundles win at scale: every instance converges on the same artifact by polling, without anyone tracking who has been written to. **One owner per subtree.** A bundle's declared roots are owned by that bundle and are rewritten on every activation, so a path inside a bundle root is not a safe home for pushed data. Decide per subtree whether the bundle or the push job owns it, and keep pushed facts outside the bundle's roots. ### Choosing between them A workable rule of thumb: if a human reviews changes to the fact and it moves on the same cadence as policy, put it in the bundle — you get review, versioning and a revision you can name when explaining a decision. If a machine owns the fact and it changes far faster than a bundle build, push it, and accept that you now own the freshness and the re-push-after-restart problem yourself.

  • If OPA restarts, what happens to data you pushed over the Data API?
    It is gone. OPA's default store is in memory, so a restart leaves only what loads at boot: startup files and the first bundle fetch. Whatever owns the fact has to push it again, which is why push pipelines need to be re-drivable on demand rather than fired once when the source of truth changes. Facts that must survive a restart unattended belong in the bundle.
  • You already ship the approved instance-type list in the bundle. What breaks if a job also pushes to that same path?
    You have given one subtree two owners. A bundle's declared roots belong to that bundle and are erased and rewritten on every activation, so a path inside a bundle root is not a safe home for pushed values. Pick one owner per subtree: either the list is built into the bundle, or the job pushes it to a path outside the bundle's roots.
  • Why can the approved list not simply be sent in as part of input?
    Because the caller would then be supplying the standard it is judged against, and every caller would have to know and correctly assemble the whole list. Standing facts belong in data, where the platform controls them; input carries only the change under decision. The split also keeps requests small and makes the facts auditable in one place.

saying these in an interview costs you the question

  • Confusing input with data; putting standing facts in input
  • Assuming data pushed over the API survives a restart
  • Thinking OPA replicas share one data store
  • Believing a bundle can only carry policy, not data

context