skip to content

In Rego, what is the difference between the input document and the data document?

level: juniorimportance: must knowfreq 80%

answer

  1. two document roots, not one
  2. one of them arrives per query
  3. loaded JSON and rule output share a tree
  4. package path becomes a data path

basics

~20 s

input is the document the caller sends with a single query, the thing being judged. data is the tree the engine already holds: JSON loaded alongside the policy, plus the virtual documents that rules themselves define.

solid answer

~40 s

Rego evaluates against two document roots. `input` is supplied by the caller with each query and lives only for that evaluation — in an infrastructure gate it is the artifact under judgement, such as an exported plan. `data` is everything the engine holds independently: base documents loaded from JSON or YAML files, and the virtual documents that rules produce. A rule named `violations` in `package estate` publishes its value at `data.estate.violations`, so rule output and loaded reference data share one addressable tree and a path alone cannot tell them apart. Rules read both by path — `input.resource_changes`, `data.catalog.families` — with no load or fetch call, and they can define documents only under `data`; `input` is read-only.

go deeper

for a junior

Be ready to say in one breath which root the caller supplies and which the engine holds, and to name a concrete path in each for a gate you have seen.

for a middle

Explain that rules define virtual documents under data at their package path, that a base document and a rule cannot share a path, and that references are plain path lookups with no fetch step.

for a senior

An interviewer expects you to reason about what belongs in each root when wiring a gate, and to spot that a policy needing a fact the caller never sends is a data-loading problem, not a rule problem.

for a principal

Own the contract: which facts the calling system is obliged to put in input, which the platform loads into data, and what the blast radius is when either drifts out of step with the rules.

### Rego has two document roots, and only one of them comes from the caller A Rego policy never "receives arguments" the way a function does. It evaluates against a document tree, and that tree has exactly two roots: `input` and `data`. **`input` is the document under evaluation.** The caller supplies it with the query, it exists only for the duration of that one evaluation, and nothing in the policy can write to it. In an infrastructure gate, `input` is typically the thing being judged — the JSON a plan export produced, a manifest, a request payload. Policies read it by path: `input.resource_changes`, `input.resource_changes[0].change.after.machine_type`. **`data` is everything the engine already holds.** Two different kinds of thing live in this single tree, and the query path cannot tell you which is which: - *Base documents* — plain JSON or YAML loaded alongside the policies. A file placing `{"catalog": {"families": ["n2", "c3"]}}` into the tree makes `data.catalog.families` readable by any rule. This is where slow-moving reference facts live: an approved machine-family catalogue, an owner registry, a list of sanctioned regions. - *Virtual documents* — the outputs of rules. A rule named `violations` declared in `package estate` publishes its value at `data.estate.violations`. It is computed on demand, when something asks for that path, not stored. That last point is the one candidates most often miss. Rules do not "return"; they **define documents inside `data`**, addressed by package path plus rule name. This is also why one rule can consume another simply by naming its path — there is no call, just a reference to a document that happens to be computed. ### Reading either root Both roots are read the same way, with dotted or bracketed paths: ```rego package estate approved_family if { family := input.resource_changes[0].change.after.machine_type family in data.catalog.families } ``` `import data.catalog` lets the body write `catalog.families` instead of the full path — pure shorthand, no loading semantics attached. There is no `read()`, `lookup()` or `fetch()`: if the path is present the reference evaluates to its value, and if it is absent the reference is undefined, which makes the expression — and therefore the rule body — produce nothing rather than raise an error. ### The boundary between the two The mechanical split is fixed: | | `input` | `data` | |---|---|---| | Who supplies it | the caller, per query | loaded with the policy, or computed by rules | | Lifetime | one evaluation | until the engine is reloaded | | Writable by a rule | never | yes — every rule defines a document under it | | Typical content in a gate | the artifact being judged | reference facts plus rule output | Two collisions are worth knowing. First, a base document and a rule cannot occupy the same path: if JSON is loaded at `data.catalog.families` *and* a rule named `families` exists in `package catalog`, the engine reports the conflict instead of choosing one. Keep loaded-data paths and rule packages disjoint. Second, `input` is a reserved root — you cannot declare a package or a rule under it, so any fact the policy needs but the caller did not send has to arrive through `data`. ### Why interviewers open here Every later question about rule shapes, query paths and gate wiring assumes this model. A candidate who thinks `data` is "the config file" and `input` is "the request body" is half right and will still be surprised the first time they query `data.estate.violations` and get a computed answer back — because the rule they wrote is, from the caller's point of view, indistinguishable from a JSON file someone loaded.

  • Can a rule write into input?
    No. `input` is read-only for the policy: the caller supplies it with the query and nothing in the policy can add to it or persist it. Rules define documents only under `data`, at their package path. So any fact a policy needs but the caller did not send must arrive through `data` — loaded alongside the policy — rather than being manufactured into the input.
  • What happens if a loaded JSON file and a rule both claim data.catalog.families?
    The engine refuses it. A base document and a virtual document cannot occupy the same path, so JSON loaded at `data.catalog.families` colliding with a rule named `families` in `package catalog` is reported as a conflict rather than silently resolved in favour of one. Keep loaded-data paths and rule package paths disjoint.
  • How does a rule reference a fact that was loaded into data?
    By its path — `data.catalog.families` — or via `import data.catalog`, after which the body can write `catalog.families`. The import is pure shorthand; there is no loading or lookup step. The whole tree is addressable as a document, so a path that does not exist is simply undefined rather than an error.

input is the envelope handed over for this one decision. data is the filing cabinet behind the desk: partly folders someone filed there in advance, partly notes the standing rules write themselves.

saying these in an interview costs you the question

  • Swaps them: calls data the request and input the config
  • Thinks a rule must call a function to load data
  • Believes input is stored in the engine between queries
  • Cannot say where a rule's own output becomes readable
  • Assumes a rule can add fields back into input

context