skip to content

An auditor asks which licence list your Rego policy enforced six months ago. How do you answer?

level: seniorimportance: nice to knowfreq 32%

answer

  1. Rule is in git; facts often not
  2. Three homes for one fact, three answers
  3. A fetched value exists only during that evaluation
  4. The clock is an input in disguise
  5. Pure function of documents you can name

basics

~20 s

You can answer only if the fact was pinned. A licence list shipped in the policy's data document has a version you can name and replay; one fetched by http.send at decision time was never captured.

solid answer

~50 s

The rule text is only half the decision. A verdict is a function of `input`, the loaded `data`, and anything the rule fetched while evaluating. If the licence classification lived in `data` and that document was versioned alongside the policy, you name the version that was in force and re-run the same SBOM against it to reproduce the verdict. If the rule called `http.send` to a licence service, the fact existed only in that response, the service has changed since, and nothing recorded what it said, so you have the verdict but not the evidence behind it. `time.now_ns()` in a rule body creates the same gap. The fix is to make the decision a pure function of documents you can name: fetch outside evaluation and pass the result in as `input`, pass the evaluation time in rather than reading the clock, and capture the response if you must fetch inline.

go deeper

for a junior

Know that a policy's answer depends on more than the rule text, and that a value fetched while the rule runs is not saved anywhere by default.

for a middle

Explain the three places a fact can live and what each implies for repeating an old decision, and why reading the clock inside a rule breaks replay.

for a senior

Diagnose the disagreeing-runs case out loud: cache lifetime, a changed upstream list, or the clock, and say which change you would make first to make the gate reconstructable.

for a principal

Be ready to state the gap honestly and bound it. Decide what a decision record must carry, and who owns versioning the reference facts the gate depends on.

## The question behind the question An auditor asking "which list was in force" is not asking about Rego. They are asking whether your control is **reconstructable**: can you show, for a specific decision on a specific date, what the rule was and what facts it used. The rule is easy — it is in version control. The facts are where this falls apart. ## Three homes for a fact, three different answers Take a licence-class check over an SBOM. The component inventory is `input`. The classification of `AGPL-3.0` as restricted is a fact that has to come from somewhere, and there are three plausible homes: 1. **In `data`, loaded as a versioned artifact.** The list is content the engine holds; it was produced by a release with an identifier and a timestamp. Six months later you name that version, load it, replay the SBOM, and get the same verdict. This is the only one of the three where the auditor's question has a clean answer. 2. **As an extra field in `input`.** Whatever assembled the query looked up the classification and passed it in. Now the fact is captured wherever the query is captured, and the burden shifts to whatever built that query being able to say where it got the value. 3. **Fetched with `http.send` during evaluation.** The fact existed only inside that one evaluation. Nothing in the policy repo says what the service returned. Six months later the service has been updated, possibly several times, and there is no way back. Option 3 is the one that feels most correct while you are authoring — always fresh, never stale — and it is the one that leaves you unable to answer. ## The symptom that precedes the audit This usually surfaces before an auditor asks, as a support ticket: the same SBOM produced a violation on one run and none on another. Same artifact, same policy commit, different answer. There are two mechanisms and both live in this leaf. **The fetched fact moved.** Between the two runs the licence service reclassified something, or a `force_cache` entry expired and the second run refetched while the first was served from a cache populated before the change. `force_cache_duration_seconds` is a knob that makes the answer depend on wall-clock timing of cache fills — invisible in the policy, invisible in the diff. **The clock moved.** `time.now_ns()` returns the current time in nanoseconds. It is consistent within a single evaluation, so all calls in one query agree with each other, but across evaluations it is by definition different. A rule that denies when an approval or an exception has expired reads the clock, and its verdict changes at a moment nobody committed. That is *correct behaviour* for an expiry rule and simultaneously a reproducibility problem, which is why the fix is not to remove the clock but to move it. ## Making the decision replayable The target is a decision that is a **pure function of named documents**. Concretely: - **Move the fetch out of the evaluation.** Have the caller resolve the fact and pass it in as part of `input`. The gate becomes a pure evaluation, the network failure moves to a place with normal retry and error handling, and the fact is captured with the query. - **Ship reference facts as versioned `data`.** A list that changes weekly is still a release. Give it a version, record which version the engine had loaded, and "which list was in force" becomes a lookup rather than an investigation. - **Pass the time in.** `input.now` instead of `time.now_ns()`. The expiry rule behaves identically in production, and in a test or a replay you set the value and get a deterministic answer. - **If you must fetch inline, capture the response.** Whatever record the decision produces should carry the fetched value, not only the verdict. A verdict alone is an assertion; a verdict plus the fact it rested on is evidence. ## What to say when the answer is "we cannot" Sometimes the honest answer to the auditor is that the decision is not reconstructable, and the useful move is to say so precisely: the rule is in version control at this commit, the artifact is retained, but the classification was resolved live and was not captured, so we can tell you what the policy would say today and not what it said then. Follow it with the specific change that closes the gap. An auditor is far better served by a bounded, named gap with a fix attached than by an inference dressed up as a record — and an interviewer asking this question is usually listening for exactly that distinction between what you know and what you are guessing.

  • Why can two runs of the same policy over a byte-identical SBOM disagree?
    Because the rule reads something outside the artifact. A licence classification fetched with `http.send` may have changed between runs, or a `force_cache` entry may have expired so one run was served stale and the other fresh. A rule calling `time.now_ns()` disagrees with itself the moment an expiry passes. None of it shows up as a diff in the policy repo.
  • How do you write an expiry check without making the rule unreplayable?
    Pass the evaluation time in as part of `input` rather than calling `time.now_ns()` in the rule body. Production behaviour is unchanged because the caller supplies the real clock, but a test or a replay can supply a fixed value and get a deterministic verdict. The rule becomes a pure function of its documents again.
  • The classification is passed in as an extra field on input instead. Has the audit problem gone away?
    It has moved, not vanished. The fact is now captured with the query, so whatever retains queries retains the fact. But the burden shifts to whatever assembled that query being able to say where the value came from. It is a better position than an inline fetch, because at least the value that was used is recoverable.

Keeping the rule but not the fact is like keeping the recipe and throwing away the label off the ingredient you actually used.

saying these in an interview costs you the question

  • Says the policy commit alone proves what was enforced
  • Treats a live lookup as inherently more trustworthy than versioned data
  • Claims time.now_ns differs between calls in one evaluation
  • Assumes a decision record's verdict implies its inputs
  • Proposes deleting the expiry check to regain determinism

context