skip to content

An auditor wants last quarter's policy evaluation re-run to reproduce the same verdict — what makes that possible?

level: seniorimportance: should knowfreq 51%

answer

  1. evaluation is a function of its arguments
  2. archive all three: rule, input, engine
  3. external data is an argument too
  4. no I/O during evaluation
  5. replay proves computation, not the observation

basics

~20 s

Reproduction needs the archived input, the exact rule version, the engine version and any external data the rule used, all pinned by digest. It becomes impossible when the rule evaluated live state instead of a captured input.

solid answer

~50 s

Re-running an evaluation is a pure function call, so it reproduces only if every one of its arguments was archived: the input document the rule saw, the exact rule version by digest, the engine version, and any external data the rule consulted — a threshold list, an exemption set, an inventory — captured and hashed as part of the input rather than fetched at evaluation time. The three things that make reproduction impossible are all design choices made months earlier: the rule was only ever stored as a moving reference so the old text is gone; the input was never captured and only the verdict was kept; or the rule performed a live lookup during evaluation, so a replay queries today's world and answers a different question. Be precise about what a successful replay proves: that the verdict follows from the archived input, not that the archived input described reality on 14 November. That second claim rests on how the input was collected.

go deeper

for a junior

Understand the basic idea: the same rule plus the same input gives the same answer, so both have to be kept. Recognise that re-checking the resource today is a different thing from reproducing an old result.

for a middle

Be able to enumerate everything that must be archived — input, rule version, engine version, entry point, external data — and explain why a rule that reads live state cannot be replayed even though it works perfectly in production.

for a senior

Show the architectural commitment behind this: no I/O during evaluation, content-addressed inputs, immutable rule bundles retained as long as their evidence, and a replay tool you exercise before an auditor asks. State clearly what replay proves and what it does not.

for a principal

Own the cost of reproducibility as a design constraint across an estate: retaining every rule version, archiving inputs at volume, and refusing convenient rule designs that reach out to live data. Decide which controls are worth that and defend the line.

## Why this question is asked It is the sharpest test of whether automated compliance evidence is real. Anybody can produce a record saying a check passed. Reproducing the evaluation in front of an auditor — feeding the archived input back through the archived rule and getting the identical verdict — demonstrates that the record was computed rather than asserted. The question is asked in senior interviews because the answer reveals whether the candidate designed for it up front, which is the only time it can be designed. ## Evaluation as a pure function Model the evaluation as `verdict = evaluate(rule, input, engine)`. Reproduction is simply calling the same function again with the same arguments. That framing tells you exactly what must be archived, and each argument has its own failure mode. **The input.** The document the rule was given: the description of the volume and its backup configuration as it was at that moment. If it was never captured, there is nothing to replay against, and the evidence record degrades into an assertion. Store it content-addressed and record its digest in the evidence record. **The rule.** Not "the backup-retention rule" but the exact text that ran. Rules get edited — thresholds move, exemptions get added, a helper is refactored. If the only reference is a branch name or a package name that always resolves to the current version, last quarter's logic is gone. Pin by digest and keep every version retrievable for as long as the evidence it produced is retained. **The engine and its configuration.** Behaviour can change between engine versions, and evaluation is often parameterised — which parts of the ruleset were loaded, which entry point was queried, how a missing field is treated. Record the engine version and the entry point. **Everything the rule reached for.** This is the one people miss. If the rule consults external data — an approved-exemption list, an asset inventory that defines which volumes count as production, a per-environment threshold table — that data is an argument too. It must be resolved before evaluation, folded into the archived input, and hashed with it. Otherwise you have archived two of three arguments. ## The three ways reproduction dies 1. **A moving rule reference.** The record names the rule but not its version, or names a version that is no longer retrievable. The replay uses today's logic, which may pass or fail for reasons unrelated to last quarter. 2. **An input that was never captured.** The pipeline evaluated a live description and kept only the verdict. There is nothing to re-run against, only a re-observation of today's world — which answers a different question. 3. **A live lookup inside evaluation.** The rule itself called out during evaluation. This is the subtle one, because such a rule looks fine and works fine; it just cannot be replayed, and the failure only surfaces at audit time. The fix is architectural: evaluation takes a document and returns a verdict, with all collection done before it and all reporting done after it. A fourth, smaller cause is non-determinism — a rule that reads the wall clock, samples, or emits messages whose ordering varies. Verdict-affecting non-determinism must be eliminated; parameterise the clock by passing an evaluation time in the input rather than reading it. ## What a successful replay actually proves Here is the distinction that separates a good answer from a great one. A matching replay proves the verdict follows from the archived input under the archived rule: the computation was honest and the record was not fabricated. It does **not** prove the archived input truthfully described the resource on the day. That is a separate claim, and it rests on the collection path: which credential read the resource, what API it read, and whether that read is itself recorded. Present these as two links in a chain — collection fidelity, then computational reproducibility — and be explicit that only the second is what replay tests. ## Designing for it before you need it In practice this means a small number of commitments made early: evaluation performs no I/O; inputs are archived content-addressed alongside every verdict; rule bundles are immutable and addressable by digest, retained as long as their evidence; the evidence record names the engine version and entry point; and there is a replay tool that takes an evidence record and re-derives the verdict without the original pipeline. Test that tool on old records periodically — a replay path that has never been exercised tends to be broken by the time somebody needs it, and a walkthrough is the worst place to discover a rule version was garbage-collected.

  • Why does a rule that queries a cloud API during evaluation break reproducibility?
    Because the query is an unarchived argument. On replay it returns today's state, so you are re-observing rather than reproducing, and the verdict can differ for reasons unrelated to the original run. The fix is to collect state before evaluation, fold it into the input document, and hash it — evaluation should take a document and return a verdict.
  • The replay reproduces the verdict exactly. What has that not established?
    That the archived input was a truthful description of the resource at the time. Replay tests the computation, not the observation. Tying the input to reality is a separate claim resting on the collection path: which identity read the resource, through which API, and whether that read left its own record.
  • How do you keep an approved-exemption list from breaking reproduction?
    Resolve it before evaluation and fold it into the archived input rather than letting the rule fetch it at decision time. Then the exemption set in force that day is captured, hashed, and replayable. If the rule reads it live, a replay applies today's exemptions to last quarter's resource and quietly changes the answer.

saying these in an interview costs you the question

  • Assumes re-running the check today reproduces last quarter
  • Pins the rule by name or branch instead of by digest
  • Lets the rule fetch external data during evaluation
  • Keeps only the verdict and discards the input
  • Claims a matching replay proves the input was truthful
  • Never exercises the replay path until an audit

context