skip to content

Machine-Generated Evidence

Evidence is a record with fields, not a screenshot: which ruleset ran, over which input, when, and who approved it. Interviewers probe it because most teams cannot reproduce last quarter's result.

on this pageshow

questions

4

What must an automated backup-retention check record so its result works as audit evidence later?

level: juniorimportance: must knowfreq 63%

answer

  1. a verdict alone is not evidence
  2. one record per resource per period
  3. which control does this prove
  4. rule version and input digest
  5. keep the input, not just its hash

basics

~20 s

Each record needs the verdict plus everything that makes it re-checkable: which control it proves, which resource was evaluated, the exact rule version and input it saw, when it ran, and which identity ran it.

solid answer

~50 s

A stored `pass` on its own is worthless a year later. Write one record per evaluated resource per period, and put five things in it: the identifier of the control it proves, the identifier of the volume it evaluated, the verdict together with the observed value (`retentionDays: 35`), a digest of the exact rule version that produced the verdict, and a digest of the input the rule saw — plus the evaluation timestamp and the identity of the job that ran it. Archive the input itself alongside the record, or the digest attests to something you can no longer produce. Then write all of it to storage the producing job can append to but not overwrite. The test is not whether the record looks convincing; it is whether somebody who does not trust you can re-derive the verdict from what is stored.

code

json · 12 lines
json
{
  "controlId": "BKP-02",
  "statement": "backup retention >= 35 days",
  "resource": "vol-prod-ledger-01",
  "verdict": "pass",
  "observed": { "retentionDays": 35 },
  "ruleVersion": "sha256:9f2c...",
  "inputDigest": "sha256:41ab...",
  "inputRef": "evidence/2026-q1/vol-prod-ledger-01/input.json",
  "evaluatedAt": "2026-03-31T02:14:07Z",
  "runner": "compliance-scan@run-8842"
}

go deeper

for a junior

Be ready to list what goes in the record and say why each field is there: control, resource, verdict with the observed value, rule version, input digest, timestamp, runner. Knowing that a pipeline log is not evidence is most of the answer.

for a middle

Explain why the unit of evidence is one record per resource per period rather than one per run, and what the input digest is actually for. Expect to be pushed on what happens when the input was never archived.

for a senior

Show that you have designed the schema against how the record will be used: sampling by an auditor, drift questions months later, and a walkthrough where someone doubts the record's integrity. Talk about where it is stored and who can write there.

for a principal

Own the decision to treat evidence as a first-class artifact with a schema, a store and a retention period, rather than a by-product of CI. Be able to justify the cost of that against collecting evidence manually at quarter-end.

## What "evidence" means in this context An auditor testing a control is not asking whether you believe it held. They are asking to be shown an artifact, produced at the time, that a person who does not trust you can check for themselves. A green pipeline run is not that artifact. It is a UI state: the log rotates, the console link expires, the run identifier means nothing outside your tooling, and nothing in it states which control the run was supposed to prove. Machine-generated evidence is the deliberate act of turning a check's output into a durable, self-describing object. ## The control fixes the shape of the record Take a concrete control: every production volume must have at least 35 days of backup retention. That one sentence already determines the shape of the evidence. There is a **population** (production volumes), a **period** (the cadence at which the control is expected to operate), and an **assertion about each member of the population**. So the unit of evidence is one record per volume per period — not one record per pipeline run, and not one summary saying "all volumes compliant". ## The fields, and what each one buys you - **Control identifier.** Binds the record to the thing it proves. Without it, you have a technical scan result that somebody must later argue is relevant. With it, the record answers a specific control on its own. - **Resource identifier.** Names which member of the population this record is about. Use the stable machine identifier (a cloud resource ID or URI), not a display name someone can rename. - **Verdict plus the observed value.** `pass` alone hides the difference between a volume sitting exactly on the 35-day boundary and one retained for a year. The observed number lets you answer questions about margin and drift later without re-running anything. - **Rule version.** A digest of the exact policy bundle that produced the verdict. "Rule: backup-retention" is a name, and names get edited; a digest identifies one immutable text. - **Input digest, and the archived input.** The digest lets anyone confirm the stored input is the one that was evaluated. The archived input is what makes the digest usable — a hash of a document you threw away proves nothing. - **Evaluation timestamp.** When the observation was made, not when the file was written. Auditors test whether the control operated throughout the period, so the timing of each observation matters. - **Runner identity.** Which workload, job or service account produced the record. This is what lets the evidence store's own access records corroborate the claim. ## Why one record per resource Audit testing is sampling. The auditor picks items out of the population and asks you to show the evidence for those specific items. A single aggregate record forces you to re-derive per-item results at audit time from data that may no longer exist, and it hides the population: if three volumes were created and destroyed inside the period, an aggregate says nothing about them. Per-resource records also survive scope changes — when the population grows, you get more records rather than a redefined summary. ## The two questions every record has to survive 1. **Was it true when it was written?** That is a question about the input: where the observation came from, whether it was captured, and whether it faithfully described the resource at that moment. 2. **Has it been changed since?** That is a question about integrity: whether the stored object can be edited or deleted after the fact without anyone noticing. A record designed only to look good answers neither. The field list above exists because each field is what somebody reaches for when one of those two questions is asked in a walkthrough. ## Where the record goes Evidence belongs in a store the producing pipeline can append to and nothing routine can overwrite, with a retention period at least as long as the framework requires you to keep evidence for that control. If the same credential that writes the record can also delete it, the record's persistence is a matter of trust rather than a property of the system — and that is exactly the property an auditor is testing. ## The common failure The common failure is not a missing field; it is treating the CI system as the evidence store. Logs are retained for weeks, are mutable by anyone who can re-run a job, and are organised by pipeline rather than by control. The moment you decide the record is an artifact with its own schema, its own store and its own retention, most of the rest follows.

  • Why one record per volume rather than a single record saying all volumes passed?
    Audit testing is sampling: the auditor picks specific items and asks for the evidence for those items. A summary cannot be broken back down, hides which resources existed during the period, and says nothing about volumes created and destroyed mid-period. Per-resource records also keep working when the population grows, instead of needing a redefined aggregate.
  • The record stores a digest of the input but not the input itself. What breaks?
    Everything the digest was supposed to buy. A digest only proves that a document you can still produce is unchanged; with the document gone, there is nothing to hash and compare. You also lose any ability to recompute the verdict, so the record degrades into an unverifiable assertion that a check once passed.
  • The verdict already says pass — why also store the observed retention value?
    It separates a volume sitting exactly on the 35-day boundary from one retained for a year, which is the difference between a control with no margin and a comfortable one. It also lets you answer later questions about drift and near-misses from the archive rather than by re-running the check against today's world.

A green pipeline run is somebody telling you the shop was locked last night. An evidence record is the timestamped photo of that specific door, filed where nobody can swap it out later.

saying these in an interview costs you the question

  • Treats the CI job's green tick as the evidence
  • Records only pass or fail, with no resource identifier
  • Offers a dashboard screenshot as the audit artifact
  • Stores the rule name but not the rule version
  • Writes evidence where the producing job can overwrite it
  • Keeps an input digest after discarding the input

context

open as a page

How do you make a stored compliance evidence record tamper-evident to an auditor?

level: middleimportance: should knowfreq 42%

basics

~20 s

Hash each record, chain each hash to the previous record's, and sign the chain head with a key the producing job cannot reuse. Store the records in append-only storage locked for the retention period the framework requires.

open as a page

An auditor wants last quarter's policy evaluation re-run to reproduce the same verdict — what makes that possible?

level: seniorimportance: should knowfreq 51%

basics

~20 s

Reproduction needs the archived input, the exact rule version, the engine version and any external data the rule used, all pinned by digest. It becomes impossible when the rule evaluated live state instead of a captured input.

open as a page

Your team both operates the backup-retention control and produces its evidence — why should an auditor trust that?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

You cannot assert trustworthiness; you engineer it. Split the paths so the team that operates the control cannot silently alter, delete or forge its evidence, and let someone outside the team verify a sample independently.

open as a page