skip to content

Assessing Live Systems

Once a control has a check, the open questions are what runs it, how often, and what to do when a host that passed in March fails in June. Interviewers probe cadence and decay together.

on this pageshow

explore

questions

12

A compliance check passed on March 3rd - what does that result claim about the system on March 20th?

level: juniorimportance: must knowfreq 62%

answer

  1. a verdict is bound to an instant
  2. state can move after the check returns
  3. some controls flip on the clock alone
  4. publish freshness next to the verdict
  5. stale means unknown, not pass

basics

~20 s

Only that the resource satisfied the rule at the moment it was evaluated. A pass is a timestamped observation, not a standing property, so every verdict has to be published together with the time it was taken.

solid answer

~40 s

A compliance evaluation reads live state at one instant and applies a rule to it, so the result means "this resource satisfied this rule at 03:14 on March 3rd" and nothing more. It says nothing about March 20th. Two things can invalidate it: the state moves, or the clock moves. The second one surprises people - a control like "access keys rotated within 90 days" flips from pass to fail with nobody touching anything, because the key simply ages. So a verdict is only usable next to its `evaluatedAt` timestamp, and the evaluation interval is the maximum age any claim on your dashboard can have. I define a staleness limit: a result older than that limit is reported as unknown, never carried forward as a pass.

go deeper

for a junior

Be ready to say plainly that a check result belongs to the moment it ran, and that you would always show the evaluation time beside it. Naming one control that decays on the calendar alone will get you full marks here.

for a middle

Explain the mechanics: the evaluation interval is the upper bound on the age of every claim, results should expire into an unknown state, and a run of passes says nothing about the gaps between them.

for a senior

Show the operating consequence - you set a staleness limit, you surface the oldest observation behind any aggregate figure, and you know which of your controls rot by the clock so you can schedule around them.

for a principal

Own the reporting contract: what the organisation is allowed to assert to a customer or a regulator from a set of observations, and why a percentage without a freshness distribution behind it is a claim you cannot defend.

## A verdict is an observation, not a property When a compliance check runs against a live resource, it reads that resource's state at one instant and applies a rule to it. The honest reading of the output is a sentence with a timestamp in it: *at 03:14 on March 3rd, this resource satisfied this rule*. It is not a statement about the resource, and it is not a statement about any other instant. Everything built on top - a dashboard tile, a quarterly report, an evidence pack - is an aggregation of such instants, and aggregation is exactly where the timestamps get lost unless you design against it. This matters because the most common way a compliance programme misleads its own organisation is not a wrong rule. It is a correct rule whose result is read as current when it is weeks old. ## Two different ways a pass stops being true **The state moves.** Someone or something changes the resource after the check ran, and the next evaluation would now fail. Why live state drifts away from a baseline is its own subject; for this question the point is only that the check has no visibility into anything that happens after it returns. **The clock moves and nothing else does.** This is the case people miss. Consider the control "IAM access keys and service-account credentials are rotated within 90 days". A key created on 15 December is 78 days old on March 3rd - a clean pass - and 95 days old on March 20th - a fail. Same key, same rule, same untouched account, opposite verdict. For controls of this shape, the age of the *evaluation* is not a minor caveat; the evaluation's freshness is nearly the whole of its value, because the truth is a function of the current date. A useful habit is to sort your controls into those whose truth changes only when someone acts (a bucket is public, encryption is off, a tag is missing) and those whose truth changes with the calendar alone (rotation windows, certificate expiry, recency-of-activity requirements). The second group decays on a schedule you can predict, which means you can also predict when your reporting will be wrong if you do not re-run. ## Freshness belongs next to the verdict Practically, this means three things: 1. **Every result carries the time it was taken.** A verdict without a timestamp is not evidence; it is a rumour. Store it, render it and export it. 2. **Results expire.** Define a maximum acceptable age - usually a small multiple of the evaluation interval - and render anything older as *unknown* or *stale*. Not as a pass. An expired pass rendered green is a fail-open reporting design: the display looks best exactly when you know least. 3. **The claim is phrased with its interval.** If you evaluate nightly, the strongest true sentence is "compliant as observed within the last 24 hours". If you sweep once per audit period, the same green tick can be eighty-nine days old and the sentence has to say so. ## What consecutive passes do not cover There is a second, quieter limit. Two passes on either side of a gap do not prove the gap was clean. If a resource violated the rule for six hours between two nightly evaluations, both runs return pass and no finding is ever raised. Sampling state at intervals reliably detects conditions that *persist* longer than the interval; it does not detect conditions shorter than it. If a control genuinely has to catch short-lived violations, the answer is not a tighter schedule but tying evaluation to change events or to an authoritative record of what happened, because you cannot shorten an interval to zero. ## How to talk about it An interviewer asking this is checking whether you understand that continuous compliance is a sampling problem. The answer they want has three moves: the result is bound to an instant; the interval bounds the age of every claim you make; and stale results are reported as unknown rather than inherited as green. If you add the observation that some controls decay by the clock alone, you have shown you know which of your checks rot fastest. The trap to name explicitly: a headline figure like "98% compliant" is a mix of observations of wildly different ages, so a resource that has not been looked at since two sweeps ago is silently counted as healthy. Aggregate the freshness alongside the verdict - the oldest observation contributing to a number is often more interesting than the number.

  • How should a dashboard render a resource whose last evaluation was six weeks ago?
    As unknown or stale, in its own bucket - never as a pass. Carrying an expired green forward is fail-open reporting: the display is most reassuring exactly where you have the least information. Show the age, and let the oldest contributing observation surface next to any aggregate percentage.
  • Which is riskier to trust two weeks later - a pass on "no publicly readable storage bucket", or a pass on "credentials rotated within 90 days"?
    The rotation one. A bucket only becomes public if somebody acts, so an old pass is wrong only if a change happened. A rotation window fails through the passage of time alone, so an old pass is guaranteed to go wrong on a date you could have calculated in advance.
  • What can you say about the period between two consecutive passing runs?
    Very little. Interval sampling catches conditions that persist longer than the interval; a violation that opened and closed between two runs leaves both passes intact and produces no finding. If catching short-lived violations matters, you need evaluation driven by change events or an authoritative activity record, not a slightly tighter schedule.

saying these in an interview costs you the question

  • Reads a passing check as "the system is compliant" with no time qualifier
  • Publishes verdicts without the timestamp of the evaluation
  • Treats a three-month-old pass and last night's pass as equal evidence
  • Assumes nothing changed because no one reported a change
  • Thinks a control can only flip after somebody edits something

context

open as a page

Why can a host that passed a security baseline check last quarter fail the same check today?

level: juniorimportance: must knowfreq 64%

basics

~20 s

The host moved, not the rule. Three ordinary causes: an operator hand-edited configuration during an incident, a package upgrade shipped new vendor defaults over the compliant settings, or an agent converged the host back to its own state.

open as a page

In an InSpec profile run, what is the difference between a failed control and a skipped one?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A failed control ran and the host did not satisfy it. A skipped control never ran at all - wrong platform, a guard, or a waiver - so it yields no evidence about that requirement, and it is not a pass.

open as a page

How do quarterly, nightly and change-triggered compliance evaluations differ in what they detect?

level: middleimportance: should knowfreq 57%

basics

~20 s

They differ in detection lag and in coverage. A per-audit-period sweep bounds lag at a quarter, a nightly run at a day, and change-triggered evaluation at seconds - but change-triggered only sees resources that emit an event, so it never fires for controls that expire on the calendar.

open as a page

A host that passed a baseline control now fails it: how do you tell whether the host moved or the baseline moved?

level: middleimportance: should knowfreq 48%

basics

~10 s

Compare both sides. Pin the revision of the check content that produced each result; if it is identical across both runs, the host moved. Then diff the recorded compliant state against the measured state.

open as a page

In InSpec, how do you override an upstream benchmark profile's sshd rule without forking it?

level: middleimportance: should knowfreq 56%

basics

~20 s

Write a wrapper profile: declare the upstream profile under depends in inspec.yml, pull its controls in with include_controls, skip_control the rule you are replacing, and add your own stricter control or feed the upstream one a different input value.

open as a page

How do you evaluate the control "a restore from backup was tested this quarter" on a continuous schedule?

level: seniorimportance: should knowfreq 38%

basics

~20 s

You cannot make the activity continuous, so you evaluate the record of it instead. Turn the control into a recency predicate over an authoritative test record - most recent successful restore test is under 90 days old - which any schedule can then check.

open as a page

On-call widened a host's egress firewall rule at 03:00 to end an incident and the baseline now fails: what did that cost the control?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The host is easy to fix; the period is not. The control can no longer be claimed continuously effective, and the gap is bounded only by the two measurements unless something independently timestamped the change.

open as a page

Why does kube-bench have to run on the node with host PID and host file access?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Most Kubernetes benchmark checks read the flags a component was actually started with and the ownership and mode of files on disk. Neither is exposed by the Kubernetes API, so the runner needs the node's process table and filesystem.

open as a page

Your first continuous compliance run returns 400 findings against one team - how do you land that with them?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Treat the 400 as the pre-existing state finally becoming visible, not as 400 new problems. Freeze them as a known baseline, hold the team only to what appears after today, group by root cause, and validate a sample before anyone sees the number.

open as a page

Host baseline controls across your fleet re-fail every quarter after legitimate work: do you engineer the decay out or accept it?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Decide per control, using its decay rate. A few decay fast enough to justify moving them off the host or rebuilding rather than repairing; for the rest, accept decay and claim periodic verification, not continuous enforcement.

open as a page

How do you run a host benchmark profile fleet-wide when it needs privileged access on every target?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Choose between a central runner holding credentials to every host and a local agent that ships only results, scope the privilege to reads, keep assessing separate from remediating, and measure coverage against an inventory so unassessed hosts never read as compliant.

open as a page