A host that passed a baseline control now fails it: how do you tell whether the host moved or the baseline moved?
answer
- two operands, one comparison
- pin the check revision per result
- diff observed values, not verdicts
- fleet-wide versus one host
- scope and inputs move too
basics
~10 sCompare both sides. Pin the revision of the check content that produced each result; if it is identical across both runs, the host moved. Then diff the recorded compliant state against the measured state.
solid answer
~50 sA re-failure has two possible origins and you cannot tell them apart from the finding alone. First establish whether the control content is the same: every result should record the revision of the check that produced it, so you can see whether a rule was tightened, a new rule added, or a check's measurement point changed between runs. If the content is identical across both runs, the machine moved, and the next step is a diff of the recorded compliant state against what was just measured, cross-referenced with anything that timestamps a change — the incident timeline, the converging agent's log, the modification time on the ruleset or audit rule files. Two traps: both sides can move at once, and the machine can look unchanged on disk while the state loaded in the kernel differs, so confirm which state the check reads.
go deeper
Know that a failing comparison has two sides, and that the first question is whether the check itself changed before anyone blames the machine.
Be ready to walk the elimination in order: pin the content revision, then diff observed values, then attribute the change using timestamps from the incident, the patch window or the agent's log.
Demonstrate that you look for the pattern first — one host versus the whole fleet — and that you know host state is layered, so a bit-identical file does not clear the machine.
Own the requirement that results carry provenance: content revision, observed values and scope at measurement time. Without it, every quarter's re-failures cost an investigation nobody budgeted for.
## Two sides, one finding A control failing today after passing last quarter tells you a comparison came out differently. It does not tell you which operand changed. There are exactly three possibilities and a good engineer eliminates them in order: 1. The **recorded baseline** moved — the control was tightened, a new rule was added to the profile, a threshold changed, or the control's applicability now covers this host when it did not before. 2. The **measured state** moved — the machine drifted, through a hotfix, a package upgrade or an agent writing state back. 3. The **measurement** moved — the check now reads a different place (the running kernel ruleset instead of the file on disk, the effective merged configuration instead of one file), or it reads something non-deterministic. Most teams debug only the second and lose an afternoon when the answer was the first or the third. ## Eliminating the baseline side The cheap discriminator is provenance on the result. Every measurement should carry the identity of the content that produced it: which revision of the control set ran, and ideally a hash of the specific control's logic. With that, *did the rule change* is a lookup rather than an argument. Without it, you are reduced to weak signals — for example, a whole fleet re-failing the same control in the same run is far more consistent with a content change than with every machine drifting simultaneously, whereas a single host failing a control the rest of the fleet still passes points at that host. Beware the softer forms of a baseline move. The rule text can be untouched while its *scope* changed: a host newly labelled as in-scope for a stricter role, an input or variable that supplies the allowed egress destinations updated to a shorter list, a control whose impact was raised so a result that used to be informational is now a failure. All of those move the recorded side without anyone editing the logic, and all of them produce a finding that looks exactly like decay. ## Eliminating the host side If the content is stable, diff the two states. What you want is not the pass/fail but the *observed values*: the egress ruleset the check saw last time against the one it sees now, the audit rule set as recorded against as measured. That diff names the field that moved, which is what makes the next conversation possible. Then attribute it. The three decay sources leave different traces: - A **human hotfix** usually correlates with an incident. The modification time on the ruleset, the shell history on the host, and the incident timeline will bracket the same window. This is where the on-call engineer who widened an egress rule at 03:00 shows up. - A **package upgrade** correlates with the patch window. The package database records when the owning package changed; a vendor copy of the file left alongside yours is a strong tell that the local file was preserved while the vendor's moved on. - An **agent write-back** correlates with a convergence run. Its log will say it enforced the resource, and the file's modification time will land on a convergence boundary rather than an incident. ## The measurement trap Host state is layered, and the two layers disagree more often than people expect. Audit rules are compiled from rule files and loaded into the kernel's audit subsystem; a rule loaded at runtime is in force but absent from disk, and a rule written to disk does nothing until it is loaded. Firewalls have the same split between the running ruleset and the saved one restored at boot. Proxy configuration can be set in a service's environment, in a system-wide file, and in the user's shell, with only one of them actually in effect for the daemon you care about. So a re-failure on a host whose files are bit-identical to last quarter is not a paradox. Either the check changed which layer it looks at, or a reboot flushed a runtime change into or out of existence, or the effective configuration is assembled from more inputs than the one file being diffed. When the file and the finding disagree, trust the layer that is actually in force and ask why the check and the machine disagree about where truth lives. ## What the answer should end with Both sides can move in the same interval, so *the rule changed* does not clear the host. Confirm each side independently rather than accepting the first explanation that fits, and record enough with each result — content revision, observed values, timestamp — that next quarter's version of this question takes minutes instead of a day.
- The whole fleet fails the same control in one run. What does that pattern suggest?Far more likely a change on the recorded side than simultaneous decay on every machine: a tightened control, a new rule in the profile, a changed input such as the list of allowed egress destinations, or a broadened scope that now includes these hosts. Confirm it by comparing the content revision that produced this run against the previous one before opening tickets against host owners.
- The control's logic is unchanged, but the result went from pass to fail. What non-logic changes could explain it?Scope and inputs. The host may have been newly labelled into a stricter role, a variable feeding the control may now list fewer allowed destinations, or the control's impact may have been raised so an informational result is now a failure. All move the recorded side without touching the rule text.
- What do you record with each result to make this diagnosis cheap next time?The revision of the control content that ran, the observed values and not just the verdict, the timestamp, and the host's identity and scope labels at the moment of measurement. With those four, distinguishing a baseline move from a host move is a comparison rather than an investigation.
saying these in an interview costs you the question
- Assumes a re-failure always means the host drifted
- Cannot say which version of the check produced a result
- Compares verdicts instead of observed values
- Forgets that scope and inputs move the baseline too
- Trusts the config file over the state actually in force