Why can a host that passed a security baseline check last quarter fail the same check today?
answer
- a pass describes one moment
- the host moved, not the rule
- hotfix, package update, agent write-back
- on-call widened it at 03:00
- disk state versus state in force
basics
~20 sThe host moved, not the rule. Three ordinary causes: an operator hand-edited configuration during an incident, a package upgrade shipped new vendor defaults over the compliant settings, or an agent converged the host back to its own state.
solid answer
~50 sA pass is a measurement of one moment, not a property the machine keeps. Between the two runs the host's state drifted away from the recorded baseline, usually for a legitimate reason. The three classic sources are a human hotfix (on-call widens an egress firewall rule at 03:00 to end an incident and never narrows it back), a package or vendor update (upgrading the audit daemon replaces or supersedes the rule files, so the machine now runs the vendor's defaults instead of yours), and write-back by a converging agent or controller that re-asserts its own version of a file and drops the setting nobody told it about. None of those require negligence or an attacker. The point an interviewer wants is that decay is the expected consequence of a machine being operated, so a control has to be re-measured rather than assumed to hold.
go deeper
Be ready to name the three decay sources without hesitating: human hotfix, package or vendor update, and an agent writing state back. Say clearly that a pass describes the moment it was taken.
An interviewer expects you to explain the mechanics: how a package upgrade can either overwrite your file or strand it, and why the state on disk and the state loaded in the kernel can disagree.
Show that you reason about decay as an operational fact of the fleet, not a per-host incident, and that you ask which state was measured before you argue with a host owner about a finding.
Own the framing that decay rate is a property of how the platform is built. If a control decays within days on every host, the control is being fought by normal work and the design, not the operators, is what needs changing.
## What baseline decay names A host baseline is a recorded statement of what a compliant machine looks like: the egress firewall ruleset it carries, the proxy settings it uses, the audit-daemon rules that must be loaded, the services that must not listen. A check measures the live host and compares it against that statement. *Passed in March, fails in June, and nobody edited the policy* is baseline decay: the **measured state** moved away from the **recorded baseline** while the baseline stood still. The framing that matters in an interview is that decay is the normal result of a machine being used. Treating every re-failure as carelessness or as an intrusion is the classic weak answer, and it leads to the wrong conversation with the host's owner. ## The three decay sources **1. The human hotfix.** At 03:00 an outage is traced to blocked outbound traffic. The on-call engineer widens the host's egress rule, traffic flows, the incident closes. The change was correct — the alternative was a longer outage. What is missing is any path by which the recorded baseline learns that the machine now differs from it. The engineer had one job that night and it was not compliance bookkeeping. **2. The package or vendor update.** Upgrading a package that owns a configuration file changes the file's provenance. Package tooling generally does one of two things when the local copy has been edited: replace it with the vendor's new version, or keep the local copy and drop the vendor's version alongside for a human to merge. Both directions can break a control. In the first, your hardened settings are gone and the vendor defaults are in force. In the second, you keep running an old file while the vendor has renamed options, changed a default, or started reading a drop-in directory that now takes precedence. Nobody logged in; the state still moved. **3. Write-back by an agent or controller.** Anything that converges state will re-assert its own idea of a file: a configuration-management agent on a schedule, a bootstrap script that reruns at boot, a platform-managed agent that owns part of the host. If the compliant setting is not part of what that agent manages, the agent quietly removes it on its next pass. The same mechanism that repairs drift is also a source of it. ## Where you measure changes the answer For host controls, the state on disk and the state in force are two different things, and this is the detail that separates a shallow answer from a real one. Audit rules live in files that a loader compiles and installs into the kernel's audit subsystem; a rule added at runtime is in force but absent from disk, and a rule written to disk is not in force until it is loaded. The firewall has the same split: a ruleset live in the kernel and a saved ruleset restored at boot. So a check that reads only files can report a pass while the machine behaves differently, and a check that reads only running state can start failing after a reboot for a reason no file explains. When someone says the host decayed, the honest follow-up is *which state did you measure*. ## Decay is about a period, not a host A single failing host is a small problem: its state is knowable and fixable. The larger consequence is what the failure does to the claim. A control is useful because it held over a period, and a re-failure means the period now contains a gap whose start you usually cannot pin down — you know the control held at the last passing measurement and does not hold now, and everything between is unknown unless something independently timestamps the change (an incident timeline, an agent's convergence log, the modification time on the ruleset). That is why decay is discussed alongside how the machine is measured, rather than filed as a one-off ticket. ## What a strong junior answer sounds like Name the three sources, say plainly that all three are ordinary operational activity, and note that a pass describes the moment it was taken. If you can add that the recorded baseline can also move — a control tightened, a new rule added to the profile — you have covered both directions the gap can open from, which is what the next question up the ladder is about.
- Which of the three decay sources is hardest to notice, and why?The package or vendor update. Nobody logged in, no incident ticket exists, and the change arrives with routine patching that the host owner considers hygiene rather than a configuration change. The only local trace is a file timestamp or a leftover vendor copy dropped next to yours, and both are easy to miss until the next measurement fails.
- Can a host decay if no person has ever logged into it?Yes. Package upgrades and converging agents both change state without a human touching the machine. A host that is fully automated has fewer decay sources than a hand-operated one, but it is not free of them — it has traded human hotfixes for whatever its agents and its vendor decide.
- If the compliant file is on disk but the check still fails, what would you look at?The state actually in force. Audit rules must be loaded into the kernel before they take effect, and firewall rules can differ between the running ruleset and the saved one. Compare what the daemon or kernel reports as active with what the file says, and check whether anything has been loaded or flushed since boot.
A passing check is like a clean bill of health from a physical, not a vaccination. It describes the patient on the day, and ordinary living moves them off it.
saying these in an interview costs you the question
- Assumes a re-failure means someone was careless
- Treats every re-failure as an intrusion
- Says a passing host stays compliant until a human edits it
- Believes the check itself must have broken
- Ignores that on-disk config and running config can differ