Told the estate is fully patched, how do you separate hosts that received a fix from hosts running it?
answer
- two populations, not one percentage
- received is easy, running is the point
- compare against when the fix landed
- boot time, process start time, running firmware
- the tail is where the always-on hosts sit
basics
~20 sAsk for two numbers, not one. Installed count is on-disk state. Running count needs a per-host comparison: did the unit that loads this code start after the fix landed - boot time, process start time, running firmware version?
solid answer
~50 sI ask which of two populations the number describes: hosts that received the fix, or hosts executing it. They are always different and the second is the only one that matters. The separating comparison is per host and always the same shape - did the thing that must load this code start after the fix landed? For a kernel fix, boot time against the fixed kernel's install time, or simply the running kernel's build identity. For a userspace library, the start time of every process that maps it against the install time. For firmware, the running version against the staged version. And management controllers are not in the host cadence at all, so they need their own version answered separately. What I accept is per-host running-code identity across the whole managed population, including the always-on hosts - not an aggregate percentage whose denominator I cannot see.
code
text · 9 lineshost: hv-prod-114 (virtualisation host, management VLAN)
installed kernel package : 5.x.y-fixed (installed 2026-06-02)
running kernel build : 2025-11-18
uptime : 284 days
library fix installed : 2026-06-02
processes mapping it, started before that date : 341
platform firmware : running 2.11 / staged 2.14 (activates at power cycle)
management controller fw : 3.02 (own lifecycle, last changed 2024-09)
...go deeper
Know that there are two facts about every host - the fix is present, and the fix is running - and that the second needs a restart of whatever loaded the code.
Be able to name the comparison per surface: running kernel build for a kernel fix, process start times for a library, running against staged for firmware.
Show the reviewer's instinct: ask what the number counts, insist on per-host running-code identity across the whole managed population, and demand the named tail rather than the average.
Own how the estate reports patching at all. If the metric the board sees is distribution, the programme is optimising for the easy half of the problem, and the hosts in the tail are the ones an unhurried adversary is counting on.
## The two numbers Every estate has two patch percentages for any given fix: - **Received**: hosts where the fixed package, image or firmware file is present. This is what package databases, deployment systems and most quoted numbers describe. - **Running**: hosts where the fixed code is what the processor is actually executing. These are never equal, and the gap is not noise. It is concentrated in exactly the hosts you least want it in: always-on virtualisation hosts, appliances in a change freeze, machines that are only ever suspended and resumed rather than restarted. Received will sit in the high nineties because distribution is easy and automated. Running trails it by whatever share of the estate has not restarted the relevant unit since the fix landed. ## The one comparison The separating question is the same on every platform: **did the unit that must load this code start after the fix landed on this host?** Instantiated per surface: | Fix surface | The comparison | The unit that must restart | | --- | --- | --- | | Kernel | Running kernel build identity, or boot time after the fixed kernel install | Boot | | Userspace shared library | Start time of every process mapping it, against library install time | Each process | | A single service's binary | Service start time against install time | That service | | Device or platform firmware | Running version against staged version | Power cycle | | Management controller | Its own running firmware version | Its own restart, off the host cadence | Notice that four of the five rows are answerable from host state you can read without touching the flaw itself, and none of them is the installed package version. ## What to accept as proof Be explicit, because this is the part interviewers are listening for: 1. **Per-host facts, not an aggregate.** A percentage hides its denominator and hides which hosts are in the tail. The tail is the answer. 2. **The denominator is the whole managed population**, including hosts that are never restarted and appliances excluded from the usual sweep. If the always-on hosts are outside the count, the number is measuring the hosts that were never the problem. 3. **Running-code identity, not installed-code identity.** Boot time, process start time, running firmware version. 4. **A named exception list.** Hosts that cannot be brought into the running population get named individually with an owner, not averaged away. And say what you will not accept: a green tick from a distribution job, a closed change record, or the sentence "the version is current", all three of which are true of a host that is still executing the vulnerable code. ## Why the un-restarted host is the interesting one The reason to insist on the second number is not tidiness. Consider one identity's worth of access already present on a management segment, used by an operator with time and no revenue clock. That operator does not need a new flaw; the estate is generating stale hosts on its own, and the fixed-but-never-restarted population is a standing supply of them that no one counts as unpatched. Meanwhile an ordinary administrator postponing a power cycle on a hypervisor for a perfectly sound business reason produces the identical host state. From the version number alone the two are indistinguishable - which is the argument for measuring the state rather than the intent. ## The reviewer's script When handed "we are fully patched, so we are fine", the useful reply is one question and one request. The question: "is that installed or running?" The request: "give me boot time and the running kernel build for every host, and the list of hosts whose uptime predates the fix." If the second list is empty on an estate with virtualisation hosts in it, the list is wrong, not the estate.
- The aggregate says 98%. Why is that number almost useless to you?Because it hides its denominator and its tail. It is very likely counting installs, and the 2% is not random - it concentrates in always-on hosts, frozen appliances and suspended machines, which are the highest-value targets in the estate. I want the named tail and the running-code identity per host, not the mean.
- How do you answer the same question for out-of-band management controllers?Separately, because they are not in the host cadence. A controller runs its own firmware on its own processor and remains powered when the host is off, so the host's kernel and package state say nothing about it. It needs its own running-version inventory and its own update cycle, and it usually has neither.
- What would you accept as proof that a userspace library fix is in force on a host?The install time of the library and the start time of every process that maps it, with none of those start times earlier than the install. A reboot after the install proves the same thing more bluntly. What I will not accept is the installed version string, which is true of both the fixed and the still-vulnerable host.
- Is a host that installed the fix and never restarted safer than one that never received it?On that flaw, no - it is executing the same vulnerable code, so exploitability is identical. It is marginally better positioned because the fix is already local and takes effect at the next restart. But treating it as patched is the error: the difference is one restart away, and until then the risk is unchanged.
saying these in an interview costs you the question
- Accepts a single patch percentage without asking what it counts
- Treats a closed change record as proof the fix is in force
- Leaves always-on hosts out of the denominator
- Assumes host patching covers firmware and management controllers
- Quotes an average instead of naming the tail hosts