A few percent of your estate never completes a reboot cycle and the owners refuse downtime. What do you do?
answer
- a downtime budget, not a patch queue
- the tail is never random
- hypervisors, frozen appliances, controllers
- buy the redundancy or mandate the cycle
- acceptance needs a name and a date
basics
~20 sTreat it as a downtime-budget decision, not a patching one. Either fund the redundancy that makes restarts cheap, mandate a periodic power cycle as a condition of hosting, or take a dated acceptance from the owner who refuses.
solid answer
~50 sFirst I stop calling it a patching problem. Those hosts receive updates fine; what they never do is restart, and the reason is a business one, so the decision belongs to whoever owns the downtime. Second I look at who is in the population, because it is never random: always-on virtualisation hosts, appliances inside a change freeze, machines that are only suspended, and management controllers that were never in the cadence at all. That is the most valuable set of machines in the estate, and an unhurried adversary needs no new flaw while it exists. Then I make the owner choose between three real options: fund the redundancy that makes a restart cheap, accept a mandatory annual power cycle as a condition of running there, or sign a dated acceptance naming the flaws left loaded. What I will not do is leave the refusal implicit inside a green patch number.
go deeper
Know that some hosts genuinely never restart for business reasons, and that they stay on old code until they do - so their patch status is not settled by an install.
Be able to describe who is in that population and why: always-on virtualisation hosts, frozen appliances, suspended machines, and management controllers on their own lifecycle.
Show that you name the hosts individually with owners rather than averaging them, and that you know suspend-and-resume is not a restart.
Own the trade openly: fund the redundancy that makes restarts cheap, negotiate a standing power-cycle condition, or take a dated acceptance from the owner who refuses - and never make the number better by shrinking the denominator.
## Why this is a principal-level question Nothing here is technically hard. Everyone in the room already knows the fix needs a restart. The question is what you do when the restart is refused for reasons that are legitimate, by people who do not work for you, on the machines that matter most. That is a budget and ownership problem wearing a patching costume, and the failure mode is that it never gets escalated because the patch number stays green. ## Know who is actually in the population The standing few percent that never completes a cycle is not a random tail. It is composed of: - **Always-on virtualisation hosts.** Restarting one means moving or stopping every guest on it. Where there is no spare capacity to evacuate onto, the restart has a direct cost in either downtime or hardware. - **Appliances inside a long change freeze.** Frozen for a quarter-end, a migration or a regulatory event, and then frozen again. - **Machines that are only ever suspended.** Resume is not boot; the same kernel and the same mappings come back. - **Out-of-band management controllers.** These were never in the host cadence in the first place. They run their own software on their own processor, remain powered while the host is off, and can power, reimage and console the host. Read that list as an adversary would. It is the hypervisor that holds every guest, the appliance that terminates management traffic, and the controller that owns the host below the operating system - all guaranteed to stay on old code indefinitely. An operator already holding one identity's worth of access on the management segment, with time and no revenue clock, does not need a fresh flaw: the estate keeps this population stocked on its own, and nobody counts it as unpatched. ## The three honest options **1. Buy away the cost.** The reason a hypervisor never restarts is usually the absence of somewhere to put the guests. Spare capacity, an extra appliance in a pair, or a rolling-replacement pattern for control-plane nodes converts a restart from an outage into a scheduling exercise. This is the only option that actually shrinks the population, and it is a capital conversation, not a security one - which is why it needs the security case attached to it. **2. Make a periodic power cycle a condition of hosting.** Not a request per fix, but a standing rule: anything running in this estate is restarted at least on a stated cycle, and workloads that cannot survive that are architected to move or are hosted somewhere that accepts the risk explicitly. This works because it is negotiated once, in the abstract, rather than fought per advisory when everyone is defensive. **3. Accept, with a name and a date.** Where neither of the above is affordable, the correct outcome is a written acceptance from the workload owner - the person who benefits from the uptime - that names what is left loaded and when the decision is revisited. The point is not the paperwork; it is that the refusal becomes visible and attributable instead of being averaged into a percentage. A fourth, partial move sits alongside all three: reduce who can reach these machines at all, particularly the management plane, so that the standing staleness is harder to convert into access. Treat that as a mitigation that buys time, never as a substitute for the restart. ## What to say, and to whom The escalation only works if it is expressed in the owner's currency. To a workload owner: "restarting this host costs you N minutes a year; not restarting it means these specific flaws stay loaded on the machine that holds all your guests." To a finance owner: "the redundancy that makes the restart free costs X, and it also removes the unplanned outage risk you already have." To an executive asking why the patch number is green while the review is not: "the number counts distribution; here are the eleven hosts that have never run the code we distributed, and here is who declined each one." ## The trap to avoid Do not solve this by shrinking the denominator. Excluding the always-on hosts from the population being measured makes the number beautiful and the estate worse, and it is the single most common organisational response to this problem. The named exception list is the deliverable; the percentage is not.
- Why is this population disproportionately attractive to an adversary with time?Because it is guaranteed to stay on old code and it contains the estate's highest-value machines: hypervisors holding every guest, appliances on the management path, and controllers that own the host beneath its operating system. An operator with no revenue clock does not need a new flaw while the estate reliably produces stale, uncounted hosts.
- Who signs the acceptance when a hypervisor cannot be restarted?The workload owner who is refusing the downtime, not the team running patch cadence. They hold the benefit, so they should hold the risk, and the acceptance should name the specific flaws left loaded and a date it is revisited. Security's job is to make the trade legible, not to absorb it silently.
- Is excluding these hosts from the patch metric ever defensible?No. Excluding them makes the number rise while the risk stays exactly where it was, and it removes the only pressure that ever gets the restart funded. Keep them in the denominator and publish them as a named exception list with owners; the list is the useful artefact, the percentage is not.
saying these in an interview costs you the question
- Treats it as a tooling gap rather than a downtime-budget decision
- Excludes the never-restarted hosts from the measured population
- Chases each advisory individually instead of negotiating a standing cycle
- Forgets management controllers are not in the host cadence at all
- Accepts the risk without a named owner or a revisit date