Every host runs one kernel, so an upgrade is a whole-machine change — how do you set a kernel version policy for a shared fleet?
answer
- what unit a kernel change applies to
- cost scales with machines, not workloads
- uniform fleet against separate populations
- one shared version, one shared exposure
- how fast can the whole fleet restart?
basics
~20 sDecide two things: how much version fragmentation the estate will carry, and how fast the whole fleet can be turned over. One version is cheapest to certify and patch; each extra host population buys one workload's requirement at a permanent cost.
solid answer
~50 sThe unit of change is the machine, not the workload: a kernel changes by moving workloads off a host and restarting it, so the cost of any kernel decision scales with the number of hosts. That leaves two decisions. First, uniformity — one version across the fleet means one thing to certify, one patch train and unconstrained placement, while a second population satisfies a workload that needs a different kernel and lets a version be tried on a slice first, at the price of split capacity and a placement rule someone maintains indefinitely. Second, turnover speed: because every workload on a machine shares its kernel, an urgent fix leaves the estate exposed until the last machine restarts, so being able to roll the whole fleet within a known number of days is the real capability — and it is only real if it has been rehearsed.
go deeper
Know that a kernel upgrade means restarting whole machines with their workloads moved off first, which is why it is planned rather than shipped like an application change.
Explain what differs between machines when versions differ, and why a workload that requires a specific kernel now carries a placement constraint that someone has to maintain.
Show the operational plan: how workloads leave a machine, how long a full pass takes, how you verify the version reached every host, and what happens to machines that do not come back.
Make the trade explicit — fragmentation against capability — and commit to numbers: how many populations the estate carries, how quickly it can be turned over end to end, and who may add a requirement.
## Why kernel version is a fleet decision A workload can be replaced copy by copy, in minutes, with no one else involved. A kernel cannot. One machine carries exactly one kernel, serving every container on it, so changing that kernel means emptying the machine and restarting it. The cost of a kernel decision is therefore denominated in **machines**, not in workloads, and it is paid again for every machine in the estate. That single fact produces the whole policy question. Two things have to be decided deliberately, and both are the kind of trade a lead owns rather than a setting someone changes. ## Decision one: one version, or several populations | | One version across the fleet | Two or more host populations | |---|---|---| | Certification and patching | one train, one thing to qualify | one per population, forever | | Placement | unconstrained; anything fits anywhere | a constraint per requirement, maintained by someone | | Capacity | one pool absorbs every workload | fragmented; each pool sized for its own peak | | Trying a new version | the fleet moves together | a population can take it first | | Urgent fix | one pass over the fleet | one pass per population | The honest reading is that a second population is not free and is rarely temporary. It buys two real things — a home for a workload whose vendor demands a particular kernel, and the ability to expose a slice of the estate to a new version before all of it — and it charges capacity fragmentation, a duplicated patch train, and a placement rule that has to survive every future reorganisation of the platform. So the policy is a number, not a principle: **how many populations will this estate carry**, and who is allowed to ask for another one. Teams will ask; the request should come with what breaks without it and whether the requirement is permanent. ## Decision two: how fast can the whole fleet turn over This is the capability that actually matters, and it is the one most estates have never measured. When an urgent kernel fix appears, every workload on every unpatched machine is still sharing the affected kernel. The exposure does not taper as machines are done — it ends at the **last** machine. So the useful number is not what fraction is patched, it is the wall-clock time from decision to the last host restarting. That number exposes the things that will actually stop you: - machines that cannot be emptied because a workload on them has nowhere else to go; - workloads nobody is willing to move during business hours; - capacity so tight that emptying machines in parallel is impossible; - host populations that each need their own pass; - machines that do not come back, and the manual work of noticing and replacing them. A turnover that has only ever been done during an incident is not a capability, it is a hope. The fix is to exercise it on an ordinary, non-urgent version change and record how long it really took. ## What to write down 1. **The supported version set** — how many, why each one exists, and the date each stops being supported. 2. **Who may add a population**, what evidence the request needs, and who pays for the capacity it strands. 3. **The turnover target** — the time within which the whole fleet, all populations, can be restarted on a new version, measured rather than asserted. 4. **How a workload's kernel requirement is recorded**, so that a rebuilt machine or a capacity move does not silently violate it. 5. **What happens to a machine that fails to return**, because at fleet scale some will not. ## Where more populations is the wrong answer Sometimes the requirement should move instead of the fleet. A workload that wants a kernel nobody else wants may be better changed, dropped or run on a boundary that brings its own kernel — a separate decision with its own trade-offs — than serviced by a permanent island of machines. And sometimes the request is really about a host-wide setting rather than a version at all, which is a cheaper conversation with the same shape: a value that belongs to a machine, and tenants who must either agree on it or stop sharing.
- A team asks for a host population on a newer kernel for one workload. What do you ask for before agreeing?What breaks without it, and whether the requirement is permanent. Then price it out loud: an extra population is capacity that cannot absorb anyone else's workloads, a second version to qualify and patch, a placement rule maintained indefinitely, and one more pass during every urgent fix. Agreeing is a standing commitment rather than a one-off change.
- Why is the share of machines patched a weaker measure than the time to turn the whole fleet over?Because exposure ends only at the last machine. Every workload on an unpatched host is still sharing that kernel, so a fleet that is ninety per cent done still has a fully exposed population. Turnover time measures the capability actually needed during an urgent fix, and it surfaces the slow parts: machines that cannot be emptied and workloads nobody will move.
saying these in an interview costs you the question
- Treats a kernel upgrade as a per-workload rollout.
- Adds a host population per team request without counting the standing cost.
- Assumes an urgent kernel fix can skip moving workloads off first.
- Calls the estate safe at ninety per cent patched, though exposure ends at the last host.
- Never rehearses a fleet turnover and improvises it during an incident.