skip to content

Your clinical-records portal runs on a three-node on-prem kubeadm cluster two minor releases behind; how do you choose between in-place upgrades and a replacement cluster, and set a cadence?

level: principalimportance: should knowfreq 40%

answer

  1. quorum and capacity per node window
  2. rollback path differs sharply
  3. spare hardware decides a lot
  4. operators are the long pole
  5. cadence fixes the root cause

basics

~20 s

Two minors behind usually means two in-place kubeadm hops, rehearsed first, if the portal fits on two nodes; a replacement cluster pays off when spare hardware exists, rollback matters most, or the gap is larger. Then upgrade every minor.

solid answer

~50 s

Two minors behind is still inside the supported window, so in-place is usually right: two rehearsed kubeadm hops, one node at a time. The costs are real on three stacked nodes: each node window leaves etcd with no spare member and the portal on two thirds of the capacity, and there is no easy way back once a hop completes except restoring an etcd snapshot. A replacement cluster at the target version avoids both hops, gives a clean rollback by switching traffic back, and tests the recovery path - but on-prem it needs spare machines, and moving stateful data and the operators' custom resources is its own project. I decide on spare hardware, how stateful the portal is, how much downtime compliance allows, and how far behind we are. Then I fix the root cause with a cadence: one minor per release, rehearsed on a staging copy, gated on removed-API checks and operator compatibility.

go deeper

for a junior

Know that a cluster can be upgraded in place or replaced, and that upgrades go one minor at a time.

for a middle

Explain what each option costs on three stacked nodes: etcd quorum and capacity for in-place, spare hardware and data migration for a replacement.

for a senior

Show how you would de-risk the chosen path: staging rehearsal, etcd snapshots, removed-API and operator checks, one node at a time.

for a principal

Weigh hardware, statefulness, rollback needs and the support window explicitly, then turn the rescue into a standing per-release cadence with clear owners.

## The situation A clinical-records portal runs on a **three-node on-prem kubeadm cluster**; each node is a stacked control-plane node that also runs portal workloads. The operators the portal depends on manage a 612-object fleet of custom resources. The cluster is **two minor releases behind** the current one. The project maintains only the three most recent minors, with roughly a year of patch support each, and ships about three minors a year - so the cluster is on the last supported rung and loses support at the next release. There is no single right answer; the job is to make the trade-offs explicit. ## Option A: in-place, two hops The cluster moves one minor at a time with kubeadm, each hop touching all three nodes. - **Pros:** no new hardware; the cluster's identity, addresses, certificates and storage stay where they are; the procedure is well documented and scriptable. - **Cons:** during each node's window etcd runs two of three members, so a second failure loses quorum; the portal must fit on two nodes; after a hop completes, going back means restoring an etcd snapshot and accepting the gap since it was taken. - **Hidden work:** every operator must support each intermediate minor, and every caller of a removed API version must be fixed before the hop that removes it. ## Option B: replacement cluster A new cluster is built directly at the target minor, workloads are redeployed from Git, data is moved, and traffic is switched. - **Pros:** no intermediate hop; rollback is switching traffic back while the old cluster still runs; the move doubles as a real test of rebuilding the platform. - **Cons:** on-prem it needs spare machines or a temporary shrink; persistent data and the operators' custom resources must be migrated and verified; load-balancer addresses, DNS, certificates and client credentials change. - **Hidden work:** proving that Git and backups really hold everything the old cluster held. ## How to decide | Factor | Leans in-place | Leans replacement | |---|---|---| | Spare hardware | none available | a second set of nodes available | | Gap to target | one or two minors | several minors, or unsupported | | Statefulness | data outside the cluster | little state, or easy to replicate | | Rollback requirement | snapshot restore is acceptable | fast, clean rollback required | | Platform confidence | rebuild never rehearsed and risky | rebuild automated and tested | For this portal - no spare hardware assumed, two minors behind - in-place is the default, with three safeguards: 1. **Rehearse** both hops on a staging copy built from the same manifests. 2. **Snapshot** etcd before each hop and confirm the restore procedure works. 3. **Schedule** each hop in an agreed change window, one node at a time, with the portal's capacity checked against two nodes. ## Fixing the cause: cadence Being two behind is the symptom; the policy is the fix. - **One minor per release**, a few weeks after it ships, once the first patch releases have landed. - **Readiness gates** before every hop: removed-API callers cleared, operator compatibility confirmed, release notes read by the owning team. - **Decouple node work where it helps.** Kubelets may trail the API server by up to three minors, so the control-plane step can happen on every release while the drain-and-restart kubelet step is batched - never letting the gap reach the edge. - **Automate the procedure** so the steps are run the same way every time and a hop is routine rather than an event. - **Budget capacity** for one node out at any time; three nodes that are each two-thirds full cannot absorb a drain. ## What to watch for in the answer A strong answer names the constraints before choosing, treats the operators and removed APIs as the usual long pole, and ends with a cadence rather than a one-off rescue. ## Signals that the choice was wrong - In-place hops that keep needing emergency fixes suggest the platform has drifted from what staging rehearses. - A replacement that takes months suggests the cluster held state that Git and backups did not. - Either way, the review after the move should change the cadence, not just close the ticket.

  • How would you use the kubelet skew window in the cadence without creating drift?
    Upgrade the control plane every release, and batch the disruptive kubelet upgrades every second release, so kubelets are at most one or two minors behind. Track the gap on a dashboard and block any control-plane hop that would push a kubelet past three minors.
  • What would make you pick a replacement cluster even with only a two-minor gap?
    Spare hardware already available, a portal whose data lives outside the cluster, a compliance requirement for a fast clean rollback, or a need to change something kubeadm cannot change in place - such as the node operating system or cluster networking. Then the extra migration work buys a safer path.

Choosing between the two is like renovating a house you live in versus building next door and moving: renovating needs no second plot but you live with the dust, while moving needs land and a big move-in day but you can walk back if the new house leaks.

saying these in an interview costs you the question

  • Always rebuild the cluster; in-place upgrades are never safe
  • Skip straight to the newest release to save a hop
  • Upgrade two stacked nodes together to shorten the window
  • Operators and custom resources need no checks during upgrades
  • Staying two minors behind is fine indefinitely