skip to content

Your 9-node bare-metal Kubernetes cluster is 95% requested but only 20% busy, and finance wants two nodes back. How do you respond as the platform owner?

level: principalimportance: should knowfreq 30%

answer

  1. scheduler spends requests, not usage
  2. check N+1 before cutting
  3. 8.55 node-units over 8
  4. break the gap into components
  5. reduce requests, repack, then return

basics

~10 s

Not yet: the scheduler budgets requests, and at 95% requested the cluster cannot even absorb one node loss. Reduce requests first, repack, keep an N+1 ceiling near 89% requested, and only then return hardware.

solid answer

~50 s

The scheduler budgets **requests against Allocatable**, not usage, so the 20% figure buys nothing by itself. At 95% requested across nine nodes, the requests add up to 8.55 nodes' worth. Losing one node leaves 106.9% of eight nodes, so the cluster already fails N+1, and removing two nodes would push it to 122%. My answer is a sequenced plan. First, break the gap down: request inflation, reservations and DaemonSet overhead, fragmentation, and deliberate headroom. Second, get owners to lower inflated requests; that belongs to the right-sizing workstream. Third, repack so whole nodes become free, planning around pods that cannot move freely, such as the wiki's single-replica database. Fourth, return a node only when steady-state requests fit under the N+1 ceiling of 8/9 ≈ 88.9% of the smaller fleet, with a fragmentation margin. Lowering requests below real peaks is not a shortcut; it trades a hardware bill for node-pressure evictions.

go deeper

for a junior

Remember that Kubernetes schedules by requests, so a cluster can be full on paper while its machines are mostly idle.

for a middle

Show the requested-over-Allocatable arithmetic, including what happens when one node disappears.

for a senior

Break the gap into inflation, reservations, DaemonSets, fragmentation and headroom, and run the repack and node-loss test before any hardware leaves.

for a principal

Turn a cost request into a sequenced, owned plan with an explicit N+1 ceiling, and state the density, blast-radius and engineering-cost tradeoffs so the business decides with real numbers.

## Reframe the question Finance sees 20% busy and concludes 80% is wasted. The platform owner has to explain that Kubernetes does not allocate by usage. kube-scheduler admits a pod only if the node's **Allocatable** minus the **requests** already bound there covers the pod's requests. On this cluster the requests are the currency, and 95% of it is spent. The honest answer is "yes, once the requests come down", with a plan and a date. ## Do the arithmetic first Treat each node as one unit of Allocatable (for simplicity, assume identical nodes). 1. Requests in use: 0.95 × 9 = **8.55** node-units. 2. One node fails: 8.55 / 8 = **106.9%** of the surviving Allocatable. Some pods stay Pending, so the cluster **already fails N+1**. 3. Two nodes returned: 8.55 / 7 = **122.1%**. The pods do not fit even with every node healthy. 4. N+1 ceiling for nine nodes: 8 / 9 = **88.9%** requested. For seven nodes it is 6/7 ≈ 85.7%. Fragmentation means the practical ceiling is lower, because a failed node's pods must fit into per-node gaps, not the cluster total. This alone rules out returning hardware today and changes the conversation from savings to risk. ## Break down the gap between requested and busy | Component | What it is | Owner / lever | |---|---|---| | Request inflation | requests set well above real peaks | workload teams, via right-sizing (VPA recommendations, review) | | Node reservations | kube-reserved, system-reserved, eviction threshold, removed before Allocatable | platform: sized from measured daemon use, not guessed | | DaemonSet overhead | per-node agents requesting resources on every node | platform: fixed cost that does not shrink when you pack denser | | Fragmentation / stranding | free capacity in blocks too small or in the wrong shape | platform: repacking, request sizes, node shapes | | Deliberate headroom | spare capacity for node failure and bursts | platform + SRE: capacity policy | Measure each one. Usually the biggest is request inflation, and the platform team cannot fix it alone. ## The plan I would commit to 1. **Publish per-team requested-vs-used figures** (requests over Allocatable, usage from metrics) so inflation is visible where it happens. 2. **Lower requests to measured peaks plus margin**, team by team. Memory requests need the most care: if pods use more memory than the node can supply, the kubelet evicts pods. The internal wiki with an embedded database, requesting 2.6 GiB, should keep a request close to its real working set. 3. **Repack.** On bare metal no node autoscaler removes machines, so freeing a node is a deliberate cordon-and-drain. Scoring toward fuller nodes (`MostAllocated`), or a descheduler compaction pass, makes the empty node stay empty. Stateful single-replica pods, such as the wiki's database, are moved in a planned window, not by an automated evictor. 4. **Return a node only when** steady-state requests fit under the N+1 ceiling of the smaller fleet, with margin for fragmentation, and a node-loss exercise shows the pods reschedule. ## Tradeoffs to put on the table - **Density vs blast radius:** packing more pods per node means one machine failure displaces more pods and takes longer to recover. On a nine-node cluster each node is 11.1% of capacity. - **Overcommit by lowering requests:** CPU overcommit shows up as throttling and latency. Memory overcommit shows up as evictions and OOM kills. Some teams accept the first for batch work. Few accept the second for a database. - **Reservations are not waste:** shrinking kube-reserved or system-reserved to make Allocatable look bigger moves the risk onto the node, where daemons and pods compete. - **Engineering cost vs hardware cost:** two bare-metal nodes may cost less than a quarter of cross-team right-sizing work. Say so, and let the business decide with the numbers. ## What good looks like A shared dashboard shows requested and used per resource, the largest free block per node, and whether the fleet passes N+1 at current requests. Decisions to add or remove hardware are made against that, not against a single utilisation number. Forecasting demand and setting the organisation's headroom policy is capacity-planning work outside the cluster. The platform owner's part is to make the node-level numbers accurate.

  • Finance proposes simply cutting every team's memory requests in half. What happens on the nodes?
    The scheduler would pack roughly twice as many pods per node by request, but real memory use would not change. Nodes whose real use passes Allocatable hit the `memory.available` eviction threshold, and the kubelet evicts pods, starting with those using the most memory relative to their requests. Services see restarts and data-heavy pods suffer most. Memory requests must follow measured working sets, not a budget target.
  • Why is the practical N+1 ceiling lower than 8/9 of Allocatable on this cluster?
    The 88.9% figure assumes a failed node's pods can be spread over the survivors' free capacity in any split. In practice each pod needs one node with enough free room. Fragmentation, stranded CPU or memory, and constraints such as affinity or large requests like the wiki's 2.6 GiB can leave pods Pending below that ceiling. Test it by draining a node.

saying these in an interview costs you the question

  • At 20% busy, 80% of the cluster can be handed back safely
  • The scheduler will pack pods tighter automatically once nodes are removed
  • Shrinking kube-reserved is free capacity with no downside
  • Halving memory requests is a safe way to raise density
  • At 95% requested, a nine-node cluster still survives one node failure