skip to content

On a shared cluster, when is dedicating a group of hosts to one workload class worth the capacity it strands?

level: principalimportance: should knowfreq 32%

answer

  1. two rules, not one
  2. attraction does not exclude
  3. every island needs its own spare
  4. hardware fit, not importance
  5. exemptions erode the guarantee

basics

~20 s

Dedicated hosts are worth it when the workloads genuinely differ in the hardware they need, the interference they can tolerate, or the separation required of them — not when one team simply wants to be treated as more important.

solid answer

~50 s

Doing it properly takes two rules, not one: an attraction rule so the chosen workloads land only on hosts carrying a particular label, and a marker on those hosts that repels everything else unless a workload carries an explicit exemption. The attraction rule alone keeps your workloads in but keeps nobody else out, which is the common half-done version. The cost is structural: every island needs its own failure headroom, spare room in one island cannot rescue another, leftovers per island are smaller so large reservations go pending sooner, and there is a second capacity plan to own. That price is worth paying when the hardware really differs, when a licence or compliance boundary is defined in terms of hosts, when interference the platform does not model is intolerable, or for workloads that keep the cluster itself running. It is not worth paying for importance.

go deeper

for a junior

Know the shape of the mechanism: hosts can be labelled and marked so that only certain workloads land on them, and that capacity is then unavailable to everything else.

for a middle

Be able to explain why two rules are needed — one to pull the chosen workloads in, one to keep the others out — and why the attraction rule alone gives no guarantee at all.

for a senior

Show the operational cost: per-island headroom, smaller leftovers, more pending workloads, and an exemption list that quietly erodes the boundary unless somebody audits it.

for a principal

Own the trade-off and the reversal. Name the property being bought, price it in utilisation across the estate, start with the smallest island that delivers it, and set a review that can undo it.

## What dedicating hosts actually requires - A label on the chosen hosts, and an **attraction rule** on the workload class so its copies land only on hosts carrying it. - A **repel marker** on those hosts, so every other workload is filtered out unless it carries an explicit exemption. - **Both**, because the two do different jobs. An attraction rule constrains your workloads; only the marker constrains everyone else's. With the rule alone, the group looks dedicated right up to the day a busy neighbour fills it, and the first symptom is your own copies going pending on capacity you believed was reserved. - Someone who owns the group's size, its upgrades and its exemption list, because all three drift. ## What it costs | | one shared pool | several dedicated islands | |---|---|---| | spare capacity | one pool absorbs any host failure | each island needs its own spare host | | leftovers | many hosts, larger usable gaps | smaller gaps per island, more pending | | utilisation | high, because demand mixes | lower, by construction | | reasoning | one place to look when something is pending | which island, whose rule, whose exemption | Two costs are easy to underestimate. **Headroom multiplies**: a fleet that keeps one host's worth of slack needs that slack in every island, and slack in one island cannot help another. And **rules rot**: the exemption someone adds to unblock a pending workload on a Friday ends the guarantee silently, because nothing reports that a dedicated group stopped being dedicated. ## When it is genuinely right 1. **The hardware differs.** Accelerators, fast local storage, a particular processor architecture. Here the attraction rule is doing real work — it is a fit, not a privilege — and the repel marker keeps ordinary workloads from consuming scarce hosts. 2. **A licence or a compliance boundary is defined in terms of hosts.** If the obligation is written about machines, placement is the mechanism that implements it, and best-effort is not an answer. 3. **The interference is intolerable and cannot be bounded by reservations and ceilings.** Those cover what the platform models; memory bandwidth, cache and the network path are shared regardless. A workload with a hard latency commitment can be the case where a boundary at host level is the only honest one. 4. **The workload keeps the cluster running.** Components the platform itself depends on should not lose a scheduling race with tenant work. ## When it is the wrong answer - **Importance.** Importance is not a placement property, and carving out hardware to express it buys nothing except lower utilisation. - **An organisational boundary.** One island per team is an accounting wish solved with scoped budgets, not with hardware. - **An unmeasured latency complaint.** Measure first; inflated or missing reservations explain a great many of these, and cost nothing to fix. - **Fear.** A dedicated island bought against an incident nobody has had is a permanent cost against a hypothetical one. ## Deciding, and being able to reverse it 1. State the property you are buying in one sentence, and how you would detect it being violated. If it cannot be detected, it cannot be defended later. 2. Price it honestly: utilisation across the estate before and after, plus the spare host per island, plus the pending risk from smaller leftovers. 3. Start with the smallest island that delivers the property. A group of three hosts is much cheaper to be wrong about than a third of the fleet. 4. Put a review date on it and audit the exemption list, because exemptions accumulate and each one moves the group back towards being a shared pool with extra steps. 5. Keep a way back. Removing the marker and the attraction rule should be one change, and it should be obvious who is entitled to make it. ## The cheaper answers to try first Most requests for dedicated hardware are really requests for predictability, and predictability has cheaper sources. Right-sized reservations stop a host being stacked beyond what it can serve. A workload declaring no reservation at all is invisible to the scheduler's arithmetic and is a common hidden cause of the interference being complained about. Spreading a sensitive workload's copies so that no two share a host removes the worst case without removing anyone else's access to the fleet. And a scoped budget stops one team consuming the pool without carving the pool up. Exhaust those before buying hardware separation, and when you do buy it, buy it for a property you can name and measure.

  • What goes wrong if you label the hosts and add the attraction rule but skip the repel marker?
    Your workloads land on the group, and so can everything else. The guarantee looks real until a busy neighbour fills the hosts, and the first symptom is your own copies going pending on capacity you believed was reserved for them.
  • How would you measure whether a dedicated island is earning its keep?
    Compare booked against measured usage inside the island with the shared pool, count workloads that went pending in each, and check whether the property you bought actually improved — the latency percentile, the separation, the hardware utilisation. Also audit how many exemptions have been granted since.
  • Is dedicating hosts the same decision as running a separate cluster for that workload?
    No, and the cost curve is different. Dedicated hosts keep one deciding half, one upgrade path and one set of rules, and strand only capacity. A separate cluster duplicates the operational surface as well, so it is justified by a stronger requirement than interference alone.

saying these in an interview costs you the question

  • Says an attraction rule on its own reserves the hosts for that workload
  • Offers dedicated hosts as the fix for a team wanting to be prioritised
  • Ignores that each island needs its own failure headroom
  • Assumes utilisation is unaffected because the hosts are still in the cluster
  • Grants the exemption broadly to unblock a pending workload, ending the guarantee
  • Cites latency interference that nobody has measured