skip to content

A platform team proposes one cluster per domain to shrink blast radius — which recurring cost does that multiply, and which does it not?

level: seniorimportance: must knowfreq 55%

answer

  1. traffic stays, overhead multiplies
  2. per-deployment work, not per-byte work
  3. upgrades, credentials, alerts, reviews, rehearsals
  4. each cluster has a node floor
  5. skew is the late bill

basics

~20 s

Splitting multiplies operating load, not traffic. Every cluster needs its own upgrade rounds, grant set and credential rotations, alert coverage, capacity reviews, drills and on-call knowledge, plus a minimum node count and spare headroom. The records written stay the same.

solid answer

~50 s

The bytes do not multiply — the same records are produced and consumed either way, just spread across more deployments. What multiplies is the recurring work of owning a deployment: a version round per cluster, a grant set and credential rotations per cluster, alerting and dashboards per cluster, a capacity review per cluster, a recovery rehearsal per cluster, and a rota that has to know all of them. Where you run the nodes yourself there is also a floor: each cluster needs enough nodes to hold a copy set and to survive losing one, so several small clusters consume more machines than a single cluster carrying the same traffic. The effect that shows up months later is version skew — the least-loved clusters stop being upgraded, and the estate ends up running several versions with different behaviour. Operating load is the term teams forget when they argue for a split.

go deeper

for a junior

Recall that every cluster is a thing somebody has to upgrade, watch, secure and be paged for. Two clusters means two of each of those jobs, even if the amount of traffic has not changed at all.

for a middle

Explain the two piles: costs that follow traffic stay flat when an estate splits, while costs that follow the number of deployments multiply. Be able to name several items in the second pile without prompting.

for a senior

Demonstrate that you cost a split before proposing one — upgrade rounds, rotations, reviews and rehearsals per year at N clusters — and that you recognise version skew as the failure mode of an estate with more clusters than operators.

for a principal

Set the rule for which boundaries earn a deployment at all, knowing that each one is a permanent claim on the platform team's capacity, and decide what the organisation will stop doing to pay for the ones you approve.

## Two different kinds of cost The shared-against-dedicated argument is usually run on blast radius alone, which is why it is so often won by the side proposing more clusters. The counterweight is not the price of hardware. It is the **operating load**: the recurring human and machine work that a deployment demands simply because it exists, independently of how much traffic it carries. A useful way to see it is to sort every cost into one of two piles: costs that follow the **traffic**, and costs that follow the **number of deployments**. Splitting an estate leaves the first pile alone and multiplies the second. ## Costs that follow the traffic — unchanged by a split - The records produced and consumed. The same writers write the same things. - The stored bytes those records occupy, given the same retention and the same copy count. - The number of streams the organisation needs; domains need what they need. - The number of reading groups; readers follow consumers, not deployments. ## Costs that follow the number of clusters — multiplied by a split 1. **Version rounds.** Every cluster must be taken through an upgrade, with its own rehearsal, its own window and its own rollback thinking. Eight clusters means eight of those a release. 2. **Grants and credentials.** Each deployment has its own set of grants and its own credentials, each with its own rotation and its own review. 3. **Alerting and dashboards.** Signals are per cluster. So are the thresholds, the silences that were never removed, and the alert nobody has re-pointed since the split. 4. **Capacity review.** Each cluster's growth has to be looked at separately, because a cluster with room to spare cannot lend it to one that has none. 5. **Recovery rehearsals.** Whatever the standard is for proving a cluster can be restored, it is per cluster. 6. **Rota knowledge.** The same people are woken for all of them, and the deployment they are woken for at three in the morning is usually the one they touch least. 7. **Hardware floor.** Where you run the nodes yourself, a cluster needs a minimum node count to hold a copy set and to keep serving when one node is lost, plus enough spare headroom to absorb that loss. That floor is paid per cluster, so several small clusters consume more machines than one carrying the same traffic. | Cost | Follows | Effect of splitting into N clusters | |---|---|---| | Records written and read | Traffic | Unchanged | | Stored bytes | Traffic and retention | Unchanged | | Upgrade rounds per release | Deployments | Multiplied by N | | Grant sets and credential rotations | Deployments | Multiplied by N | | Alert coverage and dashboards | Deployments | Multiplied by N | | Minimum nodes and spare headroom | Deployments | Multiplied by N | ## The cost that arrives late: version skew The day-one costs of a split are visible and get planned for. Operating load is different, because it does not arrive on any particular day — it arrives as a slow degradation of the estate. Upgrade work multiplied by N is not done N times; it is done for the clusters somebody cares about and postponed for the rest. Within a year the estate is running several versions, with different behaviour, different bugs and different settings available, and every runbook has an exception list. The second late arrival is uneven capacity. Traffic never distributes itself the way the domain split assumed, so some clusters run hot while others sit idle, and the idle headroom cannot be moved. ## How to make the argument honestly The decision is not shared-or-dedicated as a principle. It is: **which boundaries earn a deployment**, given that each one costs an ongoing share of the platform team. - Start from the boundaries where blast radius is not negotiable — typically the environment boundary, and anything a regulator or a customer contract names. - Ask what the split is *for*. If the answer is 'so one team cannot affect another's throughput', that is a fairness problem with its own controls and does not require a deployment. - Count the load before committing: how many upgrade rounds, rotations, reviews and rehearsals per year does N imply, and who is doing them? - Prefer the smallest number of clusters that satisfies the boundaries you actually have to honour. Every extra one is a standing commitment, not a one-off purchase.

  • A team argues the split costs nothing because the total traffic is unchanged. What is the reply?
    Traffic is only one of the two piles. The other is per-deployment work: upgrade rounds, grant sets and rotations, alert coverage, capacity reviews and rehearsals, each multiplied by the number of clusters. Where you run the nodes, each cluster also carries a minimum node count and its own spare headroom.
  • Which symptom tells you an estate has more clusters than it can operate?
    Version skew. When clusters are running several different versions because upgrades are done for the important ones and postponed for the rest, the estate has outgrown the team operating it. Uneven capacity is the companion symptom: some clusters hot, others idle, with no way to move the headroom.
  • Does renting the clusters instead of running them remove the operating load?
    It removes some of it and leaves the rest. Node replacement and patching move to the provider, but grants, credentials, alerting, capacity decisions, stream governance and the rota that answers for the application still exist once per deployment. The multiplier applies to the half that stays yours.

One apartment building against a row of houses. The same families live in the same number of rooms either way, so the living space does not change. What changes is that each house has its own roof, its own heating and its own insurance, and somebody has to remember to service every one of them — and the house nobody visits is the one whose boiler fails.

saying these in an interview costs you the question

  • Argues the split is free because total traffic is unchanged
  • Counts only hardware and omits the recurring human work
  • Assumes idle headroom on one cluster helps another
  • Splits to stop one team crowding another's throughput
  • Plans the migration but not the standing upgrade load
  • Expects all clusters to stay on the same version by default