skip to content

questions

4

One shared cluster carries every team's streams in your organisation — what does that put in common, however the names inside it are divided?

level: middleimportance: must knowfreq 62%

answer

  1. one deployment, one bad day
  2. names divide, machines do not
  3. version and upgrade window are shared
  4. administrators reach the whole cluster
  5. blast radius equals the cluster

basics

~20 s

A shared cluster puts one set of nodes and volumes, one capacity ceiling, one deployed version and upgrade window, one administrative surface and one on-call rota behind every stream on it. Dividing the names inside changes none of that.

solid answer

~50 s

A cluster is a deployed set of nodes serving streams behind one endpoint set, and everything about that deployment is common to every stream on it: the machines and their storage, the capacity they add up to, the version they run, the window in which that version is changed, the settings that belong to the deployment rather than to a stream, and the people who are paged when it misbehaves. Dividing the estate inside the cluster — by a named space where the platform has one, or by a name prefix plus grants where it does not — separates names, grants and quotas. It does not separate hardware, capacity, versions or failures. So the honest statement of blast radius is: whatever takes the cluster down takes every team on it down together, and whatever upgrade goes badly goes badly for all of them at once.

go deeper

for a junior

Recall that a cluster is a set of machines serving streams, and that every stream on it shares those machines. If one of them is lost, the effect does not stop at one team's names.

for a middle

Explain the split: names, grants and quotas are divided inside a cluster, while nodes, storage, capacity, the deployed version and the upgrade window are not. Be able to say which side of that line a given concern falls on.

for a senior

Show that you can state blast radius concretely for a real estate — who is paged, who loses traffic, who has to be told about an upgrade — instead of listing isolation features the platform advertises.

for a principal

Frame the count of clusters as a standing policy: what must be true about a domain, an environment or a customer before it earns a deployment of its own, and who pays for the operating load each additional one creates.

## What a cluster is, in this tree's words **The estate** is every cluster and every stream an organisation runs, taken together. **A cluster** is one deployed set of nodes serving streams behind one endpoint set. The shared-against-dedicated question is about how many clusters that estate actually is: one for everybody, or one per domain, per environment, or per team or customer. Everything else in estate governance — what a stream may be called, how it is created, who answers for it — happens *inside* whichever answer you gave. The reason this question comes first is that the cluster is the unit that fails, the unit that is upgraded, and the unit that is paid for. Nothing you do inside it changes that. ## What is genuinely common to every stream on one cluster - **The nodes and their storage.** Streams from every team have their records on the same machines and the same volumes. Which stream sits where is decided by the platform, not by a name. - **The capacity ceiling.** Ingress, egress, connections and stored bytes add up to one total, and that total has one limit. - **The deployed version.** A cluster runs one version at a time, and a change of version moves the whole deployment. - **The maintenance calendar.** An upgrade, a node replacement or a settings change happens at a moment somebody chooses, and every team on the cluster lives with that moment. - **The settings that belong to the deployment** rather than to an individual stream — the ones an operator sets once for the cluster, whose names differ from platform to platform but whose scope does not. - **The administrative surface.** The credentials that can administer the cluster can, by definition, reach every stream on it. Careful grants narrow what ordinary clients may do; they do not narrow what an administrator may do. - **The operating team and its on-call rota.** One group of people knows this deployment, and one rota is woken for it. - **The incident.** If it is down, it is down for everyone on it. ## What a shared cluster does still divide This is the half candidates usually get right, and it is worth stating so the boundary is clear: - **Names.** Every stream has a distinct name, and where the platform offers a first-class named space, names live inside one. - **Grants.** A grant binds a principal to an operation on a stream or a name prefix, so one team's clients need not be able to read another's streams. - **Quotas**, where the platform can attach one to a space or a prefix. - **Per-stream settings** such as how long records are kept and how many stored copies each record has — these are chosen per stream on most platforms, not once for the cluster. Some designs make retention a property of each subscriber instead; the point is only that it is not a single cluster-wide value. | Concern | Shared cluster | One cluster per domain | |---|---|---| | A node is lost | Every team feels it, in whatever copies that node held | Only the domain on that cluster | | Volumes approach full | One pool; one team's growth is everyone's problem | Contained to one deployment | | Version upgrade goes badly | The whole organisation at once | One domain at a time | | An administrative credential leaks | The entire estate on that cluster | One domain's streams | | Recurring operating work | Paid once | Paid once per cluster | ## The distinction that trips people up A name prefix *is* the isolation on platforms that have no first-class space — but only as far as the grants written against it. It never places a stream on particular hardware, never reserves capacity, and never gives a team its own version or its own upgrade window. Capacity contention *between* clients on a shared cluster has its own controls and its own subject; those controls are worth having, and they still do not change a single row of the table above. The symmetric error is to believe a dedicated cluster divides something it does not. It divides failures, capacity, versions and administrative reach. It does not divide the work of running a cluster — that it multiplies. ## What an interviewer is listening for 1. That you name the cluster as the unit of failure, capacity, version and administration, rather than reciting isolation features. 2. That you can say, without hedging, who is affected when one node is lost on a cluster two teams share. 3. That you separate what a name buys (names, grants, quotas) from what only a second deployment buys. 4. That you treat the count of clusters as an organisational decision with a cost on both sides, not as a best practice with one right answer.

  • If grants are written carefully, what can an administrative credential on a shared cluster still reach?
    Everything on it. Grants constrain the principals they are written for; the credentials that administer the deployment exist to create, alter and delete streams and grants, so their reach is the cluster. That is a property of the deployment boundary, and the only way to shrink it is a second deployment.
  • Does a first-class named space give a team its own capacity?
    No. It scopes names, grants and, where the platform supports it, quotas. A quota limits how much a client may take, which is a fairness control and not a reservation of machines. The nodes, their storage and the cluster's total ceiling stay common to every space on it.
  • Why does the deployed version matter more than most of the other shared properties?
    Because it is the one shared property that changes on a date somebody picks. Capacity and hardware fail unpredictably; a version change is a planned event that every team on the cluster must absorb together, and it cannot be staged team by team on a single deployment.

saying these in an interview costs you the question

  • Thinks a name prefix places streams on separate machines
  • Says a named space gives a team its own capacity
  • Believes careful grants shrink an administrator's reach
  • Treats a shared cluster as isolated because names do not collide
  • Assumes each team can run its own version on one cluster
  • Says retention and copy count are cluster-wide values
open as a page

A platform team proposes one cluster per domain to shrink blast radius — which recurring cost does that multiply, and which does it not?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Splitting multiplies operating load, not traffic. Every cluster needs its own upgrade rounds, grant set and credential rotations, alert coverage, capacity reviews, drills and on-call knowledge, plus a minimum node count and spare headroom. The records written stay the same.

open as a page

Two years in, you split one shared cluster into one per domain — what does that migration cost that the same decision at birth would not have?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Everything already running has to move. A second cluster starts empty, so streams, settings and grants are recreated, every producer and consumer is re-pointed and redeployed on its own team's schedule, restart points are chosen by hand, and the old cluster stays until nothing uses it.

open as a page

Non-production traffic shares your production broker cluster — what makes environment the split most estates buy first, and what does keeping it cost?

level: principalimportance: should knowfreq 44%

basics

~20 s

A version or settings change has to be rehearsed somewhere that is not production, and that argument holds whatever a second cluster costs. Sharing also puts unpredictable non-production traffic and the irreversible operations on production's nodes, guarded only by grants.

open as a page