Two years in, you split one shared cluster into one per domain — what does that migration cost that the same decision at birth would not have?
answer
- the new cluster starts empty
- not a copy, a fresh deployment
- every client is somebody else's release
- restart points chosen by hand
- old cluster runs until proven unused
basics
~20 sEverything already running has to move. A second cluster starts empty, so streams, settings and grants are recreated, every producer and consumer is re-pointed and redeployed on its own team's schedule, restart points are chosen by hand, and the old cluster stays until nothing uses it.
solid answer
~50 sAt birth the decision costs a conversation. Later it costs a coordinated migration across every team that uses the cluster. A second cluster is not a copy of the first: it starts empty, so the streams, their settings and every grant have to be established again. Each producer and consumer needs a new endpoint and new credentials, which means a release per client, on each owning team's own schedule. The history already written stays on the old cluster — where records remain readable after delivery, it simply does not come with you, and where delivery removes the record, whatever is still undelivered has to be drained before the old cluster can go. Where readers own a stored position, that position means nothing on a cluster that never held those records, so somebody chooses a starting point for each reading group and accepts either reprocessing or a gap. And the old cluster keeps running, and keeps costing, until you can show nothing is still using it.
go deeper
Recall that a second cluster is a brand-new, empty deployment. Nothing that was written to the first one appears on it, and every client has to be told the new address before it can use it.
Explain what must be recreated — streams with their creation-time settings, grants, credentials, alerting — and why the records already written stay behind on the old deployment rather than following the streams.
Show that you plan the migration around other teams' release schedules, that you make each reading group's start point an explicit decision with a named trade, and that you switch the old cluster off on evidence of disuse rather than on a date.
Draw the policy conclusion: the layout is cheapest at birth, so decide it early and encode it where new streams are created, because every stream placed on the wrong cluster today is a client migration somebody inherits.
## The decision is cheap exactly once How many clusters an estate runs is decided at birth or paid for later. At birth it is an argument and a line in a standard. Later it is a migration, and the migration is the expensive part, not the deployment. The root of the expense is a single fact that candidates often miss: **a second cluster is not a copy of the first.** It is a new, empty deployment. Nothing about it knows that the first one exists. ## What has to be recreated on the new cluster 1. **The streams themselves**, with their settings — how long records are kept, how many stored copies each has, and the parallelism count where the platform has one. These are settings stamped at creation, and on many platforms the parallelism count is awkward to change afterwards, so getting it wrong during the move is a second migration waiting to happen. 2. **Every grant**, binding each principal to the operations it needs on each stream or name prefix. Grants are written against the cluster they live on; they do not travel. 3. **Credentials** for every client that will connect, and the distribution of those credentials to the teams that own the clients. 4. **Everything operational that names the cluster**: alert rules, dashboards, capacity records, the rehearsal that proves it can be recovered, and the rota entry that says who answers for it. ## What does not come with you - **The history.** On a platform where records remain readable after delivery, everything already written stays on the old cluster. If a consumer relies on being able to re-read the last N days, that ability does not exist on the new cluster until N days have passed on it. - **Undelivered messages, where delivery removes the record.** Here the mirror-image problem applies: the new cluster starts empty, and the old one still holds whatever nobody has consumed. It has to be drained before the old deployment can be switched off, which puts a queue-shaped migration on the consumers' clock rather than the platform team's. - **Stored read positions, where readers own one.** A position is recorded against the cluster that held the records. On the new cluster it means nothing, so each reading group needs an explicit starting point, and the choice is a trade: start at the beginning of what exists and reprocess, or start at the newest record and accept that anything written during the cut is not read. | What | At birth | Two years in | |---|---|---| | Streams and their settings | Created once, on the right cluster | Created again, and the first version must keep running | | Grants and credentials | Written once | Written again, and distributed to every client owner | | Clients | Configured with one endpoint | Re-pointed and released, team by team | | History | None exists yet | Stays behind, or must be drained | | Reading group start points | Nothing to decide | Chosen by hand, with reprocessing or a gap | | Old deployment | None | Runs and costs until proven unused | ## Why the coordination is the real bill The technical steps above are finite and can be scripted. What cannot be scripted is that **the clients belong to other teams**. Every producer and consumer moving to a new endpoint is a change in somebody else's repository, tested in somebody else's pipeline, released in somebody else's window, behind whatever else that team has committed to this quarter. A migration of thirty clients across nine teams is not a week of platform work; it is a quarter of chasing, and the last three clients take as long as the first twenty-seven. That is also why the old cluster lingers. It cannot be switched off on a date — it can only be switched off when nothing writes it, nothing reads it and, where the platform allows re-reading, nothing replays it. Until that is demonstrated rather than assumed, the estate is paying for both. ## What an interviewer wants to hear - That you say plainly a new cluster starts empty, and that this is why a late split is a migration rather than a purchase. - That you name the restart decision for readers, and the trade it forces, instead of assuming positions travel. - That you distinguish the platform's work from the clients' work, and identify the second as the schedule risk. - That you draw the practical conclusion: decide the cluster layout when the estate is small, and write the answer into the standard so that new streams land in the right place rather than being moved later.
- Where readers own a stored position, what are the two options for a reading group's start point on the new cluster, and what does each cost?Start at the oldest record present and reprocess everything the new cluster holds, which is safe only if consumers tolerate repeats; or start at the newest record and accept that anything written before that point on the new cluster is never read. There is no third option that recovers the old cluster's exact position.
- When can the old cluster actually be switched off?Only on evidence: no writer connecting, no reader consuming and, where the platform allows re-reading, no replay over an observation window long enough to contain the slowest periodic client. A date on a plan is not evidence, and a monthly job is the client that is usually missed.
- What makes the parallelism count worth getting right during the move rather than after it?On platforms that split a stream into parts, that count is fixed at creation and changing it afterwards is disruptive to ordering and to how work is assigned. Since the migration is already recreating every stream, it is the cheapest moment the estate will get to change it.
saying these in an interview costs you the question
- Assumes a new cluster starts with the old one's data
- Thinks stored read positions follow readers to another cluster
- Treats re-pointing clients as a platform-team task
- Plans to switch off the old cluster on a fixed date
- Forgets the grants have to be written again
- Recreates streams without revisiting settings fixed at creation