A cluster's default copy count was raised to three, yet a stream created earlier still keeps one copy. Why?
answer
- where does the number actually live?
- defaults are read at creation time
- per-stream value, not a live rule
- a new default protects only new streams
basics
~20 sBecause the copy count is a property of each stream, fixed when that stream was created from whatever default applied at the time. A cluster default is a template consulted at creation, not a rule re-applied to streams that already exist.
solid answer
~50 sThe number lives in two places that are easy to confuse. The **cluster default** is a template: it supplies a copy count to streams created from now on, including ones a client creates on first use. The **per-stream value** is what the stream actually keeps, recorded when it was created and unchanged since. Raising the default therefore protects future streams and no others - the old stream keeps its single copy, happily and silently, until someone changes that stream. And changing it is not a settings flip: a new copy is an empty node that has to take on the stream's existing data before it is any use, so the promised count is real only once that data has been placed. The practical consequence is that a durability standard needs an inventory of existing streams, not just a new default.
go deeper
Recall that the number belongs to the stream, not to the cluster: a default is used when a stream is created and never revisited afterwards. Changing the default changes nothing that already exists.
Explain both halves - creation-time evaluation, and the fact that adding a copy means moving real data onto another node - and name the usual sources of an under-copied stream, such as a stream created implicitly on first use.
Show that you would inventory existing streams rather than trusting the standard, triage them by what the records are worth, and fix the definition the stream is re-created from so the correction survives.
Treat it as governance: a standard that is only a default is a standard for streams created after the memo. Decide whether creation below the standard is refused outright, and who owns the exceptions.
## Two numbers, two moments There are two distinct things called the copy count, and interviews on this subject usually turn on telling them apart. - The **cluster default** - a template value consulted **at the moment a stream is created**, when no explicit number was supplied. - The **per-stream value** - the number this particular stream keeps, stored with the stream's own definition. This is a **per-stream override** when it was chosen deliberately, and a frozen echo of the old default when it was not. Only the second one has any effect on your data. The first one is an argument default, and like every argument default it is evaluated once, at the call. ## Why a default cannot reach backwards Two reasons, one conceptual and one physical. Conceptually, a stream's durability posture is meant to be a property of that stream. Streams in one cluster legitimately differ: a stream carrying payments and a stream carrying debug traffic should not be forced to the same number because they happen to share a cluster. If the default were re-applied continuously, a per-stream override could not exist. Physically, a copy count is not a flag - it is an instruction to hold bytes in more places. Raising it means an additional node must take on everything the stream already holds and keep up with what arrives while it does so. That is real traffic over real time. A cluster that silently started that work for every existing stream the moment someone edited a default would be a cluster that occasionally saturates itself on a Tuesday afternoon. The practical shape of it: | | Cluster default | Per-stream value | |---|---|---| | When it is read | at stream creation | continuously, by the cluster | | What changing it affects | streams created afterwards | this stream only | | Cost of changing it | none, it is a template | data must be placed on another node | | What it tells you about today's durability | nothing on its own | everything | ## Where under-copied streams come from Almost never from a decision. The recurring sources: - **The trial that became production.** The stream was created against a single-node cluster, where the only possible count was one, and the data outlived the experiment. - **The stream nobody created.** Where first use creates a stream implicitly, it takes the default in force on that day - which may predate the standard by a year. - **The deliberate override whose reason is lost.** Somebody lowered it once to make a load test cheaper. - **Tooling that supplies its own number.** A stream re-created by a script or an environment definition gets the script's number, not the cluster's, and that is also how a corrected stream quietly reverts. - **A stream copied from another environment**, where the standard was different or absent. ## Turning a standard into reality 1. **Inventory.** List every stream with the count it actually keeps. The cluster will not volunteer this; nothing is broken, so nothing is alerting. 2. **Triage.** Not every stream deserves the standard - some are disposable and regenerable, and saying so is part of the judgment. Decide per stream rather than sweeping. 3. **Change deliberately.** Plan each increase as a data movement with traffic and a period during which the number is recorded but not yet true. 4. **Fix the source.** Record the intended number wherever the stream is defined, so that re-creating it reproduces the decision instead of the old default. 5. **Close the door.** Where the platform allows it, prevent creation below the standard rather than auditing for it later - a rule at creation is the only fix that does not need repeating. ## What varies between platforms Do not assume the mechanism you know: - some platforms expose a per-stream count and a cluster default exactly as described; others offer **no per-stream dial at all** - durability is a property of the cluster or of a purchased tier; - some designs keep **a mirrored copy** - exactly one paired copy rather than a configurable N - so the question becomes whether mirroring is on, not how many; - on designs where **shared durable storage replaces per-node copies**, the redundancy belongs to the store underneath and a broker-side number means something different again; - a managed cluster may fix the number for you, in which case your per-stream decision is really a choice of tier. In all of them the question to ask is the same: **what number is this stream actually running with, and when was it decided?**
- How would you find the streams that are below the standard?List every stream with the count it actually keeps and compare it against the standard. Nothing will alert you: an under-copied stream serves traffic perfectly until the day a node is lost. Treat the result as an inventory to triage, because some streams are regenerable and some hold records nobody can reproduce.
- Why is raising a stream's copy count not effective the moment you ask for it?Because the new copy starts empty. It must take on the stream's existing data and catch up with new arrivals before it can stand in for anything, so the recorded number changes first and the durability it claims arrives later. Until then the stream is still running at its old depth.
- A stream is re-created every deployment from an environment definition. Where should its copy count be decided?In that definition. If the number lives only in a one-off administrative change, the next re-creation silently restores the cluster default, and the correction is lost with no trace. Writing the intended count where the stream is declared makes the decision reproducible and reviewable.
saying these in an interview costs you the question
- Expects a changed default to apply to existing streams.
- Thinks raising a stream's copy count is an instant flip.
- Never audits the counts of streams already running.
- Assumes every stream was created deliberately by a human.
- Believes a cluster restart re-reads defaults and repairs streams.