How would you define a small set of estate-wide durability tiers so teams pick a write's acknowledgement rule without re-deciding it per stream?
answer
- a menu, not a setting
- the settings only work together
- two or three tiers, no more
- the default tier is most of the estate
- enforce at creation, report drift
basics
~20 sPublish two or three named tiers, each fixing a copy count, a floor on caught-up copies, an acknowledgement rule and a latency budget, then enforce the choice at stream creation so an unreviewed stream lands on a safe default rather than on nothing.
solid answer
~40 sThe goal is to turn a per-stream argument into a per-stream choice from a short menu. Define two or three tiers — something like a critical tier, a standard tier and a disposable tier — and for each one fix the copy count, the floor on caught-up copies, what the writer waits for, the placement constraint across failure domains, and the write-latency and stored-bytes envelope that combination implies. Say explicitly which tier a stream gets when nobody chooses, because most streams are created once and never reviewed. Then enforce it where streams are created rather than in a document, and make the current tier of every stream reportable so drift is visible. The hard part is not picking values; it is keeping the menu short enough that teams actually choose from it.
go deeper
Know that durability settings travel together as a bundle — copies, the floor, what the writer waits for — rather than being tuned one at a time.
Be able to describe what one tier contains and why the settings inside it have to agree, using the write-refusal failure as the example of what disagreement produces.
Bring the operational half: where the tier is enforced, how you detect a stream that has drifted from the one it claims, and how an emergency deviation is time-boxed and restored.
Own the trade as an estate decision — how many tiers, what the default is, what each costs in write latency and stored bytes, and what the standard honestly says on platforms that do not expose these settings.
## Why a menu rather than a setting Every stream in an estate has a durability posture whether anyone decided it or not. Left to defaults, it inherits whatever the cluster was built with; left to teams, it inherits whatever was copied from the last stream someone created. Neither is a decision. A **durability tier** is a named bundle of the settings that must agree with each other, so that a team makes one choice about the *data* and the platform derives the rest. The settings only make sense together, which is the whole argument for bundling them: - **Copy count** — how many copies of the stream the cluster keeps. - **The floor on caught-up copies required for a write** — how many must be current before the server will accept a write at all. - **The acknowledgement rule** — what the writer waits for: no wait, the leader alone, a majority of copies, every caught-up copy. - **Placement** — which failure domains the copies must be separated across. - **The policy on forcing bytes to persistent media**, which is a distinct question from acceptance and belongs in the bundle for the same reason. Choosing any one of these without the others is how an estate ends up with a floor that is never consulted, or a floor equal to the copy count that stops writing the first time a machine reboots. ## A three-tier shape | Tier | Posture | What the writer waits for | Tolerates | Paid in | |---|---|---|---|---| | Critical | Highest copy count, floor one below it, copies separated across failure domains | Every caught-up copy | One copy out | Write latency set by the slowest copy and by inter-domain distance; the largest stored-byte multiplier | | Standard | Moderate copy count, floor one below it | A majority, or every caught-up copy where the design offers no majority | One copy out | A round trip beyond local acceptance | | Disposable | Low copy count, no meaningful floor | The leader alone, or no wait for pure telemetry | Nothing, by design | Accepted loss on failure | Three is usually the right number. Two collapses into "important or not" and pushes every awkward case into the wrong half; five or more restores exactly the per-stream argument the menu was meant to end. ## What makes the standard hold 1. **Name the default explicitly.** Most streams are created once, by someone shipping something else, and are never revisited. Whatever the default tier is, that is the posture of the long tail of your estate — so choose it deliberately rather than inheriting a build-time value. 2. **Enforce it where streams are created.** A standard applied only in review is applied only to streams that get reviewed. The enforcement point is whatever creates streams in your estate, and the check is that the tier is set and internally consistent. 3. **Make the current posture reportable.** You want to answer "which streams are on which tier, and which no longer match the tier they claim" without a person reading configuration by hand. Drift is silent otherwise. 4. **Verify from the writing side too.** The tier promises a rung the writer must actually request. A tier that describes only server-side values will be half-applied, and half-applied looks fully applied in an audit. 5. **Write down the emergency deviation.** Lowering a floor to restore writes during an incident is legitimate; doing it without a time box and a restoration step is how a temporary posture becomes permanent. ## What varies by platform, and what the tier must not assume The tier is a policy object, so it must be expressed in terms your platforms actually have. - Where a broker keeps **a single mirrored copy** rather than a configurable number, the menu has two rungs, not four, and the critical tier means "answered once the mirrored copy also holds it". - Where durability comes from **shared durable storage underneath the brokers**, copy count is not yours to set, and the tier reduces to what the underlying store confirms before the broker answers plus whatever the provider commits to. - Where records are **deleted once acknowledged by a consumer** rather than retained, the acknowledgement rule still applies to the write, and the cost side of the tier is expressed in undelivered stored bytes rather than in a retention window. - On a **rented broker**, several of these settings may not be exposed at all, and the honest tier records what the provider commits to instead of pretending to a setting you cannot reach. ## The judgment being tested An interviewer asking this wants to see you treat durability as an organisational standard with an enforcement point and a drift story, not as a set of values to recite. The strongest answers say out loud what the tiers *cost* — in write latency and in stored bytes — and name the default tier, because that is the one most of the estate will end up on.
- Why is the default tier the most important one in the standard?Because most streams are created in passing by someone delivering a feature, and are never reviewed again. The default is therefore the real posture of the long tail of the estate, no matter what the document says the tiers are. Choosing it deliberately — and usually conservatively — does more for durability than perfecting the top tier.
- How do you keep tiers honest on a rented broker where some settings are not exposed?Express the tier as what the provider commits to rather than as settings you cannot reach, and record which parts of the posture you no longer control. A tier that lists values the platform does not expose reads as compliance while guaranteeing nothing, which is worse than an honest tier that names the gap.
- What is the failure mode of having five or six tiers instead of three?The menu stops functioning as a menu. Teams re-open the trade-off each time, tiers accumulate near-duplicates, and the platform has to support every combination during upgrades and capacity work. A small number of coarse tiers that some streams slightly over-pay for is cheaper than a catalogue nobody can hold in their head.
Shipping options at a counter: economy, tracked and signed-for. Nobody re-derives the trade for each parcel; they pick a named service whose speed, cost and proof of delivery were decided once.
saying these in an interview costs you the question
- Publishes the standard as a document with no enforcement point
- Defines a tier by copy count alone, leaving the floor and the rule loose
- Never names which tier a stream gets when nobody chooses
- Grows the menu until every stream has a bespoke posture
- States tiers only as server settings, never as what writers must request
- Omits the latency and stored-bytes cost each tier implies