skip to content

Who decides whether a cluster may promote a copy that is behind the lost leader, and what should that policy depend on?

level: principalimportance: nice to knowfreq 34%

answer

  1. the loud cost beats the quiet one
  2. decided in advance, not live
  3. per stream, not per cluster
  4. the data owner signs it
  5. replayable upstream changes the answer

basics

~20 s

The accountable owner of each stream decides, in advance and per stream, not the operator during the incident. The policy should turn on what the stream carries, whether it can be republished from upstream, and what an outage costs relative to silent loss.

solid answer

~50 s

Made live, this decision is always the same: service is visibly down, the loss is invisible, and the fast option wins. So it has to be pre-declared per stream by whoever is accountable for that data, and encoded where the cluster actually enforces it rather than left as a sentence in a runbook. The inputs are what the stream carries and who consumes it, whether the records can be republished from the system that produced them, what an outage of an hour costs against a silent hole, and any obligation that makes unexplained gaps unacceptable. A sensible default is to forbid it with a named escalation path, and to require that the extent of any loss is recorded at the moment. If the estate keeps facing the choice, the real answer is a design change, not a better decision.

go deeper

for a junior

Know that restoring service fast can mean discarding confirmed records, and that this is not a choice an individual makes alone during an incident.

for a middle

Explain why the decision has to be made before the incident: the outage is visible and the loss is not, so the fast option always looks better under pressure.

for a senior

Show that the policy is per stream, that you would establish whether the producer can republish, and that you would require the extent of any loss to be recorded and sent downstream.

for a principal

The angle is ownership and enforcement: name who signs per stream, put the rule where the cluster obeys it, and treat a recurring need for the choice as a signal to change placement, storage design or producer replayability.

## Why this is a policy and not a decision At the moment of the incident the two costs are not comparable. Unavailability is loud, measured, visible on every dashboard and attributed to the people in the room. Losing records that were already acknowledged is silent — no error is raised, no producer is told, and the discrepancy surfaces weeks later in a reconciliation nobody connects to that night. A structure that presents a loud cost and a quiet cost to a tired operator under pressure has already chosen. So the point of the policy is not to make a wiser choice in the moment; it is to have made the choice on a calm afternoon, with the data owner present, and to have put it where the cluster enforces it. ## What the policy should turn on 1. **What the stream carries.** A stream that is the system of record for money, orders or regulated events is not the same decision as one carrying metrics or cache invalidations, even in the same cluster. 2. **Whether it can be republished.** If the producing system retains its own copy and can re-emit the interval, the loss becomes a re-publish task instead of a permanent hole. This is the single most decision-changing property and the one teams least often know the answer to. 3. **What downstream does with it.** Consumers that recompute from a durable source can absorb a gap; consumers that accumulate irreversible state from each record cannot. 4. **The cost of an hour down against the cost of an unexplained gap.** Both must be estimated, or the loud one wins by default. 5. **Obligations that make gaps unacceptable.** Where the organisation must be able to state what it holds and why, a silent shortening of a stream is not available as an option at any speed. 6. **Whether the loss can be measured.** A policy that permits the fast option without requiring the extent to be recorded produces an unbounded investigation later. ## Who signs it The accountable owner of the data, not the platform team and not the on-call operator. The platform team owns the mechanism, the constraint and the evidence; the data owner owns the consequence. A useful test of whether the policy is real: can you name, per stream, the person who would be asked afterwards why those records are gone? If not, what exists is a preference rather than a policy. ## Making it operable - **Encode it where it is enforced.** A per-stream setting the cluster obeys beats a paragraph in a document, because the paragraph is not present at three in the morning. - **Default to forbidding it**, with a named escalation that can lift the restriction deliberately for a specific stream during a specific incident. - **Require the extent to be captured** at the moment: the promoted copy's end against the last acknowledged position of the lost leader, and the interval it covers. - **Require notification downstream**, since the consumers are the only parties who can reconcile. - **Review it when the stream's value changes**, which is usually when a second consumer starts treating it as authoritative without telling anyone. ## When the answer is to change the design If this choice keeps arriving, the policy is treating a symptom. The structural fixes sit outside this decision: - spreading copies so that a single failure domain cannot take the whole caught-up set; - keeping catch-up health visible so copies do not quietly drift out of the caught-up set and only get noticed during a failure; - moving to **detached storage**, where a unit's records live on shared or remote storage, so a replacement node sees the same records and the trade largely disappears; - ensuring the producing systems can replay, which converts the worst outcome from permanent loss into a re-publish. A principal-level answer says plainly that the last of these is usually the cheapest and least glamorous, and that it is bought in the producing services rather than in the cluster. ## Where designs differ Some platforms expose this as a per-stream setting, some only cluster-wide, and some forbid the behaviour entirely so the policy question never reaches the operator. On hosted offerings the provider may have decided already and may not say which way, which makes 'find out what our provider does when the current copies are unreachable' a genuine piece of due diligence rather than a formality. Where the choice is not exposed, the policy shifts entirely onto the producing side: replayability is then the only lever the organisation holds.

  • Your platform exposes this choice only cluster-wide, but two streams on it need opposite answers. What do you do?
    Stop sharing the cluster for those two streams. If the enforcement granularity is coarser than the policy, the estate layout is the control: put the stream that cannot tolerate silent loss where the restriction holds, and the tolerant one elsewhere. The alternative — one setting and a promise to behave differently per stream — is not enforceable during an incident.
  • What evidence would convince you to permit it for a stream that currently forbids it?
    That every consumer recomputes from a durable source or can reconcile a stated gap, that the producer can republish the interval, and that the extent of any loss will be recorded and published. Absent the second, permitting it means accepting a permanent hole, which is a decision the data owner must make explicitly rather than inherit.
  • How do you keep this policy from going stale?
    Tie the review to the stream's consumer list rather than the calendar. The value of a stream changes when a new consumer starts treating it as authoritative, which normally happens without anyone revisiting the policy. Reviewing on consumer change, and at the same time as the stream's ownership record, catches it far earlier than an annual audit.

saying these in an interview costs you the question

  • Leaves the call to whoever happens to be on call
  • Applies one cluster-wide rule to streams of very different value
  • Calls the incident closed because service returned
  • Keeps the policy in a runbook rather than where the cluster enforces it
  • Never establishes whether the producer could republish the interval
  • Permits the fast option without requiring the loss extent to be recorded