A tier refuses writes when fewer than a set number of replicas are current: what does that gate buy, and what does it not?
answer
- refuse writes when copies are behind
- loss becomes refusal, not safety
- bounded blindness, made visible
- narrows, never closes
- the threshold is the exposure you accept
basics
~10 sA minimum-healthy-copies write gate converts a silent loss risk into a loud refusal: writes are accepted only while enough copies are current. It narrows the un-propagated window and does not close it.
solid answer
~50 sThe gate is a rule on the primary: accept writes only while at least some number of copies are connected and within a lag threshold, and refuse them otherwise. What it buys is a bound on how blind you can get - you cannot keep accepting writes for ten minutes while no copy has heard anything. What it does not buy is a closed window: writes accepted in the moments before the gate trips are still un-propagated, everything inside the gate's own lag threshold is still exposed, and the gate's view of copy health is itself slightly stale. The real trade is not loss-versus-safety but **loss-versus-refusal**: the tier starts failing writes, which for reconstructable state is worse than losing a few and for claims and leases is better. Not every store in this class offers such a gate.
go deeper
Know that some stores can be told to stop accepting writes when their copies fall behind. It is a way of failing loudly instead of quietly losing writes nobody copied.
Explain the mechanism and state its limit in the same breath: it narrows the loss window and does not close it, because writes taken before it trips and writes inside its own lag threshold are still exposed.
Argue the exchange explicitly. The gate turns a silent loss risk into a write-path availability dependency, and whether that is an improvement depends entirely on whether the affected entries can be rebuilt from somewhere else.
Decide whether the organisation wants this tier to be able to refuse writes at all, and make sure callers are written for it. A gate adopted in the belief that it delivers zero loss creates a false sense of safety that costs more than it saves.
## What the gate is A **minimum-healthy-copies write gate** is a condition the primary evaluates before accepting a write: are at least *k* copies currently connected and within some lag bound? If yes, accept. If no, refuse. Stores that offer it express the condition as a count of current copies, as a lag threshold, or as both together, and several stores in this class offer no such mechanism at all - so the first honest answer to "do we have a gate" is "check, don't assume". It is worth being clear about what the gate is *not*. It is not a waiting acknowledgment. Under the gate the primary still answers the caller before any copy holds the write. The gate does not change the ordering of the acknowledgment; it changes whether the write is accepted at all. ## What it buys The failure it exists to prevent is the slow, silent one. A copy disconnects. Nothing is watching. The primary keeps accepting writes, happily, for minutes - and the whole of that period is un-propagated. Then the primary dies and the window is not a few milliseconds of writes, it is everything since the copy dropped. The gate bounds that. It puts a ceiling on how long the tier can keep taking writes while blind, and it makes the blindness **visible**, because callers start getting errors instead of silence. That second property is most of the value: a tier that refuses writes gets paged on in minutes; a tier that accepts writes nobody is copying gets discovered during the incident. ## Why it narrows and does not close 1. **Writes accepted just before the gate trips are still un-propagated.** The gate reacts to a condition; it cannot un-accept what it already accepted. 2. **The gate's own lag threshold is the floor.** If the rule permits writes while copies are within 10 seconds of current, then up to 10 seconds of writes are exposed by design. Tightening the threshold narrows this and makes spurious refusals more likely. 3. **The health view is itself delayed.** Whatever mechanism decides that a copy is "current" learns of a change after it happens, so there is always an interval in which the gate believes a stale picture. 4. **It says nothing about the copy that gets promoted.** A gate satisfied by copy A does not put the write on copy B. So the correct phrasing, and the one an interviewer is listening for, is that the gate **narrows the loss window without closing it**. Anyone who describes it as preventing data loss has misread what it does. ## The trade it actually makes The gate does not remove a risk; it **exchanges** one for another. | Without the gate | With the gate | |---|---| | Writes always accepted | Writes refused whenever copy health drops below the rule | | Loss on promotion can be unbounded in time | Loss bounded by the threshold, minus the caveats above | | Failure is silent until the promotion | Failure is loud, immediate, and lands on callers | | Availability of the write path is maximal | Availability of the write path now depends on copy health | That table is the decision. For entries that are reconstructable from a system of record, refusing writes is usually the worse outcome: you have taken a component whose job is to absorb load and made it an availability dependency in order to protect state you could have rebuilt. For claims, leases, deduplication records and quota counters with no other home, refusing is usually the better outcome: a caller that gets an error can retry or fall back, whereas a caller that gets a success for a write that later evaporates has no recourse and does not know it needs one. Either way, the caller has to be written for it. A gate that refuses writes into an application that treats every write as infallible converts a data-loss incident into an outage plus a data-loss incident. ## Choosing the numbers - The **count** should reflect how many copies you actually run. Requiring more current copies than you deploy is a gate that is permanently closed, and it has been configured that way by accident more than once. - The **lag threshold** is the exposure you are explicitly accepting. Set it from the sizing work - lag times write rate - rather than from a round number. - **Test it.** Disconnect a copy under load and confirm that writes are refused, that the refusal is visible in your monitoring, and that the application does something sensible with the error rather than logging it and continuing. The gate is a good mechanism used badly more often than it is used well, and the usual cause is a team that adopted it expecting zero loss rather than bounded loss.
- Does the gate change when the caller is told yes?No. Under the gate the primary still answers before any copy holds the write; only whether the write is accepted at all changes. That is what separates it from wait-for-a-copy acknowledgment, which changes the ordering and adds a hop to every write.
- Why can a gate's threshold be tightened only so far?Because the gate's picture of copy health is delayed and lag fluctuates normally. A threshold close to the ordinary lag turns routine variation into refused writes, so tightening trades a narrower window for spurious write-path outages.
- What breaks if you enable the gate without touching the application?Callers start receiving write errors they were never written to handle. A path that assumed writes always succeed will either lose the work silently or fail in a way unrelated to the actual cause, turning a bounded-loss mechanism into an outage of its own.
saying these in an interview costs you the question
- Says the gate prevents data loss rather than bounding it.
- Thinks the gate changes when the caller is told yes.
- Sets the required copy count above the number of copies deployed.
- Ignores that refused writes must be handled by the caller.
- Assumes every store in this class offers such a gate.
- Picks the lag threshold as a round number rather than from measurement.