A tier acknowledges writes before a replica holds them: how do you put a defensible number on what a promotion would vaporise?
answer
- two numbers, not one
- tail lag, not median lag
- lag times write rate
- then say what the writes held
- test it with a real promotion
basics
~20 sMultiply the tail propagation lag at peak write rate by that write rate to get a count of writes at risk, then say what those writes were. The count alone is not an answer; the state they held is.
solid answer
~50 sMeasure **propagation lag** two ways, because stores report it differently: as time behind the primary, and as outstanding un-propagated work. Take the tail of that lag under peak load rather than the median - the promotion you are sizing for happens on the bad day, not the median day. Multiply by the write rate at that moment and you have the count of acknowledged writes a promotion would vaporise. Then do the half people skip: name what those writes held. Five thousand lost counter increments is a slow afternoon; five thousand lost deduplication records is five thousand requests processed twice. The number also has to be defended against what widens the lag - a write burst, a saturated link, a busy copy, a long-running operation occupying the primary, or a copy that has fallen far enough behind to need a whole fresh copy of the keyspace.
go deeper
Know that the amount at risk is not a fixed property of the system: it depends on how far behind the copy is and how fast writes are arriving, and both of those change minute by minute.
Be able to turn a lag measurement into a count of writes, and to name at least three conditions that widen the lag. Know which form of lag your store reports and which side reports it.
Use the tail under peak load, not the median, and explain why the failure and the widened window tend to arrive together. Then convert the count into a consequence by naming what the affected entries hold.
Treat the number as something each workload owns and periodically proves, not something the platform team asserts once. An untested estimate of what a promotion vaporises is an assumption the business has not actually agreed to.
## Two numbers, not one Asked "how much would we lose", most engineers answer with a duration - "about a second". That is half a number. A defensible answer has two parts: 1. **The magnitude**: how many acknowledged writes no copy held at the worst plausible instant. 2. **The content**: what those writes were, and whether any other system can reproduce them. The first is arithmetic. The second is what turns the arithmetic into a decision, and it is the part that distinguishes someone who has actually run this tier. ## Measuring the lag honestly Propagation lag is exposed differently across stores in this class, and each form has a blind spot: - **Time behind the primary.** Readable and intuitive, but on an idle tier it reads near zero regardless of health, because there is nothing to be behind on. It tells you least exactly when you are calm. - **Outstanding un-propagated work**, as a count of operations or a volume of bytes not yet confirmed. This is closer to what you want, because it is the window itself, but it needs the write rate to become a duration. - **Which side reports it** matters. A number computed on the primary reflects what the primary has sent. A copy that received changes but is slow to apply them can look healthy from the primary's side while holding less than the primary thinks. Where both views are available, take the pessimistic one. Whatever form you get, sample it continuously and keep the distribution. A single spot reading during a calm hour is not evidence about the bad hour. ## From lag to a count of writes The estimate is deliberately crude: ``` writes at risk ~= tail propagation lag (at peak) x write rate (at peak) ``` Use the tail, not the median: 0.4ms of typical lag and 180ms of tail lag are the same deployment, and it is the tail that coincides with the failure you are sizing for - a node under memory pressure, a saturated link, or a burst is exactly the condition that both widens the window and makes a node die. ## What widens it - **A write burst.** Arrivals exceed what propagation can carry, and the backlog grows for as long as the burst lasts. - **A saturated or shared link** between the nodes, including one shared with backup traffic or with another tenant. - **A busy copy.** The lag includes the copy's own conditions - its load, a pause, its own housekeeping. Propagation is not free at the receiving end. - **A long-running single operation on the primary**, which occupies it while writes queue behind it and nothing is sent on. - **A copy that disconnected.** Many stores hold a bounded buffer of recent changes so a copy that drops for a moment can resume from where it stopped. When that buffer overflows, the copy must take a whole fresh copy of the keyspace instead - and for the duration of that transfer it holds nothing recent, so the effective window is the whole transfer. Not every store in this class works this way, and where it does the buffer's size is a capacity decision on the primary, not a free feature. ## The half that decides whether the number matters | What the entries held | What losing the window actually is | |---|---| | Values reconstructable from a system of record | Extra load and a slow period while they are rebuilt | | Advisory counters and rate figures | A small inaccuracy nobody can distinguish from noise | | Deduplication records | Requests processed a second time, with whatever that means downstream | | Claims and leases | Two holders of something meant to have one | | Quota and limit counters | Allowances reset, or consumed twice | The same mechanism and the same count produce an annoyance in the first row and an incident in the last. So the number you defend is not "we would lose 1,200 writes"; it is "we would lose up to 1,200 writes, of which the deduplication records are the ones that matter, and here is what a second delivery of those requests does". ## Defending it A number nobody has tested is a guess. The way to make it defensible is to have caused a promotion deliberately, under representative load, and compared what the new primary held against what callers were told. That exercise also tends to surface the thing the arithmetic misses: whether the write rate at the moment of failure is the steady-state rate at all, since the event that kills a node often arrives together with the burst that overwhelmed it. Finally, keep the frame right. This is a count of writes that were **acknowledged and then ceased to exist** - not a gap to be filled in later from anywhere. That is the property of this tier that makes the number worth computing at all.
- Why is a lag reading of zero on an idle tier not evidence of health?Time-behind measures how far a copy trails the stream of changes. With no changes arriving there is nothing to trail, so the reading is near zero regardless of whether propagation would keep up under load. Size the window from measurements taken at peak, not at rest.
- Your estimate says 1,200 writes. What would make that number an underestimate?A write rate at the moment of failure higher than the steady-state rate used in the arithmetic, which is common because the burst and the failure often share a cause. Also a copy mid-resynchronisation, which holds nothing recent, and a lag number reported from the primary's side rather than the copy's.
- Is the window smaller if you run the replica on the same machine as the primary?The lag is, but the number stops meaning anything. The window is sized against a failure that takes the primary; if the copy shares the machine, the same failure takes both and the window becomes everything either of them held.
saying these in an interview costs you the question
- Quotes a median lag figure as the size of the window.
- Gives a count of lost writes without saying what they held.
- Reads a near-zero lag on an idle tier as proof of health.
- Assumes the write rate at failure equals the steady-state rate.
- Trusts a primary-side lag number as what a copy actually holds.
- Has never caused a promotion under load to check the estimate.