What does a replication-lag figure for an in-memory store actually measure, and how can it look healthy while a copy is useless?
answer
- distance, not consequence
- position difference versus seconds behind
- zero on an idle tier means nothing
- a disconnected copy leaves the figure
- alert on the count of connected copies
basics
~20 sReplication lag measures how far a connected copy trails the primary's write stream, in stream position or in seconds. A copy that disconnects drops out of the figure entirely, so the signal goes blind rather than red.
solid answer
~50 sThe figure is computed by the primary about the copies currently attached to it, or by a copy about itself, and it is usually expressed one of two ways. **A position difference in the write stream** tells you how much data is outstanding but cannot be turned into seconds without knowing the current write rate. **A time difference** is readable directly but reads as zero on an idle tier whether or not the link works, because nothing is outstanding when nothing is being written. Both share the same blind spot: a copy that disconnects stops being reported at all, so the lag figure for the remaining copies can look perfect while the deployment is down to one. That is why the lag number must always be alerted on alongside **the count of connected copies** - a drop in that count is the failure that lag alone will never show you.
go deeper
Remember that replication lag describes only the copies currently connected, and that a copy which has dropped out disappears from the figure rather than showing up as badly lagged.
Explain the two frames the number comes in - a position in the write stream and an elapsed time - and why neither converts into the other without knowing the current write rate.
Demonstrate that you alert on the connected-copy count first and the lag second, over a sustained window, and that you treat a copy's own report disagreeing with the primary's as a finding in itself.
Decide up front what the copies are for - availability, read capacity, or a promotion you can afford - because that decision, not the lag graph, determines what value is tolerable and what the number will mean on the day it matters.
## What the number is computed from Replication lag is the distance between what the primary has accepted and what a copy has received. Two frames are in common use and they are not interchangeable: | Frame | What it says | What it cannot say | |---|---|---| | **Position in the write stream** (a byte or sequence difference) | exactly how much data is outstanding for this copy | how long that will take to close, or how old the copy's data is - both need the current write rate, which is not in the figure | | **Elapsed time** (seconds behind) | how stale the copy's view is right now | anything useful on an idle tier, where nothing outstanding means zero seconds whether the link is alive or dead | A useful habit: quote lag **in the unit it is measured in**, and say which vantage point produced it. A copy's own report and the primary's report about that copy are two different measurements, and they diverge exactly when the connection between them is the problem. ## The three ways this signal lies 1. **It goes blind, not red.** A copy that has disconnected entirely usually stops appearing in the primary's list of copies. Its lag is not reported as enormous - it is not reported at all, and a dashboard that plots the maximum lag across copies will happily show a healthy line computed over the one copy that is still there. This is the single most important property of the signal. 2. **Zero lag on an idle tier means nothing.** Nothing outstanding is not evidence of a working link. A tier that takes writes in bursts can report a perfect number for the entire quiet period between them. 3. **Received is not necessarily applied and servable.** Depending on the store, the figure may track how much of the stream has arrived at the copy rather than how much it has processed and can now answer reads from. Where the two differ, the copy can be simultaneously 'caught up' by the number and behind by what a reader observes. ## Where stores in this class differ - **When the caller's write is acknowledged varies.** Some stores acknowledge as soon as the primary holds the write and stream it onward afterwards; some acknowledge only once a copy also holds it; some let the caller choose per call. The lag figure means a different thing under each: on the first it measures a window of data you would lose, on the second it measures a delay you are already paying on every write. Read the number without knowing which posture you are on and you will misprice a promotion. - **What counts as a copy varies.** Some deployments chain copies behind other copies, in which case the primary's figure describes only the first hop and the lag of the furthest copy is not in it at all. - **Some managed deployments publish only a derived health indicator** rather than the raw distance, and derived indicators usually smooth exactly the short spikes that matter when a promotion is about to happen. ## Alerting on the pair, not on the number - **Alert on the count of connected copies first.** It falling below the intended number is an unambiguous, immediately actionable event, and it is the failure that lag cannot represent. - **Alert on lag second**, as prose arithmetic against a sustained window - for example, a copy more than a few seconds behind, sustained over several minutes, rather than on a single spike, since a brief burst of writes will produce spikes that close by themselves. - **Alert separately on a copy's own report disagreeing with the primary's.** That disagreement is the shape of a link that is up in one direction only. - **Distinguish the dashboard from the pager.** A lag number that rises and falls with write bursts belongs on a dashboard. A copy count below target, or lag that has not closed within the window, belongs on a pager. ## What this number does not license you to conclude The lag figure tells you the **distance** and nothing about the **consequences**. It does not tell you whether an acknowledged write would survive a promotion - that depends on the acknowledgement posture above. It does not tell you whether reads served from the copy are acceptable to the caller, which is a question about the workload rather than the tier. And on a tier whose copies exist for availability rather than for read capacity, a few seconds of lag may be entirely uninteresting right up until the moment a promotion happens, at which point the same number retrospectively describes what was lost. Reading it well means knowing, before the incident, which of those questions your deployment is actually asking.
- Why is a dashboard plotting the maximum lag across all copies an unsafe primary alert?Because the maximum is taken over the copies still being reported. A copy that disconnects leaves the set entirely, so the maximum is recomputed over the survivors and can drop at the exact moment the deployment gets less resilient. The count of connected copies is the signal that moves the right way in that event.
- A copy reports itself as caught up while the primary reports it several seconds behind. What does that disagreement suggest?That the two vantage points are measuring different things, or that the link is healthy in one direction only - the copy has processed everything it received, while the primary still has data it has not managed to hand over. Trust neither figure alone; treat the disagreement itself as the finding and check whether the copy can actually serve current data.
- Why can a lag figure of zero on a low-traffic tier be worthless as reassurance?Because with nothing outstanding there is nothing for the figure to measure, so a dead link and a perfect one report the same value. Reassurance has to come from a signal that moves without traffic - the count of connected copies, or a periodic write whose arrival at the copy is observed - rather than from the lag number itself.
saying these in an interview costs you the question
- Believes a disconnected copy shows up as a very large lag value
- Reads zero lag on an idle tier as proof the link is working
- Quotes lag without saying whether it is a stream position or a time
- Assumes every store in this class acknowledges writes before a copy holds them
- Assumes data received by a copy is already servable from it
- Alerts on the maximum lag across copies as the only replication signal