skip to content

In a multi-site LDAP directory whose replicas run minutes to weeks behind, how do you set a staleness window and decide which reads may not take it?

level: principalimportance: should knowfreq 30%

answer

  1. a number, not an aspiration
  2. classify reads by their use
  3. access decisions never take the window
  4. refer rather than answer stale
  5. silence and currency look identical

basics

~20 s

Publish a measured bound, not an aspiration: state how far behind a replica may be, measure it continuously, alarm when it is exceeded, and say what happens then. Then classify reads, and route every read that decides access or follows the reader's own write to the writable copy.

solid answer

~50 s

A replicated directory cannot be made synchronous over a link that is down, so the real decision is what you promise and to whom. Set a window per class of replica — seconds for a well-connected site, a voyage for a vessel — and make it measurable: compare a canary entry's `modifyTimestamp` on each replica against the writable copy, alarm on the gap, and state in writing what a consumer should do when the window is exceeded. Then classify the reads. Anything that decides whether an operation is permitted, and anything that must observe the reader's own recent write, goes to the writable copy or is answered with `referral (10)`; everything else may take the window. The uncomfortable part is naming the reads you will refuse to answer from a lagging replica, because that is where the cost lands.

code

pseudocode · 10 lines
pseudocode
function route_read(request, replicas, writable):
    # class 1 and 2 never take the window, whatever the lag is
    if request.decides_access or request.follows_own_write:
        return writable

    nearest = choose_nearest(replicas)
    if nearest.lag > published_window(nearest.band):
        return referral_to(writable)      # refer, do not answer stale

    return nearest

go deeper

for a junior

Recall that a copy of a directory can be behind the original, so a change you just made may not be visible everywhere at once.

for a middle

Explain that the lag is real and bounded only if it is measured, and that a read used for an access decision is different in kind from a read used to display a name.

for a senior

Demonstrate routing by read class, referring instead of answering stale, and the measurement that distinguishes a current replica from a silent one.

for a principal

Own the written contract: bands with numbers, the reads you refuse to serve from a replica, the behaviour on breach, and the honesty that a disconnected copy cannot be made current by any amount of design.

## A staleness window is a number you can be held to "Replication is usually a few seconds" is not a contract. A contract states a bound, the population it applies to, how it is measured, and what happens when it is breached. For a crew roster split across a shore office and vessels that reconnect when they dock, one number cannot cover both, so the contract is banded: - **Shore replicas** — behind by no more than a stated number of seconds, measured continuously, alarmed on breach. - **Vessel replicas** — behind by as long as the vessel has been offline, with the *last successful synchronization time* published alongside every answer they give. - **The writable copy** — current by definition, and the only place that claim is true. The second band is the one that teaches the discipline. A vessel's replica is not slow; for three weeks it is **absent**, and any design that quietly treats absence as latency will eventually answer a question from a copy that predates the fact. ## Classify the reads before you choose the mechanism The useful split is not by caller but by **what the answer is used for**. 1. **Reads that decide access.** Is this seafarer still a member of the group that may sign off a cargo manifest? Is the account still present? A removal that has not arrived is a grant you did not intend to keep, and here staleness is the failure, not a side effect of one. 2. **Reads that must see the reader's own write.** Someone changes a password or a contact detail ashore and immediately reads it back through a nearer replica. Answering from a lagging copy makes the system look broken even when nothing is wrong. 3. **Everything else.** A printed roster, a directory lookup, a report. These take the window happily and are most of the traffic. ## The mechanisms you actually have - **Route by class.** Send class 1 and class 2 reads to the writable copy. This is the primary control and the only one that is simple. - **Refer rather than answer.** A replica that knows it is outside its window can return `referral (10)` naming the writable server instead of a stale answer. An answer with a caveat nobody reads is worse than a redirect. - **Pin after a write.** Route a client's reads to the writable copy for a bounded period after it writes, which fixes class 2 without touching class 3. - **Chain the sensitive read.** The replica fetches it from the writable copy itself. This hides the hop, so it also hides the latency and the failure — acceptable for rare reads, dangerous as a default. - **Shorten the window.** Move a replica from a poll to a streamed session where the link supports it, so its window becomes replication latency rather than a polling interval. None of these makes a disconnected replica current. They decide **what a disconnected replica is allowed to say.** ## Measure it, or you do not have a window | What you measure | How | What it tells you | |---|---|---| | Lag per replica | a canary entry written on a schedule, its `modifyTimestamp` compared per replica | the actual distribution, not the design goal | | Last successful synchronization | when each replica last completed a session and stored a cookie | whether a quiet replica is current or simply silent | | Breach duration | time spent outside the published window | whether the number is honest or aspirational | | Full-refresh events | sessions that had to start with no cookie | that a cookie was refused, and the link paid for it | A replica reporting no lag because it has stopped synchronizing altogether is the failure this table exists to catch: silence and currency look identical from the outside. ## What you cannot promise, and should say so - You cannot promise that a revocation is effective everywhere at the instant it is written. You can promise that reads which depend on it do not go to a replica. - You cannot promise a window while the link is down. The contract must state the behaviour on breach — serve stale and say so, or refuse and refer — and that choice belongs to whoever owns the risk, not to whoever owns the servers. - You cannot promise that nobody will read around your routing. Publishing the window is how the rest of the estate learns the read they are about to make is not one the replica should answer. The deliverable is short: a table of bands with numbers, a list of read classes and where each is routed, the measurement that proves it, and one sentence on what happens when the number is missed. Its value is that the argument about a stale group membership happens now, in writing, and not during the incident.

  • Why is a membership check the read you route away from a replica first?
    Because its staleness is asymmetric. A membership that has not yet arrived merely delays a new permission; a membership that has been removed and not yet arrived keeps a permission alive that someone has decided to end. The second is the one with a cost, and only the writable copy is guaranteed to have seen the removal.
  • A replica reports zero lag. Why might that be the worst signal on the dashboard?
    Because a replica that has stopped synchronizing has nothing new to be behind on, and a naive lag measure can read as current. Pair lag with the time of the last completed session: a quiet replica and a current one look identical until you measure when it last finished one.
  • Why not simply chain every sensitive read from the nearest replica?
    Because chaining hides the hop. The read is correct, but its latency and its failure mode now belong to the writable copy while appearing to the client as the replica's behaviour, and a link problem shows up as an unexplained slowdown. It is a fine exception and a poor default.
  • How do you write the window for a replica that is offline for weeks by design?
    Not as a lag number at all. Publish the last successful synchronization time with every answer that replica gives, state which read classes it is permitted to answer at all, and require the others to reach the writable copy or fail. Absence is a different condition from lag and deserves different words.

saying these in an interview costs you the question

  • Promises eventual consistency without ever stating a number
  • Treats an offline replica as merely a slow one
  • Answers group membership checks from whichever replica is nearest
  • Assumes a lag metric of zero means the replica is current
  • Thinks adding replicas shortens the staleness window
  • Believes chaining a sensitive read removes the latency rather than hiding it