skip to content

Beyond simply returning older rows, what user-visible anomalies appear when an application spreads its reads over a pool of asynchronous replicas, and how do you prevent them?

level: middleimportance: should knowfreq 44%

answer

  1. pool = many different points in the past
  2. monotonic reads broken by round robin
  3. sticky replica by session hash
  4. message overtakes the change stream
  5. one unit of work per connection; eject on lag

basics

~20 s

Different replicas are at different positions, so a user can see time go backwards between two requests, see an effect without its cause, or get inconsistent halves of one page. Prevent it by pinning a session (or a page render) to one replica, and by ejecting replicas whose lag exceeds a threshold.

solid answer

~50 s

Each replica applies the change stream independently, so a pool of replicas is a set of different points in the past. Three anomalies follow. Non-monotonic reads: request one hits a caught-up replica and shows a new comment; request two lands on a laggier replica and the comment is gone. The user sees time run backwards, which is far more alarming than plain staleness. Broken causality: a job writes a parent row and enqueues a message; the consumer reads a replica that has the child but not the parent, or vice versa, and sees an effect without its cause. Dangling references appear here. Read skew across one render: a page issues several queries; different queries land on different replicas, and the totals do not match the rows. Prevention: make replica choice sticky per session (hash the session id) and per page render (reuse one connection), keep reads that must agree in one transaction, eject replicas past a lag threshold, and use the write position token for causality across services.

go deeper

for a junior

Know that different replicas are behind by different amounts, so a user can see a value appear and then disappear; sticking a user to one replica helps.

for a middle

Name monotonic reads, causality, and read skew across a render, and give the matching fix for each.

for a senior

Discuss lag fencing with hysteresis, applied versus received lag, unit-of-work pinning, and passing replication positions between services.

for a principal

Set the policy: which anomalies the product accepts, which read paths get a hard guarantee, and how that policy is enforced in routing infrastructure rather than per feature.

## Why a pool is worse than a single replica A single replica is simply behind by some amount. A pool of replicas is worse: each one applies the primary's change stream at its own pace, so at any instant they occupy different points in the primary's history. Load-balancing reads across them means consecutive requests from the same user can be answered from different moments in the past, in any order. ## Anomaly 1: reads that go backwards Monotonic reads is the property that a session never sees the data move backwards in time. Round-robin routing breaks it: the first request is served by a replica that is 20 ms behind and shows the item; the refresh is served by one that is 3 s behind and the item is missing. Users read this as the system lost my data, and support cannot reproduce it because their request lands elsewhere. Fix: make replica selection sticky. Choose the replica by a hash of the session or user id rather than round robin, so a given user keeps talking to the same copy of the past. If that replica is ejected for lag, the session moves forward in time (to a fresher node), which is harmless; the damage comes from moving backwards. ## Anomaly 2: an effect without its cause Causal anomalies appear when two pieces of related state are read from different places, or when a message travels faster than replication. Typical shapes: a worker receives a queue message referring to row 42, reads a replica, and finds no row 42, because the message overtook the change stream. Or service A writes an order and calls service B, which reads a replica and cannot see the order. Or a UI reads a comment whose parent post is not yet visible, producing a dangling reference even though the primary is perfectly consistent. Fix: carry the write's replication position (LSN or GTID) in the message or the request, and have the reader wait for or verify that position before reading; or read those specific paths from the primary; or make the message carry the payload instead of a pointer to a row. ## Anomaly 3: read skew across a single page A page that runs several queries can have them land on different replicas: the header count is computed from one snapshot, the list from another, and they disagree. Even on one replica, two separate statements see two different apply points, because the replica keeps applying between them. Fix: run the queries that must agree inside a single read-only transaction on one connection, which pins them to one snapshot; or fetch the numbers in one query; or accept the mismatch explicitly for cosmetic counters. ## Anomaly 4: mixing primary and replica in one flow When part of a request reads the primary and part reads a replica, you can observe a newer value and an older value simultaneously (for example a total from the primary against line items from a replica). Keep a logical unit of work on one node. ## Lag fencing All of the above get dramatically worse in the tail. A practical guard is to measure each replica's lag and remove any replica beyond a threshold from the read pool, sending its traffic to the primary or to healthier replicas. Two cautions: measure applied lag, not merely received bytes, since a replica can have the data but not yet have made it visible; and make ejection hysteretic so a replica does not flap in and out. ## Where the boundary is These are operational anomalies of one deployment shape (one writable primary, several read-only copies). The broader question of which consistency model a distributed system offers, and how to route reads at architecture level, belongs to system design; here the answer is concrete: sticky selection, unit-of-work pinning, position tokens, and lag fencing. ## How to present it Do not answer merely the data is a bit old. Name the three concrete shapes (backwards in time, effect without cause, halves of a page disagreeing), tie each to a routing decision that caused it, and give the specific counter-measure for each.

  • Why does hashing the session to a replica help, and what does it not fix?
    It keeps one user reading one copy of history, so their view only ever moves forward and non-monotonic reads disappear. It does not make that user's reads fresh, it does not help two different sessions or services that must agree, and it can skew load if a few sessions are very heavy.
  • A worker consumes a message about a row it cannot find on the replica. What is the fix?
    The message overtook replication. Either include the write's replication position in the message and have the consumer wait until its replica has applied that position, read that lookup from the primary, or put the needed data in the message payload. Blind retries with backoff are a weaker fallback that hides the ordering problem.

Reading from a replica pool is like asking three friends who each left the meeting at different times; ask them in the wrong order and the story goes backwards.

saying these in an interview costs you the question

  • Treating staleness as only slightly old data, missing that reads can go backwards
  • Round-robin routing across replicas for user-facing sessions with no stickiness
  • Running the queries of one page render on different connections and expecting consistent totals
  • Measuring lag as bytes received rather than changes applied and visible
  • Assuming a replica is internally inconsistent, when it is actually consistent but simply at an earlier point in history

context