A network partition leaves a Redis master reachable by some application clients but cut off from all Redis Sentinel processes, so Sentinel promotes a replica on the other side. What happens to the writes those clients keep sending to the old master, what actually bounds how long that can go on, and which Sentinel-side settings or hooks (down-after-milliseconds, client-reconfig-script, replica-priority) let an operator observe or shorten that window?
answer
- Sentinel is an observer — it cannot fence the old master
- Old master never told; keeps returning +OK on the minority side
- Window bounded by client re-resolution, not any server timeout
- Cluster self-demotes at cluster-node-timeout; Sentinel has no equivalent
- Rejoin → full resync → divergent writes gone, no error
basics
~20 sSentinel never tells the isolated master anything, so it keeps accepting and acking writes from clients that still reach it. Only client re-resolution ends that window, not a server timeout. On rejoin it full-resyncs and those writes are discarded silently.
solid answer
~50 s**Nothing fences the old master.** Sentinel is an external supervisor: it agrees a master is down and promotes a replica, but it has no channel to tell the isolated node to stop — and that channel is exactly what the partition broke. The old node still reports `role:master` and still returns `+OK`. **What bounds the window is the client, not the server.** It ends when clients stop talking to the old master: `+switch-master` pub/sub propagation, connection teardown, DNS/VIP/proxy lag. A client with a hardcoded master address never leaves. Contrast Redis Cluster, where a partitioned primary that cannot reach a majority self-demotes after `cluster-node-timeout`. **The writes die silently.** On heal, Sentinel issues `REPLICAOF` to the old master; its diverged dataset is replaced by a full resync — no merge, no conflict log. **Levers:** Sentinel-aware clients on `+switch-master`, `client-reconfig-script` to flip a VIP/proxy for clients that aren't, `down-after-milliseconds` for detection, `replica-priority` to steer promotion. Server-side write fencing is covered separately.
code
text · 6 linessentinel monitor mymaster 10.0.1.10 6379 2
# how long unresponsive before this Sentinel calls it SDOWN (detection clock)
sentinel down-after-milliseconds mymaster 5000
# runs on the Sentinel leader at failover, with old and new master addresses;
# use it to move a VIP / reload a proxy for clients that don't speak Sentinel
sentinel client-reconfig-script mymaster /opt/redis/move-vip.shgo deeper
Know the core fact: Sentinel promotes a replica, but the old master is never told, so for a while two nodes think they are master and one side's writes will be thrown away.
Explain the mechanism — Sentinel is an external observer with no fencing channel, clients on the minority side keep getting +OK, and on rejoin the old master full-resyncs and loses its diverged data. Name +switch-master as how well-behaved clients learn about the change.
Own the operational framing: the window is a client-side property (event propagation, connection teardown, DNS/VIP lag), not a server timeout; contrast Redis Cluster's self-demotion at cluster-node-timeout; use client-reconfig-script for clients that don't speak Sentinel; measure the window with fault injection rather than assuming it.
Frame it as the consequence of putting agreement outside the shard: failover automation without consensus means a divergence window that no server setting bounds, whose real bound is your client fleet's worst case. Decide what data may live there at all, budget the window explicitly, and treat detection tuning and client-cutover latency as two separate clocks whose gap you are engineering.
## The shape of the problem Redis Sentinel is a set of *separate* processes that watch a master, agree it is unreachable, and promote one of its replicas. That separateness is the whole of this question. The shard itself — the master and its replicas — has no vote in the decision and no self-awareness about role changes. A Redis master does not track whether it can still reach a quorum of anything; it serves whoever connects to it. So when a partition puts the master on the minority side, two things happen at once. On the majority side, Sentinels reach agreement (subjectively down → objectively down → leader election) and promote a replica; that side now has a legitimate master. On the minority side the old master is *untouched*. Nobody has told it anything, because the only channel by which Sentinel could is the one the partition broke. It still answers `INFO` with `role:master`, still accepts `SET`, still returns `+OK`. ## Split brain, and who ends it Every client that shares the minority side with the old master keeps writing successfully. Those writes are doomed, but at the time they are indistinguishable from real ones — the client got a success reply. The critical operational fact: **nothing on the server ends this window.** In a design where the data nodes themselves run the agreement, the node can fence itself. Redis Cluster is the in-family example: a primary that cannot reach a majority of masters for `cluster-node-timeout` enters an error state and stops serving its slots. The loss window there is bounded by a server-side timer you configure. Sentinel-managed Redis has no equivalent. The window ends only when the *clients* stop talking to the old master, so it is bounded by: - **`+switch-master` propagation** — Sentinel publishes this event on its own pub/sub channel. Only Sentinel-aware clients connected to a reachable Sentinel learn from it, and clients stranded on the minority side may reach no Sentinel at all. - **Connection teardown** — an idle TCP connection to a still-alive host is not broken by a partition elsewhere in the network. A pooled connection to the old master can stay usable indefinitely. - **Indirection lag** — if clients find the master through DNS, a VIP, or a proxy, the window is that layer's TTL, health-check interval, or script latency. - **Client configuration** — a client with a hardcoded master `host:port` never re-resolves at all. Its window is unbounded. Which is why the honest answer to "how much can we lose?" is a *client-fleet* number, not a server setting. ## Why the writes vanish silently When the partition heals, the Sentinel leader sends `REPLICAOF <new-master>` to the old master. It becomes a replica, and because its replication history has diverged from the promoted node's, a partial resync is impossible — it performs a **full resync**, loads the new master's RDB, and *replaces its own dataset*. The divergent writes are not merged, not logged, not surfaced. Clients that got `+OK` were told the truth about local state and a lie about durable state, and no artifact of the discrepancy survives anywhere. ## What the operator can observe and configure - **`down-after-milliseconds`** (per monitored master in sentinel.conf) is how long a master must be unresponsive before a Sentinel marks it subjectively down. Lowering it shortens time-to-promotion — but it does **not** shorten the minority-side write window, and can widen the overlap: promoting sooner while clients still write to the old node means more divergent data. Detection latency and client-failover latency are two different clocks; the loss window is the gap between them. - **Sentinel-aware clients** subscribed to `+switch-master` that drop pooled connections and re-resolve are the primary lever. Measure it as a real number: failover event → first successful write against the new master, per client fleet. - **`client-reconfig-script`** runs on the Sentinel leader during failover with the old and new master addresses. Use it to move a VIP, reload a proxy backend, or update service discovery, so clients that do not speak Sentinel are cut over instead of being left pointing at a live-but-orphaned node. - **`replica-priority`** on each replica steers *which* node gets promoted (`0` = never promotable). It does not bound the window, but it keeps a deliberately-lagging or cross-region replica out of the running, which is part of the same failover-quality budget. - **Server-side write fencing** — making the isolated master refuse writes when too few replicas are keeping up — is the one true server-side fence available with Sentinel, and it is covered under `cluster-failover-min-replicas-to-write`. - **Observability**: alert on the Sentinel channel's `+sdown`, `+odown`, `+switch-master`, and on any period where two nodes of one shard both report `role:master`. That last signal is invisible from a single node's metrics — you have to compare across the shard. ## The judgment Sentinel buys automatic failover, not a consensus-backed shard. The divergent-write window is a property of your client fleet and your indirection layer, not of Redis; budget for it explicitly, drive it down with Sentinel-aware clients plus `client-reconfig-script`, and keep data that cannot survive it out of a Sentinel-managed Redis.
- If lowering down-after-milliseconds does not shrink the divergent-write window, what does it actually buy you, and what does it cost?It shrinks the *unavailability* window — the time clients spend unable to write anywhere because no master has been promoted yet. The cost is spurious failovers: a GC pause, a slow BGSAVE fork, or a transient blip gets read as a dead master, and every needless failover pays the same divergence and resync cost for nothing. Tune it against your observed p99 latency spikes, not against a wish for fast recovery.
- How would you measure the real size of this window in production rather than guessing at it?Instrument the two clocks separately. Subscribe a monitor to the Sentinel pub/sub channel and timestamp `+switch-master`; then, from each client fleet, record the timestamp of its first successful write against the new master. The spread between those is that fleet's failover-detection latency, and the worst fleet defines the loss window. Fault injection is the only honest way to get the numbers — drop traffic between the master and the Sentinels while leaving clients able to reach it, because a clean SHUTDOWN breaks connections and hides the problem entirely.
- A team says they are safe because clients use a DNS name for the master and client-reconfig-script updates the record. What would you push back on?DNS is a lagging indirection: record TTL, resolver caches, and runtime-level DNS caching (some JVM configurations cache forever) all sit between the script running and a client re-resolving. Worse, an already-open pooled connection to the old master never consults DNS at all, so a busy client can keep writing to the deposed node long after the record changed. The script is a reasonable fallback for clients that don't speak Sentinel, but it should point at something that actively kills connections — a proxy or VIP that drops the old backend — rather than relying on clients to notice a name change.
A branch office loses its phone line to head office. Head office appoints a new branch and carries on; nobody can call the old branch to tell it, so it keeps serving walk-in customers and stamping receipts. When the line comes back, head office simply couriers over the authoritative ledger and the branch's own book is thrown away — the customers who got a stamped receipt are never informed.
saying these in an interview costs you the question
- Saying Sentinel demotes or fences the old master at promotion — it has no channel to do so, and the partition is exactly that channel
- Claiming the split-brain window is bounded by down-after-milliseconds or some other server-side timeout, rather than by client re-resolution
- Assuming Sentinel behaves like Redis Cluster, where a partitioned primary self-demotes after cluster-node-timeout
- Believing the old master's divergent writes are merged, replayed, or logged somewhere on rejoin — they are erased by a full resync with no error to anyone
- Treating a hardcoded master host:port plus 'monitoring will catch it' as acceptable, when such a client's window is effectively unbounded
- Thinking replica-priority or client-reconfig-script prevents divergent writes; they steer promotion and cut over clients, they do not stop the old node from accepting writes