skip to content

Walk through how Redis Sentinel detects that a master has failed and promotes a replica. What do the terms SDOWN and ODOWN mean in that flow?

level: seniorimportance: must knowfreq 50%

answer

  1. SDOWN = my opinion; ODOWN = quorum's opinion
  2. Only masters reach ODOWN
  3. Leader needs majority, not quorum
  4. Rank: not-down → priority → offset → run ID
  5. priority 0 = never promote

basics

~20 s

A Sentinel marks the master SDOWN (subjectively down) after it fails to reply for down-after-milliseconds. It asks the other Sentinels; when quorum agree it becomes ODOWN (objectively down). Sentinels then elect a leader by majority, which picks the best replica, sends REPLICAOF NO ONE, and repoints the others.

solid answer

~50 s

**Detection.** Each Sentinel pings the master every second. If it gets no valid reply for `down-after-milliseconds`, that Sentinel flags the master **SDOWN — subjectively down**: one process's private opinion, possibly just its own network problem. **Agreement.** The Sentinel asks the others with `SENTINEL is-master-down-by-addr`. Once at least `quorum` Sentinels report SDOWN, the master is **ODOWN — objectively down**. Only ODOWN starts a failover, and only for masters (replicas never go beyond SDOWN). **Leader election.** Failover is performed by exactly one Sentinel. Sentinels run a Raft-like term-based vote; a candidate needs a **majority of the whole Sentinel set** — not just the quorum — so a 3-Sentinel deployment needs 2 votes. **Promotion.** The leader ranks replicas by: not SDOWN/disconnected, `replica-priority` (0 = never), largest replication offset, then lowest run ID. It sends `REPLICAOF NO ONE` to the winner, waits for role change, then reconfigures the remaining replicas `parallel-syncs` at a time, publishes `+switch-master`, and rewrites configs. The old master returns as a replica.

code

text · 9 lines
text
$ redis-cli -p 26379 PSUBSCRIBE '*'
+sdown master mymaster 10.0.0.10 6379
+odown master mymaster 10.0.0.10 6379 #quorum 2/2
+try-failover master mymaster ...
+vote-for-leader 7f3c... 12
+selected-slave slave 10.0.0.11:6379 ...
+promoted-slave slave 10.0.0.11:6379 ...
+switch-master mymaster 10.0.0.10 6379 10.0.0.11 6379
+failover-end master mymaster

go deeper

for a junior

Know the two words: SDOWN is one Sentinel's opinion, ODOWN is when enough of them agree, and only then is a replica promoted.

for a middle

Add the timeline — down-after-milliseconds, is-master-down-by-addr polling, quorum, promotion via REPLICAOF NO ONE, and +switch-master for clients.

for a senior

Explain leader election and epochs, the replica-ranking rules, parallel-syncs, failover-timeout, and how to read the Pub/Sub event log when diagnosing a spurious failover caused by a stalled master.

for a principal

Discuss the tuning tradeoff explicitly: down-after-milliseconds trades detection latency against false failovers, and every false failover costs data because replication is asynchronous; set it from measured p99 stalls and pair it with removing the stall sources.

## Step 0: what Sentinel is measuring Every Sentinel opens two connections to each monitored instance: a command connection (used for `PING`, `INFO`, `PUBLISH`) and a subscription connection to the `__sentinel__:hello` channel. `PING` goes out every second; `INFO` every ten seconds (more often during a failover). "Not replying" means: no `+PONG`, `-LOADING`, or `-MASTERDOWN` reply within the timeout — an error reply still counts as being alive but unhealthy for some cases, while `-BUSY` (a long-running Lua script) is treated as down after the timeout. ## Step 1: SDOWN — subjectively down If a Sentinel receives nothing valid for `down-after-milliseconds` (5–30 s typically; the tuning knob for how twitchy your cluster is), it flags the instance **SDOWN**. The word *subjective* is the point: this is one Sentinel's local belief. The Sentinel's own NIC, a one-sided partition, or a GC-frozen host can all produce SDOWN on a perfectly healthy master. A second SDOWN trigger exists for masters: a node that reports itself as a *replica* in `INFO` for long enough is also treated as down, because it is no longer serving the master role. SDOWN alone changes nothing about the topology. It emits a `+sdown` event, which is exactly what you want your alerting on. ## Step 2: ODOWN — objectively down The SDOWN-ing Sentinel now asks its peers `SENTINEL is-master-down-by-addr <ip> <port> <epoch> <runid>`. Each peer answers with its own view. When the number of Sentinels agreeing reaches the configured **quorum** (the last number on the `sentinel monitor` line), the master is promoted to **ODOWN** and `+odown` is published. Two important asymmetries: - **Only masters reach ODOWN.** Replicas and other Sentinels that are down stay at SDOWN, because nothing needs to be coordinated to stop using them. - **ODOWN is a fresh, non-sticky computation** based on replies received within a short window, so a flapping master does not accumulate stale votes. ## Step 3: leader election ODOWN authorizes a failover but does not perform it — if every Sentinel promoted a replica you would get several new masters. Sentinels therefore elect one leader using a **Raft-inspired algorithm with configuration epochs**. A Sentinel that sees ODOWN increments the epoch, votes for itself, and asks others for their vote (piggybacked on the same `is-master-down-by-addr` command, which returns the voted leader's run ID). Each Sentinel grants at most one vote per epoch, first come first served. The crucial rule: the winner needs **a strict majority of all known Sentinels** (`num_sentinels/2 + 1`), regardless of how small the quorum is. That is why quorum can never substitute for having enough live Sentinels — more on that in the sizing question. If no one wins the epoch (split vote), Sentinels back off a random delay and retry with a higher epoch. ## Step 4: choosing the replica The leader filters and ranks candidates: 1. **Discard** replicas that are SDOWN, disconnected, not seen recently, or whose link with the master has been down longer than roughly `down-after-milliseconds * 10 + (time master was down)`. 2. **`replica-priority`** (formerly `slave-priority`), lowest wins; **priority 0 makes a replica ineligible for promotion** — the standard way to pin a DR-site copy out of the running. 3. **Largest replication offset** — the replica that processed the most of the master's stream, minimizing data loss. 4. **Lexicographically smallest run ID** as a deterministic tiebreak. ## Step 5: promotion and reconfiguration The leader sends `REPLICAOF NO ONE` to the winner and polls its `INFO` until `role:master` appears. It then updates its own configuration with a new epoch, publishes `+switch-master <name> <old-ip> <old-port> <new-ip> <new-port>` — the event application clients subscribe to — and sends `REPLICAOF <new-master>` to the remaining replicas, at most **`parallel-syncs`** at a time so they are not all busy doing a full sync simultaneously and unavailable for reads. The other Sentinels learn the new configuration through hello messages and adopt it because it carries a higher epoch; each rewrites its own config file. When the old master returns, Sentinel makes it a replica of the new master — its unreplicated writes are discarded during resynchronization. `failover-timeout` bounds each phase; a stalled failover is aborted and retried, and it also rate-limits how soon the same master can be failed over again. ## Diagnosing it afterwards Sentinel's Pub/Sub log is the timeline: `+sdown` → `+odown` → `+try-failover` → `+vote-for-leader` → `+selected-slave` → `+promoted-slave` → `+failover-state-reconf-slaves` → `+switch-master` → `+failover-end`. If you see repeated `+sdown`/`-sdown` pairs with no failover, `down-after-milliseconds` is too aggressive for your network or your master is stalling (long-running commands, fork pauses, swap).

  • Why does Sentinel need a leader election at all once quorum agrees the master is down?
    ODOWN is only a shared diagnosis; the failover itself is a mutating action that must happen exactly once. If every Sentinel that saw ODOWN promoted a replica, you would end up with multiple masters and divergent data. The epoch-based majority election guarantees a single actor per failover attempt, and the epoch lets other Sentinels recognize and adopt the newest configuration.
  • How does Sentinel choose which replica to promote, and how do you influence it?
    It discards replicas that are down or whose link to the master was broken for too long, then orders the rest by `replica-priority` (lowest first), then by largest replication offset, then by smallest run ID. You influence it by setting `replica-priority`: a lower number makes a replica preferred, and 0 makes it permanently ineligible — the usual setting for a cross-region or backup-only copy.
  • A master pauses for 8 seconds running a heavy Lua script or forking for a background save. What does Sentinel do?
    If the pause exceeds `down-after-milliseconds`, Sentinels see no valid PING reply, flag SDOWN, reach ODOWN, and fail over a master that was never actually dead — and the promoted replica is missing whatever the old master had not shipped. The fixes are to remove the stalls (avoid O(N) commands and long scripts, tune persistence and fork behavior) and to set `down-after-milliseconds` above your realistic worst-case pause.

saying these in an interview costs you the question

  • Saying one Sentinel noticing the master is down is enough to trigger a failover.
  • Confusing quorum with the majority needed to elect a leader — they are separate thresholds.
  • Believing replicas also go ODOWN and get failed over.
  • Claiming Sentinel picks a replica at random, or ignoring `replica-priority` and replication offset.
  • Assuming SDOWN always means the master is really dead rather than possibly a one-sided network or a stalled event loop.

context