skip to content

How many Redis Sentinel processes should you deploy, and what is the difference between the configured quorum value and the majority a Sentinel needs to actually perform a failover?

level: seniorimportance: should knowfreq 40%

answer

  1. Odd N, minimum 3, independent hosts
  2. quorum → ODOWN diagnosis
  3. majority N/2+1 → leader election, not configurable
  4. Effective need = max(quorum, majority)
  5. `SENTINEL CKQUORUM` in monitoring

basics

~20 s

Run an odd number, at least three, on hosts that fail independently. Quorum is how many Sentinels must agree the master is down to declare ODOWN. Performing the failover additionally requires a leader elected by a majority of all Sentinels (N/2+1). Quorum can lower but never bypass that majority.

solid answer

~1 min

**Sizing.** Three Sentinels minimum, on three independently failing hosts (different machines/AZs), and an **odd** count so a partition cannot produce two equal halves. Five is common when you want to survive two failures or spread across three zones. Do not co-locate all Sentinels on the Redis nodes only — with a two-node Redis setup that gives you no majority when the master's host dies. **Two thresholds, and they are not the same.** - **`quorum`** — last argument of `sentinel monitor` — is the number of Sentinels that must independently see SDOWN before the master is declared **ODOWN**. It tunes *how easily failure is diagnosed*. - **Majority (`N/2 + 1` of all known Sentinels)** is what the failover-performing Sentinel must win in the leader election. This is **not configurable** and exists to prevent two Sentinels on opposite sides of a partition from both promoting a replica. So with 5 Sentinels and quorum 2: two Sentinels can declare ODOWN, but nothing happens unless at least three are alive to elect a leader. Setting quorum below the majority makes detection more sensitive, never failover more available. Setting quorum higher than the majority makes failover *harder* — it then needs `max(quorum, majority)` reachable Sentinels.

code

text · 7 lines
text
$ redis-cli -p 26379 SENTINEL CKQUORUM mymaster
OK 3 usable Sentinels. Quorum and failover authorization can be reached

$ redis-cli -p 26379 SENTINEL set mymaster quorum 2   # apply on every Sentinel

# after decommissioning a Sentinel host, forget it everywhere:
$ redis-cli -p 26379 SENTINEL RESET mymaster

go deeper

for a junior

Know the rule of thumb: at least three Sentinels, odd number, on different machines, quorum usually set to a majority.

for a middle

Explain the two thresholds separately — quorum gates ODOWN, majority gates who performs the failover — and why an even count adds nothing.

for a senior

Reason about failure domains and partitions concretely (which side keeps a majority), tune quorum against the cost of a spurious failover, and monitor with CKQUORUM including the stale-Sentinel trap.

for a principal

Treat it as a quorum-placement problem: two symmetric sites cannot self-heal safely, so either buy a third failure domain for the tiebreaker or accept manual failover, and state which class of outage each choice covers.

## Why there are two numbers at all Sentinel separates **diagnosis** from **action**. Diagnosis ("is the master down?") should reflect several independent viewpoints so one bad NIC does not cause a failover — that is `quorum`. Action ("promote this replica") must happen at most once across the whole system, even during a network partition — that requires a majority, because only one side of a partition can hold a majority of a fixed-size set. This is the same reason Raft and other consensus protocols use majorities. Candidates who blur the two produce the classic wrong answer: "set quorum to 1 so failover always works." With quorum 1 in a 3-Sentinel deployment, a single Sentinel can declare ODOWN — and then still has to get 2 votes to lead the failover. If it is alone on its side of a partition, it gets nothing. All quorum 1 achieved was making spurious `+odown` events likely. ## Choosing N **Odd numbers.** With an even N, a symmetric partition splits into two equal halves and neither has `N/2 + 1`; you gain nothing over N-1 and add a host that can fail. **Independent failure domains.** The unit of failure is not the process, it is the host, rack, availability zone, hypervisor, and power feed. Three Sentinels in one AZ tolerate one VM loss but not the AZ. A common layout is one Sentinel per AZ across three AZs, with Redis master+replica in two of them — the third Sentinel is the tiebreaker that makes a majority possible when an entire zone goes dark. **Do not use N = 2.** Majority of 2 is 2, so losing either Sentinel makes failover impossible — strictly worse than a single Sentinel plus a human. **Sentinels can share hosts with Redis** and often do, but if you have exactly two Redis nodes you need a Sentinel on a third machine (an app server, a small tiebreaker instance) or the death of the master's host leaves one Sentinel out of two. ## Choosing the quorum value The standard setting is the majority: **quorum = N/2 + 1** (2 of 3, 3 of 5). Then the two thresholds coincide and behaviour is easy to reason about. Lower it (quorum 2 of 5) when you want faster or more sensitive detection and you accept that a minority's opinion is enough to *start* the process — the majority election still gates the action. This is legitimate; it just does not increase availability. Raise it above the majority (quorum 4 of 5) when a spurious failover is more expensive than a delayed one — for example, when the replica is in another region and promoting it means a painful traffic shift. You are trading recovery time for confidence, and you should say that explicitly in an interview. ## Worked examples | N | majority | quorum | Effect | |---|---|---|---| | 3 | 2 | 2 | Standard. Survives one Sentinel loss. | | 3 | 2 | 1 | Twitchy diagnosis, same availability. Rarely useful. | | 3 | 2 | 3 | Any single Sentinel loss disables failover. Bad. | | 5 | 3 | 3 | Standard. Survives two losses. | | 5 | 3 | 2 | Faster ODOWN; failover still needs 3 alive. | | 4 | 3 | 2 | Even N: tolerates only one loss, same as N=3. Avoid. | ## Partition behaviour, made concrete Three Sentinels: S1 with the master in AZ-A; S2 and S3 with a replica in AZ-B. AZ-A is cut off. - S2 and S3 see SDOWN, reach quorum 2 → ODOWN, elect a leader with 2 of 3 votes, promote the replica. Correct: the majority side wins. - S1 sees the replicas as down but there is no failover of replicas, and S1 alone can never gather 2 votes, so it cannot promote anything. Correct: the minority side is inert. - Clients stuck with the old master in AZ-A keep writing until either they notice, or `min-replicas-to-write` on the old master makes it refuse writes because it has zero healthy replicas. Those writes are lost on rejoin — which is why the durability knobs are a separate, complementary control. ## Operational notes - Sentinels **auto-discover** peers; adding one is just starting it with the same `sentinel monitor` line. Removing one requires `SENTINEL RESET` on the others (or `SENTINEL REMOVE`) so the majority is not computed against a Sentinel that will never return — a forgotten decommissioned Sentinel silently raises your majority threshold. - `SENTINEL set mymaster quorum <n>` changes the quorum at runtime, per Sentinel; apply it everywhere. - `SENTINEL CKQUORUM mymaster` reports whether the current live set can both reach quorum and elect a leader. Run it in monitoring — it is the cheapest way to catch "we have been one Sentinel away from unable-to-fail-over for three months".

  • Is it ever useful to set quorum to 1?
    Almost never. It only makes ODOWN easier to declare, and the failover still requires a majority to elect a leader, so availability does not improve. What you get instead is more spurious `+odown` events from single-host network blips, and a higher chance of failing over a master that was merely stalled. If detection feels too slow, tune `down-after-milliseconds` rather than dropping quorum to 1.
  • You have two data centres. Where do the Sentinels go?
    Two symmetric sites cannot form a majority when the link between them breaks, so no automatic failover is safe — whichever side you weight will either fail over spuriously or never fail over. The standard fix is a third site (or a cheap third host, even just a small VM) holding one Sentinel as a tiebreaker, giving 2+2+1; then a site loss still leaves a majority. Without a third location, accept manual failover.
  • Why does a decommissioned Sentinel that was never removed cause problems?
    Sentinels remember peers they have discovered, and the majority is computed over the known set. A ghost Sentinel raises N without ever voting, so a 3-Sentinel deployment that is really 2 live plus 1 dead now needs both live ones to agree. `SENTINEL RESET` (or `SENTINEL REMOVE`) clears the stale entry, and `CKQUORUM` is what tells you it happened.

Quorum is how many inspectors must sign off that the building is on fire; the majority vote is the rule that only one fire chief may order the evacuation — you can lower the number of signatures required, but you still cannot have two chiefs giving opposite orders.

saying these in an interview costs you the question

  • Treating quorum as the only threshold and forgetting the majority requirement for leader election.
  • Claiming a low quorum makes failover more available.
  • Deploying two or four Sentinels, or all of them in one availability zone.
  • Putting one Sentinel per Redis node with only two Redis nodes and expecting failover to work when the master's host dies.
  • Forgetting that removed Sentinels stay in the known set until `SENTINEL RESET`, silently raising the majority.

context