In a Redis deployment fronted by Sentinel, how should an application client find the current master's address, and what goes wrong if it just configures the master's IP directly?
answer
- Configure Sentinels + master name, never a master IP
- `SENTINEL get-master-addr-by-name`
- Subscribe `+switch-master`, then CLOSE old connections
- Old master returns as replica → `-READONLY` on writes
- Fallback: `client-reconfig-script` moves VIP/DNS
basics
~20 sThe client is given the list of Sentinel addresses plus the master's logical name. It asks a Sentinel with SENTINEL get-master-addr-by-name, connects there, and subscribes to the +switch-master event so it drops connections and re-resolves after a failover. A hardcoded master IP keeps pointing at the dead or demoted node.
solid answer
~60 s**Configure Sentinels, not the master.** Every mainstream client (Lettuce, Jedis, redis-py, go-redis, ioredis, StackExchange.Redis) has a Sentinel mode where you pass a list of Sentinel host:port pairs and a **master name** such as `mymaster`. The client then: 1. Connects to a Sentinel and issues `SENTINEL get-master-addr-by-name mymaster`, trying the next Sentinel if one is unreachable. 2. Optionally verifies the returned node really reports `role:master` via `INFO` — protection against a stale answer. 3. Opens the data connection directly to that address. 4. **Subscribes to the `+switch-master` Pub/Sub channel** on a Sentinel so it learns about a failover immediately, tears down connections to the old address, and re-resolves. Replica-read modes similarly track `SENTINEL replicas`. If you hardcode the master IP, the failover still happens correctly at the Redis layer, but your app keeps talking to the old node: it will time out while the old master is dead, and — worse — once the old master returns as a **replica**, writes fail with `READONLY You can't write against a read only replica`. The alternatives are a VIP or DNS record moved by Sentinel's `client-reconfig-script`, but native Sentinel clients are the standard answer.
code
text · 14 lines# ask any Sentinel
> SENTINEL get-master-addr-by-name mymaster
1) "10.0.0.11"
2) "6379"
# verify before writing
> redis-cli -h 10.0.0.11 INFO replication | head -1
role:master
# stay subscribed for changes
> SUBSCRIBE +switch-master
1) "message"
2) "+switch-master"
3) "mymaster 10.0.0.10 6379 10.0.0.11 6379"go deeper
Say that the app is configured with the Sentinel addresses and the master name, and asks Sentinel for the current master rather than hardcoding an IP.
Add the two mechanisms — SENTINEL get-master-addr-by-name plus the +switch-master subscription — and the -READONLY symptom you get when the demoted old master is still in the pool.
Discuss verifying role:master against a stale Sentinel answer, dropping connections on failover events, timeout/retry policy so a failover does not exhaust thread pools, and rehearsing with a forced failover.
Own the end-to-end availability budget: Sentinel's detection time plus client re-resolution plus timeout policy is the real recovery time, and the fallbacks (VIP, DNS, proxy) each reintroduce a component whose failure characteristics you must account for.
## Why discovery is the client's job Sentinel is not in the data path. It cannot silently move traffic; all it can do is *tell* whoever asks where the master currently lives, and *announce* when that changes. So high availability with Sentinel is only as good as the client integration. A perfectly healthy Sentinel quorum plus a hardcoded IP in `application.yml` equals an outage that Sentinel dutifully logs as successfully resolved. ## The two APIs the client uses **Request/response:** ``` SENTINEL get-master-addr-by-name mymaster 1) "10.0.0.11" 2) "6379" ``` Any Sentinel can answer. Because a Sentinel that is itself partitioned may return a stale address, robust clients treat the answer as a hint and confirm with `INFO replication` that `role:master` is present before using it for writes; if not, they try another Sentinel. **Push:** ``` SUBSCRIBE +switch-master # message: mymaster 10.0.0.10 6379 10.0.0.11 6379 ``` The payload is `<master-name> <old-ip> <old-port> <new-ip> <new-port>`. On receipt the client must **close its existing connections** to the old address, not merely note the new one — an established TCP connection to a demoted node stays usable and will happily accept reads while rejecting writes. Related events worth reacting to: `+sdown`/`-sdown` and `+slave` for read-replica pools, and `+redirect-to-master`. ## What the client library actually does for you Given a Sentinel configuration, a good client: - Maintains a pool of Sentinel connections and rotates through them; a Sentinel being down must not block startup. - Reorders its Sentinel list so the last responsive one is tried first. - Keeps one subscriber connection for failover events. - Refreshes the topology periodically as a backstop in case a `+switch-master` message was missed while disconnected. - Exposes separate "master" and "replica" connection intents so reads can be routed to replicas while writes always go to the resolved master. Spring Data Redis, for example, takes a `RedisSentinelConfiguration` with the master name and Sentinel nodes; Lettuce underneath does the resolution and event subscription. ## The failure modes of hardcoding 1. **During the outage** — connections to the dead master hang until the socket timeout; without a sane `connectTimeout`/command timeout, threads pile up and the failure spreads through your service. 2. **After the failover** — the old master returns as a replica. Reads succeed (which is why this bug survives smoke tests), writes fail with `-READONLY`. Some apps interpret that as a transient error and retry forever. 3. **Split-brain window** — clients still attached to a partitioned old master write data that is discarded when it resynchronizes. Sentinel-aware clients close those connections as soon as they see `+switch-master`; a hardcoded client does not. 4. **Silent staleness** — if you point reads at a fixed replica IP that gets promoted, you have now made your "read replica" the master and are unknowingly sending read load to it. ## Alternatives when the client cannot speak Sentinel Some stacks (a legacy driver, a sidecar, a tool that only accepts one host) cannot do Sentinel discovery. Two standard workarounds: - **`client-reconfig-script`** — Sentinel invokes a script on the leader Sentinel during failover with the master name, role, state, and old/new addresses. The script moves a floating VIP, updates a DNS record with a short TTL, or reconfigures a proxy such as HAProxy. Caveats: it runs only on the Sentinel that led the failover, it must be idempotent, and DNS TTLs plus JVM DNS caching add real delay. - **A proxy layer** (HAProxy, Envoy, or a Redis-aware proxy) that health-checks `role:master` and exposes a single stable endpoint. This restores a single point of failure unless the proxy is itself made HA, and adds a hop of latency. ## Testing it The only trustworthy verification is a **forced failover under load**: run `SENTINEL failover mymaster` while your application is writing, and assert that write errors stop within your target window and that no thread pool is exhausted. Many teams discover their timeouts and retry policy — not Sentinel — are the reason a 10-second failover becomes a 3-minute outage.
- Why is subscribing to `+switch-master` better than just re-resolving on error?Error-driven re-resolution only kicks in after requests have already failed, and some of them fail slowly (socket timeouts) or not at all (reads against the demoted replica succeed). The push event lets the client invalidate its pool the moment the topology changes, so the failure window is bounded by the failover itself rather than by your timeout settings. Periodic re-resolution is still worth keeping as a backstop for missed messages.
- A team cannot use a Sentinel-aware driver. What are their options?They can have Sentinel run a `client-reconfig-script` on failover that moves a virtual IP, updates a low-TTL DNS record, or reconfigures a proxy such as HAProxy that health-checks for `role:master`. Both add latency and new failure modes — the script runs only on the leading Sentinel and must be idempotent, and DNS caching in the client runtime can outlive the TTL — so they are fallbacks, not the preferred design.
- How would you prove the client integration actually works?Trigger `SENTINEL failover mymaster` while the application is under write load and measure how long writes fail and whether they recover without a restart. Watch for connection-pool exhaustion and for `-READONLY` errors persisting, which indicate the client is not dropping connections to the demoted node. Doing this in staging on every dependency upgrade is cheap insurance.
saying these in an interview costs you the question
- Configuring the master's host in the app and assuming Sentinel redirects traffic.
- Thinking Sentinel proxies or load-balances so the client needs no changes.
- Handling `+switch-master` by noting the new address but leaving old connections open.
- Treating `-READONLY` errors as a transient failure to retry rather than a stale-topology bug.
- Assuming DNS-based failover is instant, ignoring TTLs and runtime DNS caching.