skip to content

An ElastiCache for Valkey replication group is Multi-AZ with automatic failover enabled, and its primary node fails. Walk through what ElastiCache does, what the application sees, and what you must have configured beforehand for the recovery to be clean.

level: seniorimportance: should knowfreq 46%

answer

  1. promote a replica, repoint the DNS
  2. the window is seconds, not zero
  3. asynchronous means unshipped writes vanish
  4. stale DNS keeps clients on the dead node
  5. the promoted node is already warm

basics

~20 s

ElastiCache detects the failure, promotes a replica, and repoints the primary endpoint's DNS at it. Clients see dropped connections and errors for tens of seconds, and any write acknowledged but not yet replicated is lost, because replication is asynchronous.

solid answer

~50 s

ElastiCache detects the primary is unhealthy, promotes one of the replicas, updates the DNS record behind the **primary endpoint** to the new node, and then provisions a replacement node that rejoins the group as a replica. The application sees its connections to the old primary drop and gets errors for the promotion window — tens of seconds, not milliseconds. Because replication is asynchronous, writes the old primary acknowledged but had not yet shipped are gone; the cache is not a durable store and should never be treated as one. What makes recovery clean is client-side: a short DNS cache so the primary endpoint re-resolves promptly, a connection pool that discards broken connections rather than retrying on them, and retry-with-backoff so the reconnect storm does not become the incident. The upside over a Memcached node replacement is that the promoted replica already holds the data, so the database behind the cache does not face a cold cache.

code

bash · 9 lines
bash
# Rehearse the failover rather than discovering client behaviour in an incident
aws elasticache test-failover \
  --replication-group-id app-cache \
  --node-group-id 0001

# Confirm Multi-AZ and automatic failover are actually on
aws elasticache describe-replication-groups \
  --replication-group-id app-cache \
  --query 'ReplicationGroups[0].[MultiAZ,AutomaticFailover]'

go deeper

for a junior

Know that ElastiCache promotes a replica automatically and moves the primary endpoint to it, and that the application must reconnect rather than assuming the connection survives.

for a middle

Explain the ordering — detect, promote, repoint DNS, replace the node — and why asynchronous replication means some acknowledged writes are lost in the process.

for a senior

Show the operational readiness: short DNS TTLs, pools that discard dead connections, backoff with jitter, a decided policy for what the app does while the cache is unreachable, and rehearsed failovers.

for a principal

Own the standard across services — the failure policy when the cache is down, the load the databases must be able to absorb, and whether any workload is allowed to depend on cache state surviving a failover at all.

## The sequence ElastiCache runs When a Multi-AZ replication group with automatic failover loses its primary: 1. **Detection.** ElastiCache's health checks stop getting healthy responses from the primary node. 2. **Promotion.** One of the replicas is promoted to primary. ElastiCache prefers a replica that is furthest along in applying the replication stream, which minimises — but does not eliminate — data loss. 3. **DNS update.** The record behind the **primary endpoint** is repointed at the promoted node. This is why applications must use that endpoint and never an individual node endpoint. 4. **Replacement.** A new node is provisioned to restore the replica count, and it syncs from the new primary. Depending on dataset size this continues in the background well after writes have resumed. Multi-AZ requires at least one replica, and it only means something when that replica sits in a different Availability Zone — a replication group whose replicas share an AZ with the primary survives a node failure but not an AZ event. ## What the application actually experiences - **Connection loss.** Every connection to the old primary breaks. Whatever was in flight fails. - **An error window.** From the moment the primary stops responding to the moment clients successfully reconnect to the new one, writes fail. AWS targets tens of seconds for the promotion itself; how long *your* application stays broken is usually dominated by client behaviour, not by ElastiCache. - **Read-only errors.** A client whose DNS cache still resolves the primary endpoint to the old node will, once that node returns as a replica, get a read-only error on every write. The symptom looks like a broken cluster; the cause is stale resolution. - **Possible data loss.** Replication is asynchronous. A write acknowledged by the old primary and not yet shipped to the promoted replica is simply gone. Any design that treats the cache as the record of truth for something — a counter, a lock, a queue — has to answer for that window. ## What preparation makes it clean The failover is AWS's job; surviving it is yours. - **Short DNS caching.** The classic offender is a JVM with a long or infinite DNS cache: the `networkaddress.cache.ttl` security property must be set low (AWS SDK guidance is 60 seconds or less) or the client will keep resolving to the dead node long after DNS moved. - **A pool that discards.** Connection pools must detect and drop broken connections and validate on borrow, rather than handing out sockets to a node that no longer exists. - **Retry with backoff and jitter.** Every client in the fleet reconnecting at once is a stampede. Backoff plus jitter turns a spike into a ramp. - **A failure policy for the cache being unavailable.** Decide in advance whether the application fails open to the database (and whether the database can take that load) or fails closed with a degraded response. A cache outage that silently converts into a database overload is the more expensive incident. - **Rehearsal.** `aws elasticache test-failover` triggers a real failover on demand, which is how you find out that a client library never re-resolves before an unplanned event does. ## The one genuinely good news Because the promoted replica already held a copy of the dataset, the cache is *warm* on the other side of the failover. This is the concrete operational advantage over engines with no replication, where a node replacement leaves an empty cache and the source of truth absorbs that node's entire read share at once. It is also the sharpest argument for choosing a replicated engine when the database behind the cache could not survive losing it. ## Related knobs worth knowing - **Automatic failover** is a prerequisite for Multi-AZ; without it, a primary failure waits for you. - **Planned maintenance and scaling** also trigger failovers. A node-type change on a cluster-mode-disabled group involves one. So the failover path is exercised by routine operations, not just by disasters — which is another reason to make clients handle it well. - **Cluster mode enabled** shrinks the blast radius: each shard fails over independently, so only the slice of the keyspace on that shard is disrupted. ## What weak answers sound like "It's Multi-AZ, so there's no downtime" is the headline error — there is always a window. "The replica starts empty" confuses promotion with replacement. And "no data can be lost because there's a replica" ignores that the replication is asynchronous, which is the single most important sentence in the whole answer.

  • Your fleet reconnects successfully after a failover, but the database behind the cache spikes hard for a minute. What happened, and how would you prevent it?
    During the error window the application most likely failed open and sent those requests straight to the database, and every client retried at once. Prevent it with backoff plus jitter on reconnect, a concurrency limit or circuit breaker on the fall-through path to the database, and a deliberate decision about whether a cache outage should degrade responses instead of amplifying load.
  • How does cluster mode enabled change the impact of a primary failure?
    Each shard has its own primary and fails over independently, so only the keys owned by that shard are disrupted instead of the entire keyspace. The client, being cluster-aware, refreshes its topology and keeps serving the other shards throughout. The trade is that you now operate several failover domains rather than one.
  • Can you rely on ElastiCache automatic failover to protect a distributed lock held in the cache?
    No. Replication is asynchronous, so a lock acquired on the old primary may never have reached the promoted replica, and after failover two holders can believe they own it. If the lock protects something whose double-execution matters, put the authority in a store with real durability and consensus, and treat the cache lock as an optimisation.

saying these in an interview costs you the question

  • Claims Multi-AZ means zero downtime for clients
  • Thinks the promoted replica starts with an empty dataset
  • Says no writes can be lost because a replica exists
  • Ignores client DNS caching as a source of prolonged outage
  • Treats the cache as durable storage for counters or locks

context