An ElastiCache for Valkey replication group with cluster mode disabled exposes a primary endpoint and a reader endpoint. What does each one resolve to, and what happens if the application sends its writes to the reader endpoint?
answer
- one name follows the primary
- the other spreads over replicas
- balancing happens per connection, not per command
- replicas refuse writes outright
- stale DNS produces the same symptom
basics
~20 sThe primary endpoint is a DNS name that always tracks the current primary, including after a failover. The reader endpoint resolves across the read replicas. Writes sent to a reader are rejected by the engine with a read-only error.
solid answer
~50 sThe **primary endpoint** is a DNS name ElastiCache repoints at whichever node is currently primary, so an application that uses it keeps writing to the right node across a failover. The **reader endpoint** resolves to the read replicas, spreading *new connections* across them rather than balancing individual commands — a long-lived pooled connection stays pinned to whichever replica it first resolved to. Individual node endpoints also exist, but hardcoding one ties you to a node that may be demoted or replaced. If the application writes through the reader endpoint, the engine rejects the command with a read-only error rather than forwarding it to the primary. This is a classic misconfiguration: someone points the whole client at the reader endpoint for "load balancing" and every write path fails. The fix is two connection settings — writes on the primary endpoint, reads on the reader endpoint — and accepting that replica reads can be slightly behind.
code
bash · 8 lines# Rehearse a failover and watch which node the primary endpoint follows
aws elasticache test-failover \
--replication-group-id app-cache \
--node-group-id 0001
aws elasticache describe-replication-groups \
--replication-group-id app-cache \
--query 'ReplicationGroups[0].NodeGroups[0].NodeGroupMembers[].[CacheClusterId,CurrentRole]'go deeper
Know that one endpoint is for writes and always tracks the primary, the other is for reads across replicas, and that a replica rejects writes rather than passing them on.
Explain that the reader endpoint distributes connections rather than commands, so pooled connections stay pinned, and describe the two-client configuration that keeps writes and bulk reads separate.
Diagnose a read-only error by distinguishing wrong-endpoint configuration from stale DNS after a failover, and decide which access paths may tolerate replica staleness.
Set the client standard across services: which endpoints are permitted in configuration, the DNS caching policy, and the review rule that keeps read-after-write paths off replicas.
## The endpoints ElastiCache hands you A cluster-mode-disabled replication group publishes several DNS names, and knowing which is which is most of this question. - **Primary endpoint** — a DNS name that always resolves to the node currently acting as primary. When ElastiCache promotes a replica, it updates this record. Applications that write should use only this. - **Reader endpoint** — a DNS name that resolves across the group's read replicas, distributing **connections** among them. It exists so you do not have to enumerate replicas yourself or update configuration when the replica count changes. - **Node endpoints** — one per node, useful for diagnostics and for clients that manage routing themselves, but hazardous in application configuration because the node behind it can be demoted, replaced, or removed. A cluster-mode-enabled group is different: it publishes a **configuration endpoint** that a cluster-aware client uses to discover the topology, and routing to primaries and replicas is the client's job. ```bash aws elasticache describe-replication-groups \ --replication-group-id app-cache \ --query 'ReplicationGroups[0].NodeGroups[0].[PrimaryEndpoint.Address,ReaderEndpoint.Address]' ``` ## What "balanced" actually means The reader endpoint balances at **DNS resolution time**. Each fresh resolution can hand back a different replica address, so a fleet of processes that each open a pool at startup ends up roughly spread. But within one process, a pooled connection resolves once and then stays on that replica for its lifetime. Two consequences follow: 1. **Balance is coarse.** Small fleets, or fleets that all start at the same moment behind an aggressive DNS cache, can land disproportionately on one replica. This is not a bug in the endpoint; it is what connection-level balancing means. 2. **Skew persists.** Because pools are long-lived, an uneven distribution does not self-correct until connections are recycled. If you need per-command routing, that belongs in the client library — many clients expose a read-routing preference over the discovered topology (Lettuce's `ReadFrom` setting is one example) rather than relying on the reader endpoint. ## Writing to a reader Send a write down a connection that landed on a replica and the engine answers with a read-only error along the lines of `-READONLY You can't write against a read only replica`. There is no forwarding, no proxying, and no silent success. The application sees an error on every write path while reads look perfectly healthy — which is why this misconfiguration often survives a staging environment where write volume is low or where the group had no replicas at all. The same error appears in a second, more interesting situation: a failover has happened, the client's DNS cache still points the *primary* endpoint at the old node, and that node has come back as a replica. The symptom is identical; the cause is stale resolution rather than wrong configuration. Distinguishing the two is a good diagnostic instinct — check whether the hostname in use is the reader endpoint or the primary endpoint, and check how long the client caches DNS. ## The consistency cost of using readers Choosing the reader endpoint is choosing to read from asynchronously updated copies. A value written and immediately read back through the reader endpoint may not be there yet. For a cache that is often fine, but the decision must be deliberate: read-after-write paths, anything driving a UI immediately after a mutation, and anything used for a correctness check should stay on the primary endpoint. ## How to configure it properly The pattern that survives production: - one client configured with the **primary endpoint** for writes and any read-after-write path; - one client configured with the **reader endpoint** for bulk read traffic; - **no** individual node endpoints anywhere in configuration; - a short DNS cache so both endpoints re-resolve promptly after a topology change; - retry logic that treats a read-only error as a signal to reconnect and re-resolve rather than to retry blindly on the same connection. ## The misconceptions to name out loud "The reader endpoint load-balances every request" is the most common. "Writes to a reader are forwarded to the primary" is the most damaging, because it makes people design as if the endpoint were a proxy. And "we point everything at the reader endpoint to spread load" is the misconfiguration that generates the error in the first place.
- An application reads a key immediately after writing it and sometimes gets the old value. It uses the primary endpoint for writes and the reader endpoint for reads. What is happening?Replication to the replicas is asynchronous, so the read reached a replica that had not yet applied the write. Nothing is broken. Either route read-after-write paths to the primary endpoint, or make the caller tolerate the staleness. Splitting reads by access path — fresh-critical on the primary, bulk on readers — is the usual fix.
- Why is hardcoding an individual node endpoint in application config a problem?That name points at one specific node. After a failover the node behind it may be a replica rather than the primary, so writes start failing with a read-only error; after a replacement or a scaling operation the node may not exist at all. The primary and reader endpoints exist precisely so configuration survives topology changes.
saying these in an interview costs you the question
- Says the reader endpoint balances each individual command
- Believes writes to a reader are forwarded to the primary
- Puts individual node endpoints into application configuration
- Assumes reads from the reader endpoint are always current
- Points the whole client at the reader endpoint to spread load