An Aurora cluster gives you a cluster endpoint, a reader endpoint, and one endpoint per instance. What does each of those resolve to, and what goes wrong if an application points all of its traffic at the cluster endpoint?
answer
- one name always points at the writer
- one name spreads readers round-robin
- instance names do not move on failover
- idle readers, busy writer
- DNS TTL is about five seconds
basics
~20 sThe Aurora cluster endpoint always resolves to the current writer and follows failover; the reader endpoint round-robins DNS across available readers; an instance endpoint names one fixed instance. Sending everything to the cluster endpoint puts all reads on the writer while paid-for readers sit idle.
solid answer
~50 sThe cluster endpoint is a DNS name that Aurora keeps pointed at whichever instance is currently the writer, so it survives failover — that is where every write must go. The reader endpoint resolves, round-robin, to one of the cluster's available Aurora Replicas; if the cluster has no readers it falls back to the writer. Instance endpoints address one specific instance and do **not** move on failover, so they are for diagnostics, not application traffic. Custom endpoints let you name your own subset, for example a group of large instances for reporting. If the app uses only the cluster endpoint, correctness is fine but you are paying for readers that do no work and putting all read load on the one instance that also serves writes. The fix is on the application side: separate read and write data sources, or a driver that splits them.
code
bash · 3 linesaws rds describe-db-clusters \
--db-cluster-identifier mycluster \
--query 'DBClusters[0].{writer:Endpoint,reader:ReaderEndpoint}'go deeper
Know the three names and what each points to: cluster endpoint equals current writer, reader endpoint equals some available reader, instance endpoint equals one fixed instance. Say plainly that writes only work through the cluster endpoint.
Explain that read/write splitting is the application's job — two data sources or a cluster-aware driver — and that the reader endpoint is DNS round-robin over connections, not a query-level load balancer.
Bring in the operational detail: short DNS TTLs and clients that cache them, dropped connections on failover and the retry contract, and why adding readers to a saturated pool changes nothing until connections are recycled.
Own the connection topology as a design decision — where routing lives (driver, proxy, or application), how it degrades when a reader is unhealthy, and how you would prove the failover path works rather than assuming it.
## The four endpoint kinds An Aurora cluster is a set of instances over one shared storage volume, and AWS gives you several DNS names into it. Which one you connect to determines both where your query runs and what happens to that connection during a failover. **Cluster endpoint (writer endpoint).** One per cluster, of the form `mycluster.cluster-<hash>.<region>.rds.amazonaws.com`. Aurora keeps this record pointed at the instance that is currently the writer. When a failover promotes a reader, the record is updated to the new writer. Every `INSERT`, `UPDATE`, `DELETE`, and DDL statement must go here — a reader will reject writes. **Reader endpoint.** One per cluster, `mycluster.cluster-ro-<hash>.<region>.rds.amazonaws.com`. It resolves to one of the cluster's available readers, chosen round-robin at DNS resolution time. If the cluster currently has no reader instances, it resolves to the writer, which is a useful safety net and also a common source of confusion when someone tests read routing on a single-instance cluster and sees writes succeed. **Instance endpoints.** One per instance, naming that instance and nothing else. They do not follow failover: if you hard-code the instance that happens to be the writer today and it is demoted tomorrow, your writes start failing. Their legitimate uses are diagnosis ("is *this* reader lagging?") and workloads that must be pinned to a specific machine. **Custom endpoints.** You define a named group of instances — say the three `db.r6g.4xlarge` instances you keep for analysts — and Aurora load-balances across that group. Useful when your readers are deliberately not homogeneous. ## What actually goes wrong with cluster-endpoint-only Nothing breaks. That is why the mistake survives to production. Reads are correct, the app works, and the failure is economic and operational: - The writer serves every query. Its buffer cache is shared between the write working set and whatever reporting queries the app runs, so both get slower. - Any reader you added for scale is idle but billed by the hour. - A slow read query can now consume writer CPU and, in the worst case, contribute to the writer becoming the incident. Aurora cannot fix this for you at the endpoint layer, because it cannot tell from a connection whether the statements arriving on it will be reads. Read/write splitting is an application-side or driver-side decision: two connection pools bound to the two endpoints (a Spring `@Transactional(readOnly = true)` routing data source, or Rails' `connects_to` reading role), or a driver that understands the cluster — the AWS Advanced JDBC Driver, for example, supports read/write splitting and faster failover-aware reconnection than plain DNS. ## The DNS gotchas that follow Aurora's endpoint records carry a short TTL — about five seconds — precisely so clients notice a failover quickly. Two things follow: 1. **A JVM or resolver that caches DNS forever defeats it.** If your runtime caches the resolved address indefinitely, the cluster endpoint may still be aimed at the old writer long after failover, and you will see connection errors or `read-only` write failures. Make sure DNS TTLs are honoured. 2. **The reader endpoint balances connections, not queries, and only at resolution time.** A pooled application resolves the reader endpoint when it opens each connection and then keeps that connection for hours. Add a reader to a cluster whose pools are already full and it receives nothing until connections churn. That is the classic "I scaled out and nothing got faster" report. If you need rebalancing, you have to recycle connections — many pools support a maximum connection lifetime for exactly this. ```bash aws rds describe-db-clusters \ --db-cluster-identifier mycluster \ --query 'DBClusters[0].{writer:Endpoint,reader:ReaderEndpoint}' ``` ## During a failover Existing connections to the old writer are dropped — the endpoint moving does not migrate a live TCP session. The application must reconnect, and it must be able to retry the in-flight transaction or surface a clean error. Connections to the reader endpoint also break when the reader they landed on is the one being promoted. Treat "connection dropped, reconnect and retry idempotent work" as the contract, and test it: you can trigger a controlled failover from the console or the API rather than waiting for a real one.
- You add two readers to an Aurora cluster and the reader endpoint's traffic barely shifts. Why?The reader endpoint balances at DNS resolution time, which happens when a connection is opened. A pool that already holds its full complement of long-lived connections never resolves again, so it never discovers the new readers. Recycle connections — a maximum connection lifetime in the pool — or use a driver that is cluster-aware.
- When is a custom endpoint worth defining instead of just using the reader endpoint?When your readers are not interchangeable. If you keep two small readers for the app and two large ones for analytics, the reader endpoint would spray analyst queries onto the small instances. A custom endpoint names the group you want, and Aurora balances only within it.
- An application hard-codes an instance endpoint for writes and starts failing after a failover. What is happening?Instance endpoints name one instance permanently. After promotion, the instance it names is a reader, so every write is rejected as read-only rather than being routed to the new writer. The application must use the cluster endpoint, which Aurora re-points automatically.
saying these in an interview costs you the question
- Thinks the reader endpoint automatically splits reads out of one connection
- Believes the reader endpoint balances per query rather than per connection
- Uses an instance endpoint for application writes
- Assumes existing connections survive an Aurora failover
- Claims Aurora routes writes sent to a reader to the writer for you