skip to content

Managed Databases & Caches

AWS's managed data services — RDS, Aurora, DynamoDB, ElastiCache — and what 'managed' actually changes: failover, backups, scaling knobs, and the pricing model. The database theory itself lives in the engine trees; here I care about the AWS-shaped surface.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

23

An Aurora cluster gives you a cluster endpoint, a reader endpoint, and one endpoint per instance. What does each of those resolve to, and what goes wrong if an application points all of its traffic at the cluster endpoint?

level: juniorimportance: must knowfreq 62%

answer

  1. one name always points at the writer
  2. one name spreads readers round-robin
  3. instance names do not move on failover
  4. idle readers, busy writer
  5. DNS TTL is about five seconds

basics

~20 s

The Aurora cluster endpoint always resolves to the current writer and follows failover; the reader endpoint round-robins DNS across available readers; an instance endpoint names one fixed instance. Sending everything to the cluster endpoint puts all reads on the writer while paid-for readers sit idle.

solid answer

~50 s

The cluster endpoint is a DNS name that Aurora keeps pointed at whichever instance is currently the writer, so it survives failover — that is where every write must go. The reader endpoint resolves, round-robin, to one of the cluster's available Aurora Replicas; if the cluster has no readers it falls back to the writer. Instance endpoints address one specific instance and do **not** move on failover, so they are for diagnostics, not application traffic. Custom endpoints let you name your own subset, for example a group of large instances for reporting. If the app uses only the cluster endpoint, correctness is fine but you are paying for readers that do no work and putting all read load on the one instance that also serves writes. The fix is on the application side: separate read and write data sources, or a driver that splits them.

code

bash · 3 lines
bash
aws rds describe-db-clusters \
  --db-cluster-identifier mycluster \
  --query 'DBClusters[0].{writer:Endpoint,reader:ReaderEndpoint}'

go deeper

for a junior

Know the three names and what each points to: cluster endpoint equals current writer, reader endpoint equals some available reader, instance endpoint equals one fixed instance. Say plainly that writes only work through the cluster endpoint.

for a middle

Explain that read/write splitting is the application's job — two data sources or a cluster-aware driver — and that the reader endpoint is DNS round-robin over connections, not a query-level load balancer.

for a senior

Bring in the operational detail: short DNS TTLs and clients that cache them, dropped connections on failover and the retry contract, and why adding readers to a saturated pool changes nothing until connections are recycled.

for a principal

Own the connection topology as a design decision — where routing lives (driver, proxy, or application), how it degrades when a reader is unhealthy, and how you would prove the failover path works rather than assuming it.

## The four endpoint kinds An Aurora cluster is a set of instances over one shared storage volume, and AWS gives you several DNS names into it. Which one you connect to determines both where your query runs and what happens to that connection during a failover. **Cluster endpoint (writer endpoint).** One per cluster, of the form `mycluster.cluster-<hash>.<region>.rds.amazonaws.com`. Aurora keeps this record pointed at the instance that is currently the writer. When a failover promotes a reader, the record is updated to the new writer. Every `INSERT`, `UPDATE`, `DELETE`, and DDL statement must go here — a reader will reject writes. **Reader endpoint.** One per cluster, `mycluster.cluster-ro-<hash>.<region>.rds.amazonaws.com`. It resolves to one of the cluster's available readers, chosen round-robin at DNS resolution time. If the cluster currently has no reader instances, it resolves to the writer, which is a useful safety net and also a common source of confusion when someone tests read routing on a single-instance cluster and sees writes succeed. **Instance endpoints.** One per instance, naming that instance and nothing else. They do not follow failover: if you hard-code the instance that happens to be the writer today and it is demoted tomorrow, your writes start failing. Their legitimate uses are diagnosis ("is *this* reader lagging?") and workloads that must be pinned to a specific machine. **Custom endpoints.** You define a named group of instances — say the three `db.r6g.4xlarge` instances you keep for analysts — and Aurora load-balances across that group. Useful when your readers are deliberately not homogeneous. ## What actually goes wrong with cluster-endpoint-only Nothing breaks. That is why the mistake survives to production. Reads are correct, the app works, and the failure is economic and operational: - The writer serves every query. Its buffer cache is shared between the write working set and whatever reporting queries the app runs, so both get slower. - Any reader you added for scale is idle but billed by the hour. - A slow read query can now consume writer CPU and, in the worst case, contribute to the writer becoming the incident. Aurora cannot fix this for you at the endpoint layer, because it cannot tell from a connection whether the statements arriving on it will be reads. Read/write splitting is an application-side or driver-side decision: two connection pools bound to the two endpoints (a Spring `@Transactional(readOnly = true)` routing data source, or Rails' `connects_to` reading role), or a driver that understands the cluster — the AWS Advanced JDBC Driver, for example, supports read/write splitting and faster failover-aware reconnection than plain DNS. ## The DNS gotchas that follow Aurora's endpoint records carry a short TTL — about five seconds — precisely so clients notice a failover quickly. Two things follow: 1. **A JVM or resolver that caches DNS forever defeats it.** If your runtime caches the resolved address indefinitely, the cluster endpoint may still be aimed at the old writer long after failover, and you will see connection errors or `read-only` write failures. Make sure DNS TTLs are honoured. 2. **The reader endpoint balances connections, not queries, and only at resolution time.** A pooled application resolves the reader endpoint when it opens each connection and then keeps that connection for hours. Add a reader to a cluster whose pools are already full and it receives nothing until connections churn. That is the classic "I scaled out and nothing got faster" report. If you need rebalancing, you have to recycle connections — many pools support a maximum connection lifetime for exactly this. ```bash aws rds describe-db-clusters \ --db-cluster-identifier mycluster \ --query 'DBClusters[0].{writer:Endpoint,reader:ReaderEndpoint}' ``` ## During a failover Existing connections to the old writer are dropped — the endpoint moving does not migrate a live TCP session. The application must reconnect, and it must be able to retry the in-flight transaction or surface a clean error. Connections to the reader endpoint also break when the reader they landed on is the one being promoted. Treat "connection dropped, reconnect and retry idempotent work" as the contract, and test it: you can trigger a controlled failover from the console or the API rather than waiting for a real one.

  • You add two readers to an Aurora cluster and the reader endpoint's traffic barely shifts. Why?
    The reader endpoint balances at DNS resolution time, which happens when a connection is opened. A pool that already holds its full complement of long-lived connections never resolves again, so it never discovers the new readers. Recycle connections — a maximum connection lifetime in the pool — or use a driver that is cluster-aware.
  • When is a custom endpoint worth defining instead of just using the reader endpoint?
    When your readers are not interchangeable. If you keep two small readers for the app and two large ones for analytics, the reader endpoint would spray analyst queries onto the small instances. A custom endpoint names the group you want, and Aurora balances only within it.
  • An application hard-codes an instance endpoint for writes and starts failing after a failover. What is happening?
    Instance endpoints name one instance permanently. After promotion, the instance it names is a reader, so every write is rejected as read-only rather than being routed to the new writer. The application must use the cluster endpoint, which Aurora re-points automatically.

saying these in an interview costs you the question

  • Thinks the reader endpoint automatically splits reads out of one connection
  • Believes the reader endpoint balances per query rather than per connection
  • Uses an instance endpoint for application writes
  • Assumes existing connections survive an Aurora failover
  • Claims Aurora routes writes sent to a reader to the writer for you

context

open as a page

In DynamoDB, what is the difference between an eventually consistent read and a strongly consistent read, and how does that choice change the read capacity a request consumes?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Eventually consistent reads, the DynamoDB default, may return a slightly stale copy of an item and cost half a read unit per 4 KB. Strongly consistent reads always reflect the latest acknowledged write and cost a full read unit per 4 KB.

open as a page

Amazon ElastiCache lets you run Valkey, Redis OSS, or Memcached as the cache engine. What does each option give you operationally, and for which workload would you actually choose Memcached?

level: juniorimportance: must knowfreq 72%

basics

~10 s

Valkey and Redis OSS on ElastiCache add replicas, automatic failover, snapshots and authentication; Memcached has none of those but is multi-threaded and simple. Choose Memcached only for a plain, disposable, horizontally sharded key-value cache.

open as a page

Amazon Aurora replicas typically show replica lag in the tens of milliseconds, while a standard RDS MySQL read replica can fall seconds or minutes behind. What is structurally different about how an Aurora cluster stores its data and feeds its replicas?

level: middleimportance: must knowfreq 70%

basics

~20 s

Aurora separates compute from storage: every instance in a cluster attaches to one shared distributed volume that keeps six copies across three Availability Zones. Replicas replay nothing — they read the pages the writer already wrote, so lag stays in milliseconds.

open as a page

For a DynamoDB table, how do you choose between on-demand capacity and provisioned capacity with auto scaling, and what does each mode do when traffic spikes suddenly?

level: middleimportance: must knowfreq 70%

basics

~20 s

On-demand bills per request and absorbs spikes with no configuration, at a higher per-request price. Provisioned bills for reserved throughput and is cheaper when load is steady and well utilised, but auto scaling reacts in minutes, so sharp spikes throttle.

open as a page

In ElastiCache for Valkey or Redis OSS, what is the difference between a replication group with cluster mode disabled and one with cluster mode enabled, and what does that choice force on the application's client library?

level: middleimportance: must knowfreq 58%

basics

~20 s

Cluster mode disabled means one shard holding the whole keyspace, scaled up by node size. Cluster mode enabled splits the keyspace across many shards, each with its own primary, and requires a cluster-aware client that follows slot redirections.

open as a page

On Amazon RDS, what is the difference between automated backups and manual DB snapshots, and what actually happens when you run a point-in-time restore?

level: middleimportance: must knowfreq 60%

basics

~20 s

Automated backups are RDS-managed daily snapshots plus transaction logs, kept for a retention period of up to 35 days and deleted with the instance. Manual snapshots persist until you delete them. A restore never overwrites the source — it creates a new DB instance with a new endpoint.

open as a page

In Amazon RDS, how does a Multi-AZ deployment differ from a read replica, and which of the two would you add to relieve a database whose CPU is saturated by reporting queries?

level: middleimportance: must knowfreq 78%

basics

~20 s

Multi-AZ keeps a standby copy in another Availability Zone purely for failover, and that standby serves no traffic. A read replica is an asynchronous, separately addressable copy you can query. Reporting load needs a replica; Multi-AZ buys availability, not read capacity.

open as a page

You need to change an engine setting such as PostgreSQL's log_min_duration_statement on an Amazon RDS instance, but RDS gives you no shell and no true superuser. How do you change it, and why might your change not take effect right away?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Engine settings on RDS live in a DB parameter group attached to the instance. You cannot edit the default group, so you create a custom one, set the value, and attach it. Dynamic parameters apply right away; static ones need an instance reboot.

open as a page

What is a DynamoDB Stream, what does the StreamViewType setting control, and what ordering and retention guarantees does the stream give a consumer?

level: middleimportance: should knowfreq 56%

basics

~20 s

A DynamoDB Stream is an ordered, 24-hour log of every item-level change in a table. StreamViewType chooses whether each record carries keys only, the new image, the old image, or both. Ordering is guaranteed per item, not table-wide.

open as a page

An ElastiCache for Valkey replication group with cluster mode disabled exposes a primary endpoint and a reader endpoint. What does each one resolve to, and what happens if the application sends its writes to the reader endpoint?

level: middleimportance: should knowfreq 50%

basics

~20 s

The primary endpoint is a DNS name that always tracks the current primary, including after a failover. The reader endpoint resolves across the read replicas. Writes sent to a reader are rejected by the engine with a read-only error.

open as a page

A team wants to move an internal Aurora PostgreSQL cluster with spiky, mostly-idle traffic onto Aurora Serverless v2. What is an ACU, how does Serverless v2 change capacity at runtime, and when would you tell them to stay on provisioned instances?

level: seniorimportance: should knowfreq 45%

basics

~20 s

An Aurora Capacity Unit is roughly 2 GiB of memory plus matching CPU and network. Serverless v2 scales an instance in place between a minimum and maximum ACU setting without dropping connections. It suits spiky or idle workloads; steadily busy clusters cost more than provisioned.

open as a page

A DynamoDB table provisioned at 10,000 WCU is consuming about 2,000 WCU on average, yet writes are being throttled. What is going on, how do you confirm it, and what are your options?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Table capacity is spread across partitions, so a skewed workload can exhaust one partition's share while the table looks idle. Confirm with throttle metrics and Contributor Insights to find the hot key, then relieve it with caching, write sharding or a key change.

open as a page

An ElastiCache for Valkey replication group is Multi-AZ with automatic failover enabled, and its primary node fails. Walk through what ElastiCache does, what the application sees, and what you must have configured beforehand for the recovery to be clean.

level: seniorimportance: should knowfreq 46%

basics

~20 s

ElastiCache detects the failure, promotes a replica, and repoints the primary endpoint's DNS at it. Clients see dropped connections and errors for tens of seconds, and any write acknowledged but not yet replicated is lost, because replication is asynchronous.

open as a page

Your production Amazon RDS PostgreSQL database runs Multi-AZ in one Region. An interviewer asks how you would survive the loss of that entire Region. What AWS mechanisms do you reach for, and what does each cost you in recovery point and recovery time?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Multi-AZ never leaves its Region, so regional survival needs a cross-Region mechanism: a cross-Region read replica you promote manually, cross-Region automated backup replication, or copied snapshots. Replicas recover fastest with the smallest data loss; snapshot copies are cheapest and slowest.

open as a page

A Lambda function backed by an Amazon RDS database starts failing with "too many connections" whenever traffic spikes. What is RDS Proxy, and how does putting it in front of the database fix this?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Every concurrent Lambda execution opens its own database connection, so a traffic spike can exhaust the instance's connection limit. RDS Proxy is a managed endpoint inside your VPC that keeps a warm pool of database connections and multiplexes many short-lived client connections onto them.

open as a page

You are asked to give an Aurora-backed service a second AWS Region for disaster recovery and low-latency local reads. How does Aurora Global Database replicate between Regions, what recovery point and recovery time should you expect, and what does it force the application to handle?

level: principalimportance: should knowfreq 38%

basics

~20 s

Aurora Global Database replicates at the storage layer to secondary Regions with typical lag under a second, so cross-Region reads are local and the writer pays almost no CPU for it. There is still one writer Region, and the application must handle stale reads and a Region-level promotion.

open as a page

A team wants active-active multi-Region writes for a DynamoDB-backed service and proposes global tables. What do global tables actually give them, what do they not, and how should that shape the design?

level: principalimportance: should knowfreq 36%

basics

~20 s

Global tables replicate a DynamoDB table across Regions with writes accepted in every replica, converging asynchronously with last-writer-wins conflict resolution. They do not give cross-Region transactions, cross-Region read-after-write, or a backup — the application must be designed to tolerate convergence.

open as a page

You are choosing between ElastiCache Serverless and a node-based ElastiCache cluster you size yourself for a new service. How do you make that call, and what do you give up either way?

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide on workload shape and control. Serverless removes node sizing, shard layout and capacity planning, and bills for data stored plus processing consumed — good for spiky or unknown traffic. Node-based clusters cost less at steady high load and keep node-level tuning.

open as a page

When does putting Amazon DynamoDB Accelerator (DAX) in front of a DynamoDB table actually help, and which reads does it not accelerate?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

DAX is a write-through, in-VPC cache cluster for DynamoDB that turns repeated reads of the same items into microsecond responses. It helps read-heavy workloads with strong key reuse, and does nothing for strongly consistent reads, writes, or read patterns with poor locality.

open as a page

An ElastiCache for Valkey cluster holds session data for a web application. What mechanisms does ElastiCache give you to control who can connect and to protect the traffic, and which layer does an IAM policy actually govern?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

Four layers: the cluster is VPC-only behind security groups, encryption in transit protects the wire, an AUTH token or RBAC users authenticate clients, and RBAC access strings limit commands and key patterns. IAM policies govern the management API, not the data commands.

open as a page

A bad migration script corrupted data in an Aurora MySQL cluster twenty minutes ago. What does Aurora Backtrack do that restoring a snapshot to a new cluster does not, and what must have been true beforehand for Backtrack to be available at all?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Aurora Backtrack rewinds the existing Aurora MySQL cluster in place to an earlier timestamp in minutes, keeping the same endpoints, instead of restoring into a new cluster. It only works if backtracking was enabled with a target window when the cluster was created.

open as a page

Amazon RDS offers both a Multi-AZ DB instance deployment and a Multi-AZ DB cluster deployment. What is the difference between them, and what would make you pick the cluster?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A Multi-AZ DB instance has one hidden standby in a second Availability Zone. A Multi-AZ DB cluster has a writer plus two readable standbys across three AZs, with a reader endpoint and faster failover. Pick the cluster when you need shorter failover and read capacity from the HA copies.

open as a page