skip to content

RDS

RDS runs Postgres, MySQL, MariaDB, Oracle, or SQL Server for me, turning failover, backups, and read replicas into configuration rather than code. Interviewers check that I know Multi-AZ buys availability, not read scaling.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

On Amazon RDS, what is the difference between automated backups and manual DB snapshots, and what actually happens when you run a point-in-time restore?

level: middleimportance: must knowfreq 60%

answer

  1. one expires, one does not
  2. the instance's deletion takes one with it
  3. logs plus a daily copy
  4. restore never happens in place
  5. new instance, new endpoint

basics

~20 s

Automated backups are RDS-managed daily snapshots plus transaction logs, kept for a retention period of up to 35 days and deleted with the instance. Manual snapshots persist until you delete them. A restore never overwrites the source — it creates a new DB instance with a new endpoint.

solid answer

~40 s

RDS keeps two distinct things. **Automated backups** are taken by the service on a schedule you set, combined with transaction logs uploaded continuously, which together let you restore to any point inside the retention window — configurable from 0 (disabled) up to 35 days. **Manual snapshots** are ones you take yourself; they are full point-in-time copies with no expiry, and they survive after the instance is deleted, which automated backups do not by default. The operational detail that catches people out is that `restore-db-instance-to-point-in-time` and `restore-db-instance-from-db-snapshot` both **create a brand-new DB instance** — RDS never restores in place. The new instance has a new endpoint and, unless you specify otherwise, default settings for security groups and parameter groups, so a real recovery involves repointing the application and reapplying configuration.

go deeper

for a junior

Know that RDS takes automated backups on a retention schedule, that snapshots you take yourself do not expire, and that restoring produces a new database rather than rewinding the old one.

for a middle

Explain the daily-snapshot-plus-transaction-log mechanism behind point-in-time restore, the 0–35 day retention range, and why the restored instance arrives with a new endpoint and default configuration.

for a senior

Show that your runbook accounts for restore duration, endpoint repointing through a DNS alias, reapplying security and parameter groups, and the KMS constraints on cross-Region and cross-account snapshot copies.

for a principal

Own the recovery policy across the estate: which databases justify longer retention or cross-Region backup replication, how deletion protection and mandatory final snapshots are enforced, and how often restores are actually rehearsed.

## Two mechanisms with different lifecycles RDS backup questions are asked because the two mechanisms are easy to confuse and because the difference shows up at exactly the wrong moment — after someone deletes a database. **Automated backups** are a service feature you enable by setting a *backup retention period*. RDS takes a daily storage-level snapshot during the backup window you configure and, separately, uploads database transaction logs continuously. Retention is configurable from 0 to 35 days; setting it to 0 disables automated backups entirely, which also disables point-in-time restore. The daily snapshot and the log stream together are what make an arbitrary restore point possible: RDS restores the most recent snapshot before your target time and replays logs forward to that instant. **Manual DB snapshots** are ones you create with `create-db-snapshot`. They are full copies, they never expire on their own, and they are yours to delete. Because they outlive the source instance, they are the mechanism for anything you must keep beyond the retention window — a pre-migration safety copy, a monthly archival point, a hand-off to another account. ## The deletion trap When you delete a DB instance, its automated backups are removed with it on the normal path — they are tied to the instance's lifecycle. Manual snapshots are not; they persist. Two RDS features soften this: the delete operation offers to take a **final snapshot** (a manual snapshot created at deletion time), and RDS supports **retained automated backups**, which keep the automated backup set for the remainder of its retention period after the instance is gone. Neither is a substitute for a deliberate snapshot policy for the copies you truly need to keep. ## What a restore actually does This is the part worth saying out loud, because it changes runbooks: ```bash aws rds restore-db-instance-to-point-in-time \ --source-db-instance-identifier prod-db \ --target-db-instance-identifier prod-db-recovered \ --restore-time 2026-08-20T14:30:00Z ``` RDS **creates a new DB instance**. It does not roll the existing database back, and there is no in-place option. Consequences to plan for: - **The endpoint changes.** The recovered instance has a new DNS name, so recovery includes repointing applications — via configuration, a CNAME you control, or a Route 53 record. Teams that keep an alias in front of the RDS endpoint recover faster than teams that hardcode it. - **Configuration does not follow automatically.** Unless you pass them, the new instance gets default security groups and parameter groups. A restored database that nothing can reach, because it landed in the default security group, is a common and avoidable incident. - **Restore time scales with data volume.** Provisioning and hydrating a large instance is not instantaneous, and log replay to a distant point adds to it. That is why restore time belongs in a tested runbook rather than an assumption. - **You can restore to a different shape.** The new instance may take a different instance class, storage type, or Multi-AZ setting, which makes restores useful for more than disasters — for example spinning up a production-sized copy for a schema-change rehearsal. The latest time you can target is not "now": RDS exposes a **latest restorable time**, typically a few minutes behind the present, reflecting how recently transaction logs were uploaded. Any window you promise has to account for that gap. ## Copies, Regions and encryption Snapshots can be copied — `copy-db-snapshot` — to another Region or shared with another account, and automated backups can be replicated to a second Region so that point-in-time restore is possible there too. Encryption is where copies bite: a snapshot of an encrypted instance is encrypted with a KMS key, and a cross-Region copy must be re-encrypted with a key in the destination Region, so the copy operation needs a destination key and permission to use it. Sharing an encrypted snapshot with another account requires a customer-managed key, because AWS-managed keys cannot be shared. Discovering this during an incident is much worse than discovering it during a drill. ```bash aws rds copy-db-snapshot \ --source-db-snapshot-identifier arn:aws:rds:eu-west-1:123456789012:snapshot:prod-2026-08-20 \ --target-db-snapshot-identifier prod-2026-08-20-dr \ --kms-key-id alias/dr-rds \ --region us-east-1 ``` ## The habits that show experience Set retention deliberately rather than accepting whatever the console offered; take a manual snapshot before every risky change and label it; keep an application-controlled DNS alias in front of the database endpoint so a restore does not require a code deploy; and rehearse a restore on a schedule, timing it, so the number in your recovery plan is measured rather than hoped for.

  • What is the practical consequence of RDS's 'latest restorable time' lagging the present?
    You cannot restore to the last few moments before an incident. Transaction logs are uploaded periodically, so the most recent restorable instant typically trails wall-clock time by a few minutes. Any recovery-point promise must include that gap, and if it is unacceptable you need a different mechanism — for example a replica held for failover — rather than a tighter backup setting.
  • An engineer deleted an RDS instance and now needs data from last week. What determines whether that is possible?
    Whether a manual snapshot or a final snapshot exists, or whether automated backups were retained. Automated backups follow the instance's lifecycle on the normal deletion path, so if nobody took a final snapshot and retention of automated backups was not in play, there may be nothing left. This is why deletion protection and a mandatory final snapshot belong in the standard configuration.
  • Why can copying an encrypted RDS snapshot to another Region fail even when the copy call itself is permitted?
    Because the copy must be encrypted with a KMS key in the destination Region, and the caller needs permission to use that key. An encrypted snapshot cannot simply carry its source-Region key across. Sharing with another account adds a further constraint: only customer-managed keys can be shared, so instances encrypted with an AWS-managed key cannot have their snapshots shared directly.

saying these in an interview costs you the question

  • Manual snapshots expire with the backup retention period
  • A point-in-time restore rolls the existing instance back in place
  • Deleting an RDS instance keeps its automated backups indefinitely
  • You can restore to the current second with no lag
  • The restored instance keeps the source's security groups automatically

context

open as a page

In Amazon RDS, how does a Multi-AZ deployment differ from a read replica, and which of the two would you add to relieve a database whose CPU is saturated by reporting queries?

level: middleimportance: must knowfreq 78%

basics

~20 s

Multi-AZ keeps a standby copy in another Availability Zone purely for failover, and that standby serves no traffic. A read replica is an asynchronous, separately addressable copy you can query. Reporting load needs a replica; Multi-AZ buys availability, not read capacity.

open as a page

You need to change an engine setting such as PostgreSQL's log_min_duration_statement on an Amazon RDS instance, but RDS gives you no shell and no true superuser. How do you change it, and why might your change not take effect right away?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Engine settings on RDS live in a DB parameter group attached to the instance. You cannot edit the default group, so you create a custom one, set the value, and attach it. Dynamic parameters apply right away; static ones need an instance reboot.

open as a page

Your production Amazon RDS PostgreSQL database runs Multi-AZ in one Region. An interviewer asks how you would survive the loss of that entire Region. What AWS mechanisms do you reach for, and what does each cost you in recovery point and recovery time?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Multi-AZ never leaves its Region, so regional survival needs a cross-Region mechanism: a cross-Region read replica you promote manually, cross-Region automated backup replication, or copied snapshots. Replicas recover fastest with the smallest data loss; snapshot copies are cheapest and slowest.

open as a page

A Lambda function backed by an Amazon RDS database starts failing with "too many connections" whenever traffic spikes. What is RDS Proxy, and how does putting it in front of the database fix this?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Every concurrent Lambda execution opens its own database connection, so a traffic spike can exhaust the instance's connection limit. RDS Proxy is a managed endpoint inside your VPC that keeps a warm pool of database connections and multiplexes many short-lived client connections onto them.

open as a page

Amazon RDS offers both a Multi-AZ DB instance deployment and a Multi-AZ DB cluster deployment. What is the difference between them, and what would make you pick the cluster?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A Multi-AZ DB instance has one hidden standby in a second Availability Zone. A Multi-AZ DB cluster has a writer plus two readable standbys across three AZs, with a reader endpoint and faster failover. Pick the cluster when you need shorter failover and read capacity from the HA copies.

open as a page