skip to content

Storage & Databases

The durable half of an AWS system: S3 with its storage classes, lifecycle rules, versioning and presigned URLs, EBS versus EFS, and the database line-up from RDS Multi-AZ and read replicas to Aurora Serverless and DynamoDB. You learn the partitioning and consistency models, because picking the wrong store is exactly the failure interviewers probe for.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

82 · 4 sections

Several EC2 instances behind a load balancer must read and write the same set of files using ordinary file-system calls. Which AWS storage service is built for that, and what does it give you that a volume attached to a single instance does not?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Amazon EFS is the AWS shared file system: an NFS file share that many EC2 instances, containers and Lambda functions mount read-write at the same time, across Availability Zones, with capacity that grows and shrinks automatically.

open as a page

In AWS, what happens to data on an EC2 instance store volume, compared with data on an attached EBS volume, when the instance is rebooted, stopped and started again, or terminated?

level: juniorimportance: must knowfreq 78%
basics
~20 s

EC2 instance store data survives a reboot but is lost on stop, hibernate, or terminate, because it lives on disks physically inside the host. EBS volumes are network-attached and persist independently of the instance lifecycle.

open as a page

In Amazon EBS, how do gp3 and gp2 volumes differ in how their IOPS and throughput are determined, and why is gp3 usually the better default?

level: middleimportance: must knowfreq 75%
basics
~20 s

gp2 derives performance from size — 3 IOPS per GiB with a burst credit bucket — so you buy capacity to get speed. gp3 ships a fixed baseline at any size and lets you provision IOPS and throughput independently, usually at a lower price per GiB.

open as a page

Amazon EBS snapshots are described as incremental. What does that mean for what is stored and billed, and if you delete a snapshot from the middle of a chain, what happens to the snapshots taken after it?

level: middleimportance: must knowfreq 62%
basics
~20 s

Each EBS snapshot stores only the blocks that changed since the previous snapshot, and you are billed for the unique blocks retained. Deleting a middle snapshot is safe: AWS removes only blocks no other snapshot needs, and every remaining snapshot still restores a complete volume.

open as a page

EC2 instances in two of your three Availability Zones mount an EFS file system fine, but instances in the third hang on the mount command and eventually time out. What are the AWS-specific causes, and how do you diagnose it?

level: seniorimportance: must knowfreq 45%
basics
~20 s

EFS is reached through a per-Availability-Zone mount target — a network interface with its own security group. A hang almost always means the third AZ has no mount target, or the mount target's security group does not allow inbound TCP 2049 from the client.

open as a page

An Aurora cluster gives you a cluster endpoint, a reader endpoint, and one endpoint per instance. What does each of those resolve to, and what goes wrong if an application points all of its traffic at the cluster endpoint?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The Aurora cluster endpoint always resolves to the current writer and follows failover; the reader endpoint round-robins DNS across available readers; an instance endpoint names one fixed instance. Sending everything to the cluster endpoint puts all reads on the writer while paid-for readers sit idle.

open as a page

In DynamoDB, what is the difference between an eventually consistent read and a strongly consistent read, and how does that choice change the read capacity a request consumes?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Eventually consistent reads, the DynamoDB default, may return a slightly stale copy of an item and cost half a read unit per 4 KB. Strongly consistent reads always reflect the latest acknowledged write and cost a full read unit per 4 KB.

open as a page

Amazon ElastiCache lets you run Valkey, Redis OSS, or Memcached as the cache engine. What does each option give you operationally, and for which workload would you actually choose Memcached?

level: juniorimportance: must knowfreq 72%
basics
~10 s

Valkey and Redis OSS on ElastiCache add replicas, automatic failover, snapshots and authentication; Memcached has none of those but is multi-threaded and simple. Choose Memcached only for a plain, disposable, horizontally sharded key-value cache.

open as a page

Amazon Aurora replicas typically show replica lag in the tens of milliseconds, while a standard RDS MySQL read replica can fall seconds or minutes behind. What is structurally different about how an Aurora cluster stores its data and feeds its replicas?

level: middleimportance: must knowfreq 70%
basics
~20 s

Aurora separates compute from storage: every instance in a cluster attaches to one shared distributed volume that keeps six copies across three Availability Zones. Replicas replay nothing — they read the pages the writer already wrote, so lag stays in milliseconds.

open as a page

For a DynamoDB table, how do you choose between on-demand capacity and provisioned capacity with auto scaling, and what does each mode do when traffic spikes suddenly?

level: middleimportance: must knowfreq 70%
basics
~20 s

On-demand bills per request and absorbs spikes with no configuration, at a higher per-request price. Provisioned bills for reserved throughput and is cheaper when load is steady and well utilised, but auto scaling reacts in minutes, so sharp spikes throttle.

open as a page

In Amazon S3, how does a bucket policy differ from an IAM identity policy attached to a user or role, and which element does a bucket policy have that an identity policy never has?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A bucket policy is a resource-based policy attached to the S3 bucket itself and must name a Principal. An IAM identity policy attaches to a user, group or role, and never names a Principal because the identity it is attached to is the principal.

open as a page

Amazon S3 refuses a single PutObject request larger than 5 GB. How does an S3 multipart upload work, and what does it give you beyond raising that ceiling?

level: juniorimportance: must knowfreq 62%
basics
~10 s

Multipart upload splits one object into parts sent independently: CreateMultipartUpload returns an UploadId, UploadPart sends each numbered piece, CompleteMultipartUpload assembles them. Beyond the 5 GB ceiling it adds parallelism, per-part retry, and resumable uploads.

open as a page

What is an Amazon S3 presigned URL, and whose permissions does the person holding one actually exercise?

level: juniorimportance: must knowfreq 78%
basics
~20 s

An S3 presigned URL is an object URL carrying a SigV4 signature and an expiry in its query string. Whoever holds it performs one specific operation on one specific object using the signing principal's permissions, with no AWS credentials of their own.

open as a page

In Amazon S3, what does a bucket's Lifecycle configuration do, and what are the two kinds of action a lifecycle rule can take?

level: juniorimportance: must knowfreq 72%
basics
~20 s

An S3 Lifecycle configuration is a set of bucket-level rules that act on objects automatically as they age. A rule takes one of two kinds of action: transition an object to a cheaper storage class, or expire (delete) it.

open as a page

In an Amazon S3 bucket with versioning enabled, what happens when a client sends a DELETE for an object key without specifying a version ID, and how do you get the object back?

level: juniorimportance: must knowfreq 68%
basics
~20 s

S3 erases nothing: it adds a delete marker as the new current version, so plain GETs return 404 while every earlier version stays stored and billed. You restore by deleting the delete marker by its own version ID.

open as a page

On AWS, the name "Glacier" shows up both as a standalone archival service with vaults and archives and as a set of Amazon S3 storage-class names. Where does archival data actually live in each case, and which of the two should a new design target?

level: middleimportance: nice to knowfreq 22%
basics
~20 s

Glacier today is a set of Amazon S3 storage classes — GLACIER_IR, GLACIER and DEEP_ARCHIVE — applied to ordinary objects in ordinary S3 buckets. The older standalone S3 Glacier service stores opaque archives in vaults under its own API and is legacy; new designs should use the S3 storage classes.

open as a page