skip to content

Block & File Storage

The storage I attach to an instance instead of calling over an API: EBS block volumes, EFS/FSx shared file systems, and the ephemeral instance store. Interviews turn on picking the right one and on knowing which of them survives a stop or a host failure.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

16

Several EC2 instances behind a load balancer must read and write the same set of files using ordinary file-system calls. Which AWS storage service is built for that, and what does it give you that a volume attached to a single instance does not?

level: juniorimportance: must knowfreq 68%

answer

  1. many hosts, one mutable tree
  2. NFS share, not a block device
  3. an interface in each Availability Zone
  4. no size to provision, pay per GB stored
  5. costs latency and per-GB price

basics

~20 s

Amazon EFS is the AWS shared file system: an NFS file share that many EC2 instances, containers and Lambda functions mount read-write at the same time, across Availability Zones, with capacity that grows and shrinks automatically.

solid answer

~40 s

Amazon EFS is a managed NFS (v4.1) file system. Every client mounts the same namespace at the same time and sees the same POSIX tree — files, directories, permissions, file locking — so an app that already does `open()`/`write()` needs no code change. Capacity is elastic: you provision nothing, you are billed for what is stored, and it shrinks when you delete files. Clients reach it through a **mount target** — an elastic network interface EFS places in a subnet in each Availability Zone — so instances in different AZs share one file system, which a block volume attached to one instance in one AZ cannot do. The tradeoffs are real: per-operation latency is much higher than a local or network block device, and the per-GB price is well above object storage.

go deeper

for a junior

Be ready to name Amazon EFS the moment a question says several instances need the same files, and to say plainly that it is an NFS share mounted at a path, with capacity that grows on its own.

for a middle

Explain the mechanics: mount targets are network interfaces in each Availability Zone, traffic is NFS over TCP 2049 inside the VPC, and the file system is billed per GB stored rather than provisioned.

for a senior

Show judgment about when not to use it. Quantify the cost: higher per-operation latency than block storage and several times the per-GB price of object storage, which makes metadata-heavy workloads and bulk media the wrong fit.

for a principal

Own the architectural stance. A shared mutable filesystem is a coordination point that couples every instance to one blast radius; argue for pushing state into S3 or a database and keeping the filesystem for software that genuinely cannot be changed.

## The problem EFS solves Once a service runs on more than one instance behind a load balancer, any state written to a local disk becomes wrong: an upload handled by instance A is invisible to instance B, and it disappears when the Auto Scaling group replaces A. The three AWS answers are object storage (S3), block storage attached to one instance, and a shared file system. Amazon EFS is the shared file system. ## What EFS actually is EFS is a managed implementation of NFS version 4.1 (4.0 is also accepted). Concretely: - **One namespace, many writers.** Every client mounting the file system sees the same directory tree at the same time. Reads and writes are visible to other clients, and NFSv4 byte-range locking works, so cooperating processes on different hosts can coordinate. - **POSIX semantics.** Files have owners, groups, mode bits, symlinks, timestamps and directories. Software written for a local filesystem — a CMS writing `wp-content/uploads`, a Git server, a training job checkpointing to a path — runs unmodified. - **Elastic capacity.** There is no size to choose. The file system grows as you write and shrinks as you delete, and you pay per GB-month actually stored. That is the single biggest operational difference from provisioning a fixed-size volume you must monitor and resize. - **Regional by default.** A standard EFS file system stores data redundantly across multiple Availability Zones in the Region. (A cheaper `One Zone` variant keeps data in a single AZ and accepts the corresponding risk.) ## How clients reach it EFS lives inside your VPC. For each Availability Zone you create a **mount target**: an elastic network interface with a private IP in a subnet of that AZ, protected by a security group. Clients mount the file system's DNS name, and NFS traffic flows over TCP port 2049 to that interface. Because the network path is a normal VPC path, there is no internet gateway or NAT involved. A typical mount using the `amazon-efs-utils` package looks like: ```bash sudo mount -t efs -o tls fs-0123456789abcdef0:/ /mnt/shared ``` The `-o tls` option routes the NFS session through a local stunnel process so the traffic is encrypted in transit; encryption at rest is a checkbox set when the file system is created. EFS is not EC2-only. ECS and Fargate tasks can declare an EFS volume in the task definition, and a Lambda function can mount an EFS access point under `/mnt`, which is how people give a function a large model file or dependency set that will not fit in a deployment package. ## What it costs you Two things make EFS the wrong default rather than the obvious default: 1. **Latency.** Every file operation is a network round trip to a distributed service. Per-operation latency is an order of magnitude above a locally attached block device, and metadata-heavy work (`find` over millions of small files, an `npm install` into an EFS path, a build tree) feels dramatically slower than the same work on local storage. 2. **Price per GB.** EFS Standard storage costs several times object storage per GB-month. Storing bulk media on it because "the code already writes files" is a common and expensive mistake. ## When something else is the right answer - **S3** when the data is whole objects written once and read many times — user uploads, images, backups, data-lake files. It is cheaper, effectively unbounded, serves as a CloudFront origin, and lets browsers upload directly with presigned URLs. The cost of choosing it is that your code must speak the object API: no partial in-place writes, no rename, no directories. - **A block volume attached to one instance** when a single writer needs the lowest possible latency — most obviously a self-managed database's data directory. Running a relational database's data files on NFS is the classic misuse of EFS. - **FSx** when you need a file system EFS does not speak: SMB with Active Directory permissions for Windows applications, or a parallel file system for HPC-scale throughput. The interview test is simple: choose EFS when multiple hosts genuinely need concurrent POSIX access to the *same* mutable tree. If only one host writes, or the access pattern is really "fetch an object by key", EFS is the expensive answer to a question you did not have.

  • Can anything other than EC2 mount an EFS file system?
    Yes. ECS and Fargate tasks declare an EFS volume in the task definition and the platform mounts it into the container. Lambda functions can mount an EFS access point under `/mnt`, provided the function is attached to the VPC — a common way to give a function a large model or dependency set. On-premises hosts can mount it over Direct Connect or a VPN.
  • When would you still push user uploads to S3 rather than EFS?
    Almost always, if you can change the code. Objects written once and read many times fit S3's model: far cheaper per GB, effectively unlimited, a native CloudFront origin, and browsers can upload straight to it with presigned URLs, keeping the bytes off your servers. EFS earns its price only when software genuinely requires POSIX paths and in-place mutation.
  • What does EFS One Zone change?
    It stores the data in a single Availability Zone instead of replicating across several, at a substantially lower per-GB price. You keep elasticity and the shared-file-system semantics but take an availability and durability hit: an AZ failure takes the file system offline. It suits reproducible data — scratch space, caches, dev environments — not the system of record.

saying these in an interview costs you the question

  • Calling EFS a cheaper or shared version of a block volume
  • Putting a relational database's data directory on EFS
  • Believing you must provision an EFS file system's size up front
  • Assuming Windows clients can mount EFS over SMB
  • Reaching for EFS when a single instance is the only writer

context

open as a page

In AWS, what happens to data on an EC2 instance store volume, compared with data on an attached EBS volume, when the instance is rebooted, stopped and started again, or terminated?

level: juniorimportance: must knowfreq 78%

basics

~20 s

EC2 instance store data survives a reboot but is lost on stop, hibernate, or terminate, because it lives on disks physically inside the host. EBS volumes are network-attached and persist independently of the instance lifecycle.

open as a page

In Amazon EBS, how do gp3 and gp2 volumes differ in how their IOPS and throughput are determined, and why is gp3 usually the better default?

level: middleimportance: must knowfreq 75%

basics

~20 s

gp2 derives performance from size — 3 IOPS per GiB with a burst credit bucket — so you buy capacity to get speed. gp3 ships a fixed baseline at any size and lets you provision IOPS and throughput independently, usually at a lower price per GiB.

open as a page

Amazon EBS snapshots are described as incremental. What does that mean for what is stored and billed, and if you delete a snapshot from the middle of a chain, what happens to the snapshots taken after it?

level: middleimportance: must knowfreq 62%

basics

~20 s

Each EBS snapshot stores only the blocks that changed since the previous snapshot, and you are billed for the unique blocks retained. Deleting a middle snapshot is safe: AWS removes only blocks no other snapshot needs, and every remaining snapshot still restores a complete volume.

open as a page

EC2 instances in two of your three Availability Zones mount an EFS file system fine, but instances in the third hang on the mount command and eventually time out. What are the AWS-specific causes, and how do you diagnose it?

level: seniorimportance: must knowfreq 45%

basics

~20 s

EFS is reached through a per-Availability-Zone mount target — a network interface with its own security group. A hang almost always means the third AZ has no mount target, or the mount target's security group does not allow inbound TCP 2049 from the client.

open as a page

An Amazon EBS volume is in us-east-1a and you need the same data on an instance in us-east-1b, and later in eu-west-1. What is the scope of an EBS volume compared with an EBS snapshot, and how do you actually move the data?

level: juniorimportance: should knowfreq 55%

basics

~20 s

An EBS volume is locked to a single Availability Zone; an EBS snapshot is Region-scoped. To move data you snapshot the volume and create a new volume from that snapshot in the target AZ, and copy the snapshot to reach another Region.

open as a page

A production EC2 instance's Amazon EBS data volume is nearly full. How do you grow it without downtime, and what are the limits of an EBS volume modification?

level: middleimportance: should knowfreq 52%

basics

~20 s

Call ModifyVolume with the new size while the volume stays attached, wait for it to leave the optimizing state, then extend the partition and filesystem inside the operating system. EBS volumes can only grow, never shrink, and a cooldown applies before the next modification.

open as a page

An AWS team needs managed shared file systems for two workloads that EFS does not fit: a Windows application requiring SMB shares with Active Directory permissions, and a training job that must stream terabytes out of S3 at very high throughput. Which AWS file services fit each, and what distinguishes the FSx family members?

level: middleimportance: should knowfreq 40%

basics

~20 s

Amazon FSx for Windows File Server serves SMB shares with Active Directory and NTFS permissions; Amazon FSx for Lustre is the high-throughput parallel file system that links directly to an S3 bucket. FSx also offers NetApp ONTAP and OpenZFS file systems.

open as a page

Amazon EFS offers Bursting, Provisioned and Elastic throughput modes. What does each one do, and how would you choose between them for a given workload?

level: middleimportance: should knowfreq 50%

basics

~20 s

Bursting ties throughput to how much data is stored and accrues burst credits; Provisioned sets a fixed rate you pay for regardless of size; Elastic scales automatically and bills per GB read and written. Elastic is the default for unpredictable workloads.

open as a page

Why can you not create an EBS-style snapshot of an EC2 instance store volume, and how do you protect data that lives on one?

level: middleimportance: should knowfreq 48%

basics

~20 s

An EC2 instance store volume is a raw disk in the host, not an AWS resource with a volume ID, so no snapshot, detach, or resize API applies to it. Protection has to happen above it: copy to S3 or EBS, or replicate at the application level.

open as a page

You must back up the Amazon EBS volumes of a busy database instance without stopping it. What consistency does a snapshot of a running instance give you by default, and what would you do to get more?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A snapshot of a running instance is crash-consistent: it captures the volume as if power were cut, so recovery relies on the database's own journal replay. For application consistency you flush and quiesce writes briefly before the API call, and use CreateSnapshots to capture multi-volume instances at one point in time.

open as a page

An Amazon EFS lifecycle policy is supposed to move cold files into the Infrequent Access storage class to cut cost. How does EFS decide which files to move, and what commonly defeats the saving or makes the bill worse?

level: seniorimportance: should knowfreq 32%

basics

~20 s

EFS lifecycle management moves a file to Infrequent Access or Archive after a configured number of days without access to its contents. Infrequent Access charges per GB retrieved and adds latency, so jobs that read every file destroy the saving.

open as a page

For an EC2 workload, how do you decide whether data belongs on an instance store volume rather than an EBS volume, and what must be true of the application before you choose instance store?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Choose EC2 instance store only when the data is reproducible — scratch, spill, cache, or a replica of a shard held elsewhere — and the workload is limited by storage latency or throughput. Everything that is a source of truth belongs on EBS or a managed service.

open as a page

An EC2 instance's Amazon EBS volume is provisioned for 16,000 IOPS, but a benchmark tops out near 6,000 with high latency. What AWS-side causes would you investigate, and how does that shape your volume-type choice?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Suspect the instance before the volume: every EC2 instance type has its own EBS bandwidth and IOPS ceiling, and smaller sizes only burst to it. Then check I/O size against the throughput cap, queue depth, and whether the volume was restored from a snapshot and is still lazily loading blocks.

open as a page

Several applications must share one Amazon EFS file system without being able to read each other's directories or act as root. What does an EFS access point enforce, and how does IAM authorization fit alongside it?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

An EFS access point pins a client to a root directory and a POSIX user and group, so it cannot traverse above that directory or choose its own identity. A file system policy then grants IAM principals actions such as elasticfilesystem:ClientMount and ClientWrite.

open as a page

You operate a fleet of EC2 instances whose local NVMe instance store volumes hold each node's working data. AWS sends a scheduled retirement notice for one of them. What is the correct operational response, and what does running on ephemeral storage change about how you patch and deploy that fleet?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Treat the node as already lost: replace it rather than stopping and starting it, and let the application rebuild its local data from replicas or from durable storage. On ephemeral storage every maintenance action becomes a node replacement, so rebuild time is your real operational constraint.

open as a page