EC2 instances in two of your three Availability Zones mount an EFS file system fine, but instances in the third hang on the mount command and eventually time out. What are the AWS-specific causes, and how do you diagnose it?
answer
- one network interface per Availability Zone
- the file share has its own firewall
- NFS speaks on a single well-known port
- a dropped packet hangs, a refused one fails fast
- widening the ASG did not widen the file system
basics
~20 sEFS is reached through a per-Availability-Zone mount target — a network interface with its own security group. A hang almost always means the third AZ has no mount target, or the mount target's security group does not allow inbound TCP 2049 from the client.
solid answer
~50 sEFS is not reachable as a regional endpoint; each Availability Zone gets a **mount target**, an elastic network interface with a private IP in one of that AZ's subnets, guarded by its own security group. Two failures produce this exact symptom. First, no mount target exists in the third AZ, so there is nothing local for the client's DNS lookup to resolve to. Second — far more common — the mount target's security group has no inbound rule allowing TCP 2049 from the client's security group. A dropped packet makes `mount` hang and time out rather than fail fast, which is the tell that a security group or NACL is silently discarding traffic rather than something rejecting the connection. Confirm with `aws efs describe-mount-targets`, check the mount target's security group, and verify the VPC has DNS resolution enabled so the file-system DNS name resolves at all.
code
bash · 8 lines# List mount targets and the AZ each one serves
aws efs describe-mount-targets --file-system-id fs-0123456789abcdef0 \
--query 'MountTargets[].{AZ:AvailabilityZoneName,Subnet:SubnetId,IP:IpAddress,State:LifeCycleState}' \
--output table
# Inspect the security groups attached to one mount target
aws efs describe-mount-target-security-groups \
--mount-target-id fsmt-0123456789abcdef0go deeper
Remember that EFS is reached through a mount target in each Availability Zone, and that NFS traffic uses TCP port 2049 — so a firewall rule is almost always involved when a mount does not work.
Explain that the mount target is an ENI with its own security group, distinct from the instance's, and that inbound 2049 must be allowed from the client. Know that DNS resolution in the VPC is a prerequisite for the mount name to work.
Work the diagnosis in order — mount target present, security group, NACL return path, DNS — and read the hang-versus-refuse distinction as evidence. Explain the latency and availability cost of a client reaching another AZ's mount target.
Make the failure structurally impossible: mount targets provisioned in every AZ compute can occupy, security groups referencing groups rather than CIDRs, and a standard that widening a compute footprint and widening the storage footprint are one change, not two.
## Why EFS is per-AZ at all An EFS file system is a Regional resource, but clients never talk to a Regional endpoint. For each Availability Zone you create a **mount target**: an elastic network interface placed in a subnet you choose, with a private IP address in that subnet's CIDR and one or more security groups attached. NFS traffic goes to that IP over TCP port 2049. That design is why the failure is AZ-shaped. Everything works in the AZs you configured, and the AZ nobody tested — often one added later when the Auto Scaling group was widened, or the third subnet added during a resilience review — has no path. ## The two causes, in order of likelihood **1. The mount target's security group blocks port 2049.** The mount target's security group is a separate object from the instance's. It must allow inbound TCP 2049 from the client — ideally by referencing the client's security group ID rather than a CIDR. This is the single most common EFS mistake, and it is easy to make because the file system may have been created with one security group and the new AZ's mount target with another (or with the VPC default group, which permits only traffic from itself). **2. There is no mount target in that AZ.** Mount targets are created one per AZ, and adding a subnet to your compute does not add one. Nothing about the file system's configuration warns you. ## The diagnostic tell: hang versus refuse The distinction between a *hang* and an immediate error is worth stating explicitly in an interview, because it separates the two candidate causes: - A security group or network ACL that **drops** the packet produces silence. The client retransmits SYNs and the mount command sits there until it times out. That is what you are seeing here. - A name that does not resolve produces a fast, explicit DNS failure instead. So a hang points at the network path — security group first, then NACL — while an instant "failed to resolve" points at DNS. ## The checks, in order ```bash # 1. Is there a mount target in the failing AZ at all? aws efs describe-mount-targets --file-system-id fs-0123456789abcdef0 \ --query 'MountTargets[].[AvailabilityZoneName,SubnetId,IpAddress,LifeCycleState]' # 2. Which security groups guard it? aws efs describe-mount-target-security-groups --mount-target-id fsmt-0123456789abcdef0 ``` Then, from the failing instance, confirm the name resolves and the port opens. If DNS resolution fails, check that the VPC has DNS support and DNS hostnames enabled — EFS mount helpers depend on the file system's DNS name resolving to the local mount target's private IP, and a VPC with DNS support switched off (or a custom DHCP option set pointing at private resolvers with no forwarder) breaks it. Mounting by the mount target's IP address is a valid emergency workaround and also a fast way to prove that DNS, not the network path, is the problem. ## The other suspects - **Network ACLs.** Unlike security groups, NACLs are stateless: allowing inbound 2049 is not enough, the return traffic on ephemeral ports must be allowed outbound as well. A hardened subnet with a restrictive NACL is a classic cause. - **The client's own egress rules.** A locked-down instance security group that permits no outbound traffic to 2049 fails the same way. - **TLS mounts without the helper.** Mounting with `-o tls` requires the `amazon-efs-utils` package, which runs a local stunnel process. On an AMI without it, `mount -t efs` is simply not a known type. - **Cross-account or shared VPC setups**, where the mount target lives in a subnet you do not control. ## The design lesson Create a mount target in **every** Availability Zone your compute can land in, and treat that as part of the file system's definition rather than a follow-up task. Beyond avoiding this outage, a client that has to reach a mount target in another AZ pays a latency penalty on every file operation and takes a dependency on a second AZ staying healthy — precisely the coupling that multi-AZ deployment was meant to remove. When an Auto Scaling group is widened to a new subnet, adding the mount target belongs in the same change.
- Why does the mount hang rather than return an error immediately?Because security groups and network ACLs drop disallowed packets silently rather than sending a reset. The client retransmits TCP SYNs to port 2049 and blocks until the mount times out. An immediate error would point elsewhere — typically DNS failing to resolve the file system's name, which fails fast and loudly.
- The mount succeeds but every file operation is noticeably slower from one AZ. What is likely happening?That AZ probably has no local mount target, so the client is traversing to a mount target in another Availability Zone. Every NFS round trip crosses the AZ boundary, adding latency to metadata-heavy work, and the workload now depends on two AZs being healthy instead of one. The fix is a mount target in the local AZ.
- How would you prevent this class of failure from recurring?Treat mount targets as part of the file system's definition, not an afterthought: whenever compute is allowed into a subnet, that AZ gets a mount target and its security group references the client security group by ID rather than a CIDR. Alarm on client-side mount failures so a new AZ surfaces immediately rather than at the next scaling event.
saying these in an interview costs you the question
- Thinking EFS has one regional endpoint reachable from anywhere in the VPC
- Only opening port 2049 on the instance's security group, not the mount target's
- Assuming a NAT gateway or internet route is needed for EFS
- Forgetting that network ACLs are stateless and need a return rule
- Believing a new subnet automatically gets a mount target