A three-Availability-Zone AWS VPC has a single NAT gateway serving the private subnets in all three AZs. What breaks when that one AZ has a problem, and what is the layout costing you even when everything is healthy?
answer
- the gateway belongs to one zone
- healthy zones inherit a sick one's fate
- every byte crosses a boundary twice
- one route table, one possible target
- more hourly charges, less per-GB
basics
~20 sA NAT gateway is zonal, so losing its AZ removes outbound internet for private subnets in all three — healthy zones inherit the failure. Healthy days are not free either: traffic from the other two AZs crosses zone boundaries and is charged per GB in each direction on top of NAT data processing.
solid answer
~50 sA NAT gateway lives in one subnet in one Availability Zone. AWS makes it redundant *within* that zone, but it does not span zones — so if that AZ is impaired, private subnets in the other two lose all outbound internet even though their own zone is perfectly healthy. You have taken three independent failure domains and made them share one. The steady-state cost is the second half of the answer: every byte from the other two AZs crosses an AZ boundary to reach the gateway and comes back the same way, and **cross-AZ data transfer is charged per GB in each direction** — on top of the NAT gateway's own per-GB data-processing charge and its hourly charge. The fix is the standard pattern: one NAT gateway per AZ, with each AZ's private route table sending `0.0.0.0/0` to the gateway in its own zone. You pay two more hourly charges and delete both the cross-zone transfer and the shared failure domain.
go deeper
Know that a NAT gateway is created in one subnet and therefore lives in one Availability Zone, and that the usual production layout puts one in each AZ. Knowing the resilience reason is enough at this level.
Explain both halves — losing the AZ removes egress everywhere, and traffic from other zones is billed as cross-AZ transfer in each direction on top of NAT data processing. Be able to describe the per-AZ route table layout that fixes it.
Show that you spot the partial-failure signature during an incident: the app is up, outbound calls time out, health checks look fine. Be ready to justify the per-AZ layout on cost as well as availability, and to say when consolidating is a deliberate, acceptable choice.
Own the standard: whether every account's network baseline enforces zone-local egress, how you detect drift from it across many accounts, and where the organisation draws the line between paying for independence and paying for traffic that should not traverse NAT at all.
## The zonal property is the whole question When you create a NAT gateway you choose a **subnet**, and a subnet belongs to exactly one Availability Zone. That makes the NAT gateway a zonal resource. AWS builds redundancy for it inside that zone — you are not exposed to a single host failing — but there is no such thing as a NAT gateway that spans zones, and no automatic failover to another zone if the one it lives in is impaired. This matters because the entire point of a multi-AZ layout is that the zones fail independently. Compute in one zone should not care what is happening in another. The moment every private subnet's default route points at one gateway in one zone, that independence is gone for anything that needs outbound internet: package installs, third-party API calls, license checks, webhook deliveries, and any AWS public endpoint not reached over a private path. The failure is also unpleasantly partial. The application in the surviving zones is up, health checks against it may pass, and instances keep serving requests that need nothing outbound — while every outbound-dependent code path times out. That mixed signal is much harder to diagnose during an incident than a clean outage. ## The bill in steady state Three charges apply to NAT gateway traffic, and the single-gateway layout maximises the wrong one. 1. **Hourly charge per NAT gateway.** One gateway means one hourly charge — the only line where the consolidated layout wins. 2. **Data processing, per GB.** Everything that passes through a NAT gateway is charged per GB, regardless of where it came from. 3. **Cross-AZ data transfer, per GB in each direction.** Traffic leaving an instance in one zone to reach a gateway in another crosses a zone boundary and is charged; the response crossing back is charged again. With one gateway in a three-AZ VPC, roughly two-thirds of your egress traffic pays that third charge — and pays it twice, once each way. Two additional hourly charges are a fixed, small, predictable number. Cross-AZ transfer scales with traffic, which means the consolidated design gets *relatively* worse exactly as the system succeeds. Above a fairly modest traffic volume, one NAT gateway per AZ is both cheaper and safer, which is a rare and very quotable combination. ## The pattern to describe One NAT gateway per Availability Zone, placed in that zone's public subnet, and one private route table per AZ whose default route targets the gateway in the same zone: ``` private-subnet-a -> route table A -> 0.0.0.0/0 -> nat-a (AZ a) private-subnet-b -> route table B -> 0.0.0.0/0 -> nat-b (AZ b) private-subnet-c -> route table C -> 0.0.0.0/0 -> nat-c (AZ c) ``` The mistake to call out explicitly: a single shared private route table across all three AZs makes the per-AZ pattern impossible to express, because a route table has one target for `0.0.0.0/0`. Teams end up with one gateway *because* they started with one route table. ## When consolidating is the right call Do not present per-AZ NAT as an absolute. In a development or sandbox account with negligible egress and no availability requirement, three gateways is three hourly charges for nothing, and one is a defensible, deliberate choice. The distinction an interviewer is listening for is whether you *chose* it: "one gateway in dev because outbound traffic is trivial and an AZ outage there costs us nothing" is a good answer; discovering at 3am that production was built the same way is not. There is also a middle position for very egress-heavy systems: keep per-AZ gateways for resilience, and separately reduce what has to traverse NAT at all. Traffic to services reachable over a private path never touches the NAT gateway, never incurs data processing, and never crosses a zone boundary — but that mechanism belongs to a different discussion; the point here is that per-AZ placement and traffic reduction are complementary, not alternatives. ## Detecting it in an existing account Two signals give it away. First, a NAT gateway's traffic volume in CloudWatch being wildly disproportionate to what one zone should generate. Second, a cost report where cross-AZ data transfer is a large line sitting next to NAT data processing — those two growing together is the fingerprint of instances in one zone routing egress through a gateway in another. A cost report is, in this specific case, an availability report too. ## What the interviewer is grading They want both halves — resilience *and* cost — because candidates reliably give one. Saying "a NAT gateway per AZ for high availability" and stopping earns partial credit; adding "and because otherwise two-thirds of my egress pays cross-AZ transfer in both directions" is the senior answer. Naming the shared route table as the structural cause of the anti-pattern is the detail that shows you have actually fixed one.
- How many private route tables does the per-AZ pattern require, and why?One per Availability Zone. A route table has a single target for 0.0.0.0/0, so a route table shared across zones can only ever point at one NAT gateway. Splitting the route tables is what makes zone-local egress expressible — teams that share one table end up with one gateway by construction.
- During an AZ impairment with per-AZ NAT gateways, does traffic automatically shift to a healthy zone's gateway?No, and it should not. Each AZ's private subnets keep using their own gateway; the impaired zone's workloads lose egress along with everything else in that zone, and your scaling and load balancing shift capacity away from it. The design contains the failure rather than routing around it.
- Is one NAT gateway per AZ ever the more expensive option overall?Yes — at very low egress volume, where two extra hourly charges outweigh the cross-AZ transfer you avoid. That is why sandbox and development accounts legitimately consolidate. The crossover comes quickly with real traffic, because the hourly charge is fixed while cross-AZ transfer scales with every gigabyte.
saying these in an interview costs you the question
- Believes a NAT gateway is regional or spans Availability Zones
- Mentions resilience but never the cross-AZ data transfer charge
- Thinks traffic fails over to another AZ's NAT gateway automatically
- Shares one private route table across all AZs and expects per-AZ egress
- Assumes consolidating NAT gateways is always the cheaper option