skip to content

VPC

A VPC is the private network your AWS resources live in: subnets, route tables, internet and NAT gateways, security groups, and peering. You need to draw one confidently, because "how does this private instance reach the internet?" is a standard whiteboard question.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

27

In an AWS VPC, how do security groups and network ACLs differ in what they attach to, what a rule can express, and how return traffic is treated?

level: juniorimportance: must knowfreq 85%

answer

  1. two filters, two attachment points
  2. one keeps state, one does not
  3. allow-only versus numbered allow and deny
  4. the reply needs its own rule

basics

~20 s

Security groups attach to network interfaces, are allow-only and stateful, so return traffic is automatically permitted. Network ACLs attach to subnets, allow or deny in numbered order, and are stateless, so every direction needs an explicit rule.

solid answer

~50 s

AWS gives you two independent packet filters in a VPC. A **security group** is attached to an elastic network interface — so effectively to an instance, task, Lambda ENI or load balancer — and holds only *allow* rules; there is no deny, and there is no ordering, because every rule in every attached group is evaluated and traffic passes if any of them matches. It is **stateful**: if a request is allowed in, the reply is allowed back out regardless of the outbound rules. A **network ACL** is attached to a subnet, so it filters everything entering or leaving that subnet. Its rules are numbered, evaluated lowest number first, first match wins, and they can be `deny` as well as `allow`. It is **stateless**: the reply to an allowed request is a brand-new packet that must match a rule in the opposite direction. Both must permit a packet for it to pass.

go deeper

for a junior

Be ready to state the four-part contrast cleanly: interface versus subnet, allow-only versus allow and deny, unordered versus numbered, stateful versus stateless. Say plainly that both have to permit a packet.

for a middle

Explain the mechanics: all attached security groups evaluate as one allow set, NACL rules stop at the first numeric match before falling to the implicit deny, and a stateless filter re-checks the reply as a fresh packet.

for a senior

Show that you know which one production actually uses. Expect to justify leaving NACLs wide open and doing real segmentation in security groups, and to name the case that flips it — a subnet-wide explicit deny.

for a principal

Own the ownership question: security groups belong with the workload and its team, NACLs are a platform-level guardrail. Be ready to argue where each control lives so that two teams are not fighting over one subnet's rules.

## Two filters, two attachment points Inside a VPC, a packet on its way to an instance is checked by two mechanisms that look superficially alike and behave very differently. A **security group** is attached to an *elastic network interface* (ENI). An ENI is the virtual NIC that AWS gives to an EC2 instance, an ECS task with `awsvpc` networking, a Lambda function configured for VPC access, an RDS instance, an interface VPC endpoint or a load balancer node. So a security group is an instance-level filter: two instances in the same subnet can have completely different security groups and completely different reachability. A **network ACL** (NACL) is attached to a *subnet*. Every ENI in that subnet is behind the same NACL, and every packet crossing the subnet boundary in either direction is evaluated against it. A subnet is associated with exactly one NACL at a time; one NACL can be associated with many subnets. Both must permit a packet. They are AND-ed, not layered as an override: a NACL `allow` does not rescue traffic the security group never permitted, and a permissive security group does not rescue traffic a NACL denies. ## What a rule can say Security-group rules are **allow-only**. There is no `deny` action — the absence of a matching allow rule *is* the denial. There is also no rule ordering: all rules across all security groups attached to the ENI are evaluated as a set, and the traffic passes if any of them matches. Because of this, attaching a second security group can only ever widen access, never narrow it. A rule's source (inbound) or destination (outbound) can be a CIDR block, a managed prefix list, or another security group. NACL rules are **numbered** — you assign the number yourself — and carry an explicit `allow` or `deny`. Evaluation walks the rules in ascending numeric order and stops at the first match; anything unmatched falls through to the untouchable final `*` rule, which denies. This ordering is why people leave gaps between rule numbers (100, 200, 300) so a rule can be inserted later without renumbering. The source or destination of a NACL rule is a CIDR block only — a NACL cannot reference a security group. ## Stateful versus stateless: the part interviews hinge on A security group **tracks connections**. When an inbound rule admits a TCP connection, the reply packets are permitted back out no matter what the outbound rules say, and vice versa for connections the instance originates. This is why the very common configuration — restrictive inbound rules plus the default "allow all outbound" rule — works fine, and why removing the outbound rule does not break inbound-initiated traffic. A NACL tracks nothing. Each packet is evaluated on its own against the rules for the direction it is moving. The reply to a request you allowed inbound is a separate outbound packet, and it will be dropped unless an outbound rule matches it. Because a reply travels *from* the service port *to* the client's randomly chosen ephemeral port, the outbound rule that permits it is written over the ephemeral port range, not over the service port. Getting that wrong is the single most common way a hand-written NACL breaks connectivity. ## The defaults A VPC ships with a **default security group** whose inbound rule allows traffic from resources that have that same security group attached, and whose outbound rule allows all traffic. A security group you create yourself starts with **no inbound rules** and an allow-all outbound rule. The **default network ACL** that comes with a VPC allows all inbound and all outbound traffic — it is deliberately transparent, so that people who never touch NACLs are only ever filtered by security groups. A NACL you create yourself starts with **only** the implicit `*` deny in both directions, so associating a fresh custom NACL with a live subnet blackholes it instantly. ``` # custom NACL, inbound 100 allow tcp 443 0.0.0.0/0 * deny all 0.0.0.0/0 ``` ## Why AWS ships both The security group is the workhorse: it follows the workload rather than its location, it composes safely, and its statefulness means you write half as many rules. The NACL is a coarse subnet-wide guardrail with the one capability security groups lack — an explicit `deny`, which lets a platform team block a CIDR range across a subnet without touching any application's rules. Most production VPCs leave NACLs wide open and do all real work in security groups.

  • If I attach a second security group to an instance, can that make the instance less reachable than it was?
    No. Security-group rules are allow-only and unordered, and all groups attached to the ENI are evaluated as one set — traffic passes if any rule in any of them matches. Adding a group can only widen access. Narrowing means removing rules from the groups already attached, or putting a deny in the subnet's network ACL.
  • Which filter would you reach for to block a single abusive IP range from an entire subnet, and why?
    The network ACL, because it is the only one of the two that has an explicit deny action and it applies subnet-wide. A security group cannot express "everyone except this range" — you would have to enumerate allowed CIDRs across every group. A NACL deny rule with a low rule number short-circuits before the permissive allow rules below it.
  • You create a new network ACL and associate it with a running subnet. What happens immediately?
    All traffic stops. A custom network ACL is created with only the implicit final rule, which denies everything in both directions, so the subnet blackholes until you add rules. The default NACL created alongside a VPC is the opposite — it allows all traffic in both directions.

saying these in an interview costs you the question

  • Claiming security groups support deny rules
  • Saying network ACL rules are evaluated top to bottom by creation order
  • Thinking a permissive network ACL can override a missing security-group allow
  • Saying security groups attach to subnets and NACLs to instances
  • Assuming return traffic is automatic on both filters

context

open as a page

In an AWS VPC, what actually makes a subnet "public" rather than "private", and what else does an EC2 instance in that subnet need before it can reach the internet?

level: juniorimportance: must knowfreq 88%

basics

~20 s

A subnet is public only because the route table associated with it sends 0.0.0.0/0 to an internet gateway. The instance also needs a public IPv4 or Elastic IP address, or its packets have no return path.

open as a page

In an AWS VPC, an instance in a private subnet can download OS package updates from the internet, yet nothing on the internet can open a connection to it. What does the internet gateway do, what does the NAT gateway add, and which of them makes that asymmetry possible?

level: middleimportance: must knowfreq 78%

basics

~20 s

An internet gateway attaches to the VPC and gives two-way internet reach, mapping public IPv4 addresses one-to-one. A NAT gateway hides private instances behind one public address and keeps state only for flows that start inside, so traffic can leave but connections can never be started inbound.

open as a page

VPC A is peered with VPC B, and VPC B is peered with VPC C. Instances in A cannot reach instances in C even though every route table looks right. Why, and what are your options?

level: middleimportance: must knowfreq 80%

basics

~20 s

VPC peering is non-transitive: each connection carries traffic only between the two VPCs it joins, so A cannot reach C through B. Fix it with a direct A-to-C peering connection, or attach all three VPCs to a Transit Gateway.

open as a page

You create a new subnet in an existing AWS VPC and never associate it with a route table. Which route table governs that subnet's traffic, and why does this trip teams up?

level: middleimportance: must knowfreq 52%

basics

~20 s

It falls back to the VPC's main route table through an implicit association. If that table has a 0.0.0.0/0 route to an internet gateway the new subnet is silently public; if it holds only the local route, nothing in it can leave the VPC.

open as a page

A three-Availability-Zone AWS VPC has a single NAT gateway serving the private subnets in all three AZs. What breaks when that one AZ has a problem, and what is the layout costing you even when everything is healthy?

level: seniorimportance: must knowfreq 52%

basics

~20 s

A NAT gateway is zonal, so losing its AZ removes outbound internet for private subnets in all three — healthy zones inherit the failure. Healthy days are not free either: traffic from the other two AZs crosses zone boundaries and is charged per GB in each direction on top of NAT data processing.

open as a page

An estate has grown from three VPCs to about twenty-five, all connected with VPC peering. What breaks down at that scale, and what does moving to a Transit Gateway change — including on the bill?

level: seniorimportance: must knowfreq 62%

basics

~20 s

A peering mesh grows as n(n−1)/2 connections with a route-table entry per peer in every VPC, so each new VPC touches all the others. A Transit Gateway replaces that with one attachment and one route per VPC, routes transitively, and adds per-attachment-hour and per-GB processing charges peering does not have.

open as a page

An EC2 instance launched into an AWS VPC public subnet receives a public IPv4 address automatically. What happens to that address when the instance is stopped and started again, and how does allocating an Elastic IP change the picture?

level: juniorimportance: should knowfreq 62%

basics

~20 s

An auto-assigned public IPv4 address is borrowed from an AWS pool and released when the instance stops or terminates, so a start hands out a different one. An Elastic IP belongs to your account, survives stop/start, and can be moved to another instance or interface.

open as a page

Two VPCs in the same AWS Region have a VPC peering connection, but instances in one still cannot reach instances in the other. What has to be in place before a peering connection actually carries traffic?

level: juniorimportance: should knowfreq 60%

basics

~20 s

A VPC peering connection carries traffic only after the peer accepts the request, both VPCs add route-table entries sending the other's CIDR to the peering connection, and security groups and NACLs on both sides allow the traffic.

open as a page

In AWS, when would you run a self-managed NAT instance on EC2 instead of a managed NAT gateway, and what do you take on operationally by doing so?

level: middleimportance: should knowfreq 45%

basics

~20 s

Choose a NAT instance only for low-traffic or cost-sensitive environments, or when you need something a NAT gateway cannot do, such as a security group, port forwarding, or traffic filtering. In exchange you own patching, availability failover, and a bandwidth ceiling set by the instance type.

open as a page

You need to connect an on-premises data center to a VPC. How do you choose between AWS Site-to-Site VPN and AWS Direct Connect?

level: middleimportance: should knowfreq 58%

basics

~20 s

Site-to-Site VPN is IPsec over the public internet: available the same day, cheap, but limited per-tunnel throughput and internet-grade latency. Direct Connect is a dedicated private link with predictable latency, high throughput and cheaper egress, but weeks of lead time and no encryption by default.

open as a page

A subnet's custom network ACL in AWS has an inbound rule allowing TCP 443 and an outbound rule allowing TCP 443, and the instances' security groups allow 443 inbound. Clients still time out. What is wrong?

level: middleimportance: should knowfreq 62%

basics

~20 s

The reply packets are being denied. A response leaves from source port 443 to the client's randomly chosen ephemeral port, so the outbound network ACL rule must allow the ephemeral range — AWS recommends TCP 1024-65535 — not port 443.

open as a page

In an AWS security group rule, what does it mean to name another security group as the source instead of a CIDR block, and where does that stop working?

level: middleimportance: should knowfreq 52%

basics

~20 s

It means "any network interface that currently has that security group attached", resolved to their private addresses. Membership replaces addresses, so rules survive scaling and IP churn. It only applies to traffic arriving on private addresses inside the VPC or a same-Region peered VPC.

open as a page

In an AWS VPC, how many usable IP addresses does a /24 subnet give you, and which addresses does AWS reserve in every subnet?

level: middleimportance: should knowfreq 58%

basics

~20 s

AWS reserves five addresses in every VPC subnet — the network address, the VPC router, the DNS resolver, one held for future use, and the last (broadcast) address — so a /24 gives 251 usable addresses, not 256.

open as a page

AWS creates every Site-to-Site VPN connection with two tunnels. Why two, and what does your side of the connection have to do for that redundancy to be real?

level: seniorimportance: should knowfreq 40%

basics

~20 s

AWS terminates each Site-to-Site VPN connection on two independent endpoints so the connection survives the loss or maintenance of one. The redundancy is only real if your customer gateway is configured for both tunnels and your routing fails over — plus a second customer gateway for device redundancy.

open as a page

An EC2 instance's VPC Flow Logs show an ACCEPT record for an inbound TCP flow and a REJECT record for the same addresses with the ports reversed. Which layer is dropping the traffic, and how would you confirm it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The subnet's network ACL is dropping the reply. A stateful security group never rejects the response to a flow it admitted, so an outbound REJECT on an accepted connection points at the stateless subnet filter — almost always a missing egress rule over the ephemeral port range.

open as a page

Launches into an AWS VPC subnet start failing because it has run out of free IPv4 addresses. What are your options, and why can't you simply make that subnet bigger?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A subnet's CIDR is immutable, so it cannot grow. You add capacity instead: carve a new subnet in the same Availability Zone from unused VPC space, or attach a secondary CIDR block to the VPC and build subnets from that, then reclaim addresses held by idle interfaces.

open as a page

IPv6 addresses assigned inside an AWS VPC are globally routable. How do you give IPv6 instances outbound internet access while keeping them unreachable from the internet, and why is a NAT gateway not the answer?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

Use an egress-only internet gateway and route ::/0 to it from the private subnets. It is stateful, allowing outbound IPv6 and the replies while dropping anything initiated from outside. NAT exists to conserve scarce IPv4 addresses, a problem IPv6 does not have.

open as a page

Instances behind an AWS NAT gateway start failing to connect to one busy third-party API while every other destination works fine, and the gateway's ErrorPortAllocation metric is non-zero. What is happening, and what are your options?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

The NAT gateway has run out of source ports for that one destination. Its simultaneous-connection capacity is counted per unique destination and per associated IP address, so a single hot endpoint exhausts its port range while every other destination is unaffected. Add addresses, spread the load, or open far fewer connections.

open as a page

In an AWS Transit Gateway, what is the difference between associating an attachment with a route table and propagating that attachment's routes into one, and how do you use both to keep dev and prod isolated while both reach shared services?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Association picks the one route table a Transit Gateway consults for traffic arriving from an attachment; propagation copies that attachment's routes into any route tables you choose. Segmentation comes from giving dev and prod separate route tables that propagate shared services but never each other.

open as a page

Security groups already filter traffic per workload in a VPC. When is it worth also managing network ACLs, and what does that decision cost you operationally?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Network ACLs earn their place when you need an explicit deny or a subnet-wide guardrail that no application team can undo. The cost is real: stateless rules written twice, a small rule quota, no identity references, and a filter that fails silently.

open as a page

You are setting the IPv4 addressing standard for a new AWS organization that will hold dozens of accounts and VPCs. How do you allocate CIDR ranges, and what are you optimizing for?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Give every VPC a non-overlapping slice of one organization-wide private supernet, sized with headroom and grouped so ranges summarize per region and environment. Overlap is the one decision you cannot reverse cheaply, so a central registry governs allocation.

open as a page