A subnet's custom network ACL in AWS has an inbound rule allowing TCP 443 and an outbound rule allowing TCP 443, and the instances' security groups allow 443 inbound. Clients still time out. What is wrong?
answer
- check each direction on its own
- which port does the reply actually carry?
- source 443, destination is arbitrary
- the reply lands on an ephemeral port
- outbound allow 1024-65535
basics
~20 sThe reply packets are being denied. A response leaves from source port 443 to the client's randomly chosen ephemeral port, so the outbound network ACL rule must allow the ephemeral range — AWS recommends TCP 1024-65535 — not port 443.
solid answer
~50 sNetwork ACLs are stateless, so the reply to an admitted request is judged on its own. That reply travels from source port 443 *to* the client's ephemeral port — a high, effectively random destination port — so an outbound rule scoped to destination port 443 never matches it and the handshake dies after the client's SYN. The fix is an outbound rule allowing TCP to the ephemeral range; AWS's guidance is `1024-65535`, because the actual range depends on the client: Linux typically uses 32768-60999, recent Windows uses 49152-65535, and NAT gateways and Elastic Load Balancing use 1024-65535. The mirror image applies to connections the subnet *initiates*: outbound to port 443 plus an **inbound** rule over the ephemeral range for the replies. The security groups are innocent here — they are stateful and permit the reply automatically.
code
bash · 11 linesaws ec2 create-network-acl-entry \
--network-acl-id acl-0123456789abcdef0 \
--rule-number 100 --protocol tcp \
--port-range From=443,To=443 \
--cidr-block 0.0.0.0/0 --rule-action allow --ingress
aws ec2 create-network-acl-entry \
--network-acl-id acl-0123456789abcdef0 \
--rule-number 100 --protocol tcp \
--port-range From=1024,To=65535 \
--cidr-block 0.0.0.0/0 --rule-action allow --egressgo deeper
Recall that a network ACL is stateless and that a reply comes back on a high, randomly chosen port. Be able to say the outbound rule has to cover the ephemeral range rather than the service port.
Walk the four-tuple out loud: request to port 443, reply from port 443 to the client's ephemeral port, no matching outbound rule, implicit deny. Then give the mirrored rule set for outbound-initiated traffic.
Show you can confirm it rather than guess — flow-log records with reversed ports, an inbound ACCEPT paired with an outbound REJECT — and explain why the range must cover NAT gateway and load balancer sources, not just Linux clients.
Argue the policy: every stateless rule is written twice over a range you must guess, so subnet-level filtering carries a standing outage risk. Decide deliberately when that guardrail is worth the operational cost.
## What actually goes over the wire A TCP connection has two endpoints with two ports. The client picks a local port from its operating system's *ephemeral* range — a high-numbered, essentially arbitrary port used for the lifetime of that one connection — and connects to the server's well-known port, 443. So the flow looks like this: ``` request: client:51314 -> server:443 reply: server:443 -> client:51314 ``` Only the request carries 443 as a *destination*. In the reply, 443 is the **source** port and the destination is the client's ephemeral port. ## Why the network ACL drops the reply A network ACL is attached to a subnet and is **stateless**: it keeps no record that it just admitted a request, so the reply is evaluated from scratch against the rules for the direction it is travelling. The outbound rule in the question permits TCP with destination port 443. The reply's destination port is 51314. No rule matches, so it falls through to the implicit final rule, which denies. The server sends its SYN-ACK, the subnet swallows it, and the client sits in `SYN_SENT` until it times out — which is exactly the symptom described: a hang, not a refusal. This is the single most common way a hand-written NACL breaks a working system, and it is why interviewers reach for it: it can only be diagnosed by someone who has internalised what "stateless" means, rather than memorised the word. ## The correct rule set For a subnet hosting HTTPS servers: ``` inbound 100 allow tcp 443 0.0.0.0/0 outbound 100 allow tcp 1024-65535 0.0.0.0/0 ``` The outbound rule is deliberately wide because you cannot know which ephemeral port any given client picked. AWS's documented recommendation is `1024-65535`, chosen to cover every client you might see. The underlying ranges differ by stack: Linux kernels default to `32768-60999` (exposed as `net.ipv4.ip_local_port_range`), Windows Server 2008 and later use `49152-65535`, and AWS's own NAT gateways and Elastic Load Balancing nodes use `1024-65535`. If your traffic arrives via a NAT gateway or a load balancer, that last range is the one that matters, so narrowing the rule to the Linux range would break it. ## The mirror case The same trap exists in the other direction and catches people just as often. If instances in the subnet *call out* — to an API, to a package repository, to an RDS endpoint — you need: ``` outbound 100 allow tcp 443 0.0.0.0/0 inbound 100 allow tcp 1024-65535 0.0.0.0/0 ``` That inbound ephemeral rule is uncomfortable to write, because it looks like you are opening the subnet to the world on 65,000 ports. In practice you constrain it by CIDR where you can, and you rely on the *security group* to do the real filtering — the packet still has to get past an instance-level allow rule, and a stray packet to a high port with no listening socket goes nowhere. ## Why the security groups are not the problem Security groups attach to elastic network interfaces and are **stateful**: they track the connection, so a reply to an admitted request is permitted outbound regardless of the outbound rules, and the reply to a connection the instance originated is permitted inbound. You never write ephemeral-port rules in a security group. If a candidate proposes adding an ephemeral outbound rule to the security group as the fix, they have not understood which of the two filters is stateless. ## How to confirm it quickly VPC Flow Logs on the instance's ENI settle the argument. You will see an `ACCEPT` record for the inbound request and a `REJECT` record for the outbound reply carrying the same addresses with the ports reversed. Since a stateful security group never rejects the reply to a flow it admitted, an outbound `REJECT` on an accepted flow points at the subnet's network ACL. ## The lesson beyond the puzzle This is the strongest practical argument for leaving network ACLs at their permissive default and doing segmentation in security groups. Every stateless rule you write is a rule you have to write twice, in a direction that is not intuitive, over a port range you have to guess correctly for every client stack that will ever talk to the subnet.
- Why does AWS recommend 1024-65535 rather than the narrower range a Linux client actually uses?Because you do not control every client. Linux defaults to 32768-60999, recent Windows to 49152-65535, and AWS's own NAT gateways and Elastic Load Balancing nodes source traffic from 1024-65535. A rule scoped to one stack's range silently drops replies to the others, so the recommendation covers the union.
- Does the same problem appear for UDP or ICMP through a network ACL?Yes for UDP — it has ports, replies land on the client's ephemeral port, and the ACL still keeps no state, so you need the mirrored rule. ICMP has no ports; you allow the specific ICMP type in each direction instead, for example echo request inbound and echo reply outbound, which is a classic reason a ping into a subnet never answers.
- The symptom was a timeout rather than a connection refused. Why does that detail matter?A silent drop makes the client wait through its SYN retransmissions, which is the fingerprint of a filter discarding packets — a network ACL deny, or a security group with no matching rule. A connection refused means a RST came back, so packets are reaching a host and nothing is listening on the port. The distinction tells you whether to look at rules or at the process.
saying these in an interview costs you the question
- Adding an outbound ephemeral rule to the security group as the fix
- Believing the network ACL remembers the inbound connection
- Assuming the ephemeral range is always 32768-60999
- Thinking outbound rules matter only for connections the instance initiates
- Blaming DNS or routing because the failure is a timeout