skip to content

An EC2 instance's VPC Flow Logs show an ACCEPT record for an inbound TCP flow and a REJECT record for the same addresses with the ports reversed. Which layer is dropping the traffic, and how would you confirm it?

level: seniorimportance: should knowfreq 45%

answer

  1. two records, one per direction
  2. which filter re-checks the reply?
  3. a stateful filter cannot reject its own reply
  4. subnet egress, ephemeral ports
  5. the action field never names the rule

basics

~20 s

The subnet's network ACL is dropping the reply. A stateful security group never rejects the response to a flow it admitted, so an outbound REJECT on an accepted connection points at the stateless subnet filter — almost always a missing egress rule over the ephemeral port range.

solid answer

~50 s

Flow logs record each direction as its own record, and the pairing localises the fault. The inbound `ACCEPT` proves the request cleared both the network ACL and the security group and reached the ENI. The reply then shows up as a separate outbound record with source and destination reversed, and it is `REJECT`. Since security groups are stateful, they permit the response to an admitted flow no matter what their egress rules say — so a security group is not capable of producing that record. What remains is the subnet's network ACL, which is stateless and evaluates the reply on its own, typically dropping it because the outbound rules cover the service port rather than the client's ephemeral port. Confirm by reading the NACL's outbound rules in ascending order to find which one matches first, then widen the egress rule to `1024-65535` and watch the REJECT records stop.

go deeper

for a junior

Know that flow logs record each direction separately with an ACCEPT or REJECT action, and that ports appear reversed on the reply. Recognise the record pair as evidence rather than noise.

for a middle

Explain why a stateful security group cannot produce an outbound REJECT for a flow it admitted, so the stateless subnet ACL is the only candidate, and name the ephemeral-port egress rule as the usual fix.

for a senior

Drive the whole triage: establish direction, read the subnet's outbound rules in number order for the first match, check the far subnet too, then verify the REJECT records stop. Know which flows are never logged so their absence does not mislead you.

for a principal

Own the observability posture — which VPCs log flows, where the records land, what they cost at scale, and how much of this triage should be a runbook or an automated reachability check rather than an engineer reading records by hand.

## Reading the record pair A VPC Flow Log record covers one direction of one flow over an aggregation interval, and includes the source and destination addresses, the source and destination ports, the protocol, packet and byte counts, and an `action` of `ACCEPT` or `REJECT`. For a single TCP connection into an instance you therefore expect two records: ``` ... 203.0.113.9 10.0.7.20 51314 443 6 ... ACCEPT OK ... 10.0.7.20 203.0.113.9 443 51314 6 ... REJECT OK ``` Same pair of addresses, ports swapped. The first is the request arriving; the second is the reply leaving. `ACCEPT` then `REJECT` is a specific, diagnosable fingerprint. ## Why the security group is exonerated Security groups track connections. Once an inbound rule admits a flow, the response is permitted outbound irrespective of the group's egress rules — the check simply is not made against the rule set. So a security group can produce an inbound `REJECT` (no matching allow rule), but it cannot produce an outbound `REJECT` for a flow it just accepted. That single property is what turns the record pair into a conclusion rather than a guess. Network ACLs keep no state. The reply is a fresh packet evaluated against the subnet's outbound rules in ascending rule-number order, first match wins, falling through to the implicit final deny. Because the reply's destination is the client's ephemeral port — high and effectively random — an outbound rule written over the service port never matches it. That is the overwhelmingly common cause of this exact fingerprint. ## The confirmation walk 1. **Fix the direction first.** Check which address is the instance. If the REJECT has the instance as *source*, the reply is being dropped and the subnet's egress rules are the suspect. If the instance is the *destination* on a REJECT with no matching ACCEPT anywhere, the request itself never got in and you are looking at an inbound problem instead. 2. **Read the NACL for that subnet, outbound, in number order.** Find the first rule that matches the reply's protocol, destination address and destination port. If none does, the implicit deny is your answer. Watch for a low-numbered `deny` shadowing a correct allow further down — first match wins, so rule 90 beats rule 100 even when rule 100 is exactly right. 3. **Check both subnets when the flow crosses one.** A packet between two subnets in the same VPC is evaluated by the source subnet's egress rules *and* the destination subnet's ingress rules. Asymmetric NACLs are why a flow works in one direction only. 4. **Fix and re-observe.** Widen the egress rule to the ephemeral range `1024-65535`, then confirm the REJECT records stop rather than assuming they did. ## What flow logs will not tell you The `action` field says `REJECT`; it does not name the rule, the NACL or the security group that produced it. The layer has to be inferred, which is exactly why the stateful-versus-stateless reasoning above is the whole skill. Several flows are never logged at all, and treating their absence as evidence sends you down the wrong path: traffic to and from the instance metadata endpoint at `169.254.169.254`, traffic to and from the Amazon-provided DNS server, DHCP traffic, and traffic to the reserved VPC router address. If you are debugging "my instance cannot reach the metadata service", flow logs will show you nothing, and the cause is more likely IMDS configuration than a packet filter. Records are also aggregated over an interval — one minute or ten depending on how the flow log was created — so a record's appearance lags the event, and a very short flow shows up as a single record rather than a live packet trace. Flow logs answer "was this permitted", not "what happened in the last two seconds". ## The other pairings worth memorising - **Inbound REJECT, nothing else.** Either the security group has no matching inbound allow, or the subnet's NACL denies inbound. Compare the security group's rules against the destination port first — that is the cheaper check. - **No records at all for the flow.** The packet never reached the ENI. Look at routing, at whether the client is even resolving to the right address, and at the load balancer or endpoint in front of the instance — not at the filters. - **ACCEPT in both directions but the client still fails.** The filters are innocent; the problem is above the network — the application, TLS, or a timeout somewhere in the path. ## Why interviewers like this one It cannot be answered by reciting "security groups are stateful, NACLs are stateless". It requires using that fact as an inference rule against real evidence, which is the difference between having read the documentation and having debugged a VPC at two in the morning.

  • You see an inbound REJECT and no other record at all for the flow. What are the candidates now?
    The request itself was dropped at the ENI, so either the security group has no inbound rule matching that destination port, or the subnet's network ACL denies it inbound. Check the security group first — it is one rule set and the likelier culprit. If the ACL is in play, read its inbound rules in number order and look for a low-numbered deny shadowing the allow.
  • The flow produces no flow-log records whatsoever. What does that tell you?
    That the packets never reached the network interface, so the filters are not the story — look at routing, name resolution, or whatever sits in front of the instance. Be careful with a few flows that are never logged regardless: the metadata endpoint at 169.254.169.254, the Amazon DNS server, DHCP, and the reserved VPC router address.
  • Why can flow logs not simply tell you which rule rejected the packet?
    The action field carries only ACCEPT or REJECT — it is a record of the verdict, not of the evaluation. Attribution has to be inferred from the direction and from which filter is even capable of producing that verdict, which is why the stateful-versus-stateless property does the diagnostic work. Reachability Analyzer is the AWS tool that does name the blocking component.

saying these in an interview costs you the question

  • Blaming the security group's outbound rules for the rejected reply
  • Assuming REJECT identifies which rule dropped the packet
  • Forgetting that the far subnet's inbound ACL also evaluates the flow
  • Treating missing records as proof the traffic was denied
  • Ignoring that a low-numbered deny can shadow a correct allow

context