An AWS Lambda function worked fine until it was attached to a VPC; now every invocation hangs and times out when it calls an external HTTPS API. Explain the cause and how you would diagnose and fix it.
answer
- it hangs, it does not refuse
- interfaces hold private addresses only
- a public subnet does not help here
- default route target is the question
- endpoints keep AWS traffic off that path
basics
~20 sA VPC-attached function has private addresses only, so outbound internet traffic goes nowhere unless the subnet routes it. Fix it by placing the function in private subnets whose route table sends 0.0.0.0/0 to a NAT gateway, or by using a VPC endpoint for AWS-service calls.
solid answer
~50 sAttaching the function moved its traffic onto interfaces in your subnets, and those interfaces only ever have private addresses — Lambda cannot be given a public IP or an Elastic IP. So the packets leave, hit a route table with no path to the internet, and are dropped; the connection never completes, which is why you see a hang until the function's timeout rather than an immediate refusal. Diagnosis is a short checklist: confirm the configured subnets' route table has a `0.0.0.0/0` entry, check that it points at a NAT gateway in a *different*, public subnet, verify the security group permits egress on 443, and check the network ACLs allow return traffic. VPC Flow Logs on the interfaces confirm whether packets are being rejected or simply going nowhere. The fixes: private subnets plus a NAT route for genuine internet destinations, a VPC endpoint when the destination is an AWS service, or detaching from the VPC entirely if nothing private is actually needed.
go deeper
Recall the core fact: a function attached to a VPC has no public address, so it cannot reach the internet unless the network is set up to route it. Say that raising the timeout is not a fix.
Explain why the call hangs instead of failing fast, and name the correct layout — the function in private subnets with a default route to a NAT gateway that sits in a public subnet.
Walk the diagnosis in order: actual configured subnets, their route tables, the NAT gateway's placement, security-group egress, network ACL return traffic, then flow logs — and know that AWS service endpoints hit the same wall.
Frame the standing decision: which workloads may be attached to a VPC at all, whether egress goes through NAT or endpoints given the data-transfer bill, and how that path is made uniform across accounts rather than rediscovered per team.
## Why it hangs rather than fails fast The shape of the symptom is itself a clue. A permissions failure returns `AccessDenied` in milliseconds. A DNS failure returns a resolution error. A missing network path produces neither: the TCP SYN is emitted, nothing routes it, nothing answers, and the client sits in connect until something gives up. If the function's configured timeout is 30 seconds you get a 30-second invocation and a `Task timed out` message, and the log tells you almost nothing about why. Learning to read "hangs to the timeout on an outbound call" as *network path*, not *slow dependency*, is most of the diagnosis. ## The cause Before attachment, the function ran in AWS-managed network space with outbound internet access. After attachment, its traffic rides elastic network interfaces in your subnets, and those interfaces hold private addresses only. There is no way to attach a public IP or an Elastic IP to them, and no `AssociatePublicIpAddress` equivalent in the function's VPC configuration. This produces the single most useful counter-intuitive fact in the topic: **putting a Lambda function in a public subnet does not give it internet access.** A public subnet is one whose route table points `0.0.0.0/0` at an internet gateway, and an internet gateway performs one-to-one NAT for resources that have a public address. The function has none, so its packets reach the gateway and are dropped. The working arrangement is the opposite of intuition: place the function in **private** subnets whose route table sends `0.0.0.0/0` to a **NAT gateway** that itself lives in a public subnet. Also note that "the internet" here includes public AWS service endpoints. A VPC-attached function calling S3, Secrets Manager, DynamoDB or SQS is making an outbound call to a public endpoint and fails exactly the same way. Teams often discover this second, after fixing the third-party call. ## Diagnosis, in order 1. **Read the function's VPC configuration.** `aws lambda get-function-configuration --query VpcConfig` gives you the exact subnet and security group IDs — not the ones you think are configured. 2. **Inspect those subnets' route tables.** Look for a `0.0.0.0/0` route and what it targets. An `igw-` target on a function subnet is the bug. A `nat-` target is what you want. No default route at all is the other common finding. ```bash aws ec2 describe-route-tables \ --filters Name=association.subnet-id,Values=subnet-0a1b2c3d \ --query 'RouteTables[].Routes' ``` 3. **Check the NAT gateway is not in one of the same subnets.** A NAT gateway must sit in a public subnet with a route to an internet gateway. Putting it in the same private subnet the function uses creates a routing loop that also manifests as a hang. 4. **Check the security group's egress rules.** The default group allows all outbound, but a purpose-built group may allow nothing. Egress on 443 to the destination must be permitted; because security groups are stateful, no inbound rule is needed for the response. 5. **Check network ACLs.** They are stateless, so a NACL that allows outbound 443 but not inbound on the ephemeral port range silently kills the response half of the conversation. 6. **Confirm with VPC Flow Logs** on the function's interfaces. `REJECT` records point at a security group or NACL; the absence of any record for the destination points at routing. A quick discriminator worth knowing: if name resolution succeeds but the connection hangs, DNS is fine and the problem is routing or filtering. Resolution works through the VPC resolver regardless of whether you have an egress path, so a successful lookup proves nothing about reachability. ## The fixes, and how to choose - **Private subnets plus a NAT gateway route.** The general answer, and the only one that works for arbitrary third-party endpoints. It is a per-hour charge plus per-GB processing, and it is a per-AZ resource — one NAT gateway per AZ if you care about both resilience and avoiding cross-AZ traffic. - **A VPC endpoint** when the destination is an AWS service. This keeps that traffic off the NAT path entirely, which matters both for the bill and for environments where internet egress is disallowed by policy. - **Do not attach at all.** Ask what private resource actually justified the attachment. If the function only calls AWS service APIs and a third-party endpoint, taking it back out of the VPC removes the whole class of problem. - **Split the function.** When one function genuinely needs both a private database and a public API, that is a real design point: keep the database-touching part in the VPC and let a second, unattached function or an outbound path through the NAT handle the external call. ## What raising the timeout does Nothing useful. It converts a 30-second failure into a 15-minute one and multiplies the bill by thirty. Anyone who reaches for the timeout dial first has not understood the symptom.
- Why does moving the function into a public subnet not fix it?An internet gateway only translates for resources that already hold a public address, and a function's network interfaces never get one — you cannot assign a public IP or an Elastic IP to them. So packets reach the gateway and are discarded. The working layout inverts the intuition: private subnets for the function, with `0.0.0.0/0` routed to a NAT gateway that lives in a public subnet.
- The same function also fails calling Secrets Manager. Is that a different problem?No — it is the same one. Public AWS service endpoints are reached over the internet path, so a VPC-attached function without egress fails on them exactly as it fails on a third-party API. Either the NAT route fixes both, or you add a VPC endpoint for that service so its traffic stays inside the VPC and never needs the NAT path at all.
- How would you prove it is routing rather than a slow dependency?Look at the shape of the failure and the flow records. A dependency that is merely slow eventually returns something; a missing route hangs until the function's own timeout, every single invocation, with no partial progress. VPC Flow Logs on the function's interfaces settle it: REJECT entries implicate a security group or network ACL, while no entries at all for that destination implicate routing.
saying these in an interview costs you the question
- Raises the function timeout and calls it fixed
- Moves the function to a public subnet to get internet access
- Tries to attach an Elastic IP to the function
- Checks only security groups and never the route table
- Assumes successful DNS resolution proves the path works