Networking & CDN
The network is the part of an AWS design you have to draw: VPCs and subnets, route tables, internet and NAT gateways, peering and Transit Gateway, and security groups versus NACLs. You add Route 53 routing policies, CloudFront, and the ALB-versus-NLB choice, because every system-design round ends up here.
part ofAWSoverview, primer and where to startread it →on this pageshowhide
explore
- Elastic Load Balancing18 questions
- Application Load Balancer (ALB)6 questions
- Network Load Balancer (NLB)6 questions
- Target Groups & Health Checks6 questions
- VPC27 questions
- Subnets, CIDR & Route Tables5 questions
- Internet Gateway & NAT6 questions
- Security Groups vs NACLs5 questions
- VPC Endpoints & PrivateLink5 questions
- Peering, Transit Gateway & Hybrid6 questions
- Route 5310 questions
- Hosted Zones & Record Types4 questions
- Routing Policies & Health Checks6 questions
- CloudFront17 questions
- Distributions & Origins6 questions
- Caching, Cache Keys & Invalidation5 questions
- Edge Security & Edge Functions6 questions
- API Gateway16 questions
- REST vs HTTP APIs & Integrations6 questions
- Authorizers & Access Control5 questions
- Throttling, Quotas & Caching5 questions
questions
88 · 5 sectionsOn an AWS Application Load Balancer, what can a listener rule match on, and how does the load balancer decide which rule applies to a given request?
basics
~20 sAn ALB listener rule matches on host header, URL path, HTTP header, HTTP method, query string, or source IP. Rules are evaluated in priority order, lowest number first; the first match wins, and unmatched requests fall through to the listener's default action.
In AWS Elastic Load Balancing, what is a target group, and what happens to a registered target once it starts failing that target group's health check?
basics
~20 sA target group is the named set of backends an Elastic Load Balancing listener forwards to, plus the health check run against each member. A target that fails enough consecutive checks is marked unhealthy and stops receiving new requests.
You are putting a public HTTPS API in front of a fleet of containers on AWS. When would you choose an Application Load Balancer (ALB) over a Network Load Balancer (NLB), and what can an ALB do that an NLB cannot?
basics
~20 sPick an ALB when the routing decision depends on the HTTP request itself - host, path, header or method - or when you want WAF, cookie stickiness or built-in OIDC login. Pick an NLB for raw TCP/UDP, static IPs and lowest latency.
Your service must be reachable from outside its VPC. When would you put an AWS Network Load Balancer in front of it instead of an Application Load Balancer, and what do you give up by choosing NLB?
basics
~20 sNetwork Load Balancer is the right choice when traffic is not HTTP, when you need a fixed IP per Availability Zone, or when you need very high connection volume at minimal added latency. You give up HTTP-aware routing, WAF and authentication actions.
An Application Load Balancer is failing requests and every target in its target group shows unhealthy, yet curling the health-check path directly against a target from inside the VPC returns 200. How do you find the cause?
basics
~20 sRead the reason code from describe-target-health first. Target.Timeout means the probe never arrived — usually the target's security group does not allow the health-check port from the load balancer. Target.ResponseCodeMismatch means it arrived but the status code fell outside the matcher.
In an AWS VPC, how do security groups and network ACLs differ in what they attach to, what a rule can express, and how return traffic is treated?
basics
~20 sSecurity groups attach to network interfaces, are allow-only and stateful, so return traffic is automatically permitted. Network ACLs attach to subnets, allow or deny in numbered order, and are stateless, so every direction needs an explicit rule.
In an AWS VPC, what actually makes a subnet "public" rather than "private", and what else does an EC2 instance in that subnet need before it can reach the internet?
basics
~20 sA subnet is public only because the route table associated with it sends 0.0.0.0/0 to an internet gateway. The instance also needs a public IPv4 or Elastic IP address, or its packets have no return path.
In an AWS VPC, what is the difference between a gateway VPC endpoint and an interface VPC endpoint, and how do you choose between them?
basics
~20 sGateway endpoints add a route-table entry for S3 or DynamoDB only and cost nothing. Interface endpoints put a PrivateLink network interface with a private IP into your subnets, cover most AWS services, and bill per endpoint-hour plus per gigabyte processed.
In an AWS VPC, an instance in a private subnet can download OS package updates from the internet, yet nothing on the internet can open a connection to it. What does the internet gateway do, what does the NAT gateway add, and which of them makes that asymmetry possible?
basics
~20 sAn internet gateway attaches to the VPC and gives two-way internet reach, mapping public IPv4 addresses one-to-one. A NAT gateway hides private instances behind one public address and keeps state only for flows that start inside, so traffic can leave but connections can never be started inbound.
VPC A is peered with VPC B, and VPC B is peered with VPC C. Instances in A cannot reach instances in C even though every route table looks right. Why, and what are your options?
basics
~20 sVPC peering is non-transitive: each connection carries traffic only between the two VPCs it joins, so A cannot reach C through B. Fix it with a direct A-to-C peering connection, or attach all three VPCs to a Transit Gateway.
In Route 53, why can't you point the zone apex example.com at an Application Load Balancer with a CNAME record, and what do you use instead?
basics
~20 sDNS forbids a CNAME at a zone apex, which must also hold SOA and NS records. Route 53's alias record is an A record whose target is an AWS resource, so the apex can point at a load balancer and still return plain addresses.
You set up Amazon Route 53 failover records for an active-passive design across two regions. The primary region goes down, but users keep hitting it for several minutes before traffic moves. Walk through everything that contributes to that delay and what you would change.
basics
~20 sTwo delays stack: Route 53 must see enough consecutive failed health-check probes to mark the primary unhealthy, and then every cached copy of the old answer must expire. Shorten both with a faster health check, a lower failure threshold, and a low record TTL.
You created a public hosted zone for example.com in Route 53 and added an A record for www, but the name still does not resolve anywhere on the internet. What step is most likely missing, and where is it performed?
basics
~20 sDelegation. A new Route 53 public hosted zone is assigned four name servers, and nothing on the internet consults them until the domain's registrar publishes those four as the domain's NS records. Records exist but are unreachable until then.
An application runs in three AWS regions behind one name. In Amazon Route 53, when would you choose latency-based routing over geolocation routing, and what does each policy actually decide on?
basics
~20 sLatency-based routing sends a query to whichever AWS region Route 53 measures as fastest from the requester's network, so it optimises performance. Geolocation routing answers by the requester's mapped location — continent, country or subdivision — so it enforces where traffic must go.
Using Amazon Route 53 weighted records, you want roughly 5% of traffic for api.example.com to reach a newly deployed stack. How do weighted records produce that split, and why will the real share of users differ from 5%?
basics
~20 sWeighted records share a name, and Route 53 picks one per query with probability equal to its weight divided by the total — for example 5 against 95. The real user share drifts because each cached answer serves many clients for the whole TTL.
In a CloudFront distribution, what is the difference between an origin and a cache behavior, and how does CloudFront decide which origin serves an incoming request?
basics
~20 sAn origin is a backend CloudFront fetches from: an S3 bucket, a load balancer, any HTTP server. A cache behavior maps a URL path pattern to one origin. CloudFront tests behaviors in their configured order and falls back to the default behavior.
In CloudFront, what decides the cache key for a request, and what happens to a query string, cookie or header that is named in neither the cache policy nor the origin request policy attached to the cache behavior?
basics
~20 sCloudFront's cache key is the distribution, the matched cache behavior and exactly the query strings, headers and cookies named in the attached cache policy. Values named in neither policy are stripped, so the origin never sees them at all.
How do you configure a CloudFront distribution and an S3 bucket so the bucket's objects are readable only through the distribution, and what does Origin Access Control (OAC) do that the legacy Origin Access Identity (OAI) does not?
basics
~20 sUse the bucket's REST endpoint as the origin, keep Block Public Access on, and attach an Origin Access Control so CloudFront signs each origin request with SigV4; the bucket then grants read only to the CloudFront service principal for that distribution. OAC, unlike OAI, works with SSE-KMS objects and non-GET methods.
CloudFront supports both CloudFront Functions and Lambda@Edge. Which trigger points can each hook into, and how would you choose between them for a given piece of edge logic?
basics
~20 sCloudFront Functions run only on viewer request and viewer response, in a tiny sandboxed JavaScript runtime with no network access — ideal for header and URI rewrites. Lambda@Edge also hooks origin request and origin response, and can call other services.
In Amazon CloudFront, when would you protect private content with signed URLs versus signed cookies, and what does a trusted key group have to do with either?
basics
~20 sCloudFront signed URLs authorize one object each and change the URL; signed cookies authorize many objects and leave URLs unchanged. Both are verified by CloudFront against the public keys in the trusted key group attached to that cache behavior.
An API Gateway route uses the Lambda proxy (`AWS_PROXY`) integration and every call returns HTTP 502 with `{"message": "Internal server error"}`, yet the function's own CloudWatch logs show it completing without error. What is wrong, and what must a Lambda proxy handler return?
basics
~20 sThe function returned a shape API Gateway cannot map to an HTTP response. A Lambda proxy handler must return an object with a numeric statusCode, an optional headers map, and a body that is already a string — returning a raw object or a bare value produces 502.
Amazon API Gateway offers REST APIs and HTTP APIs (as well as WebSocket APIs). What would make you choose a REST API over an HTTP API, and what are you giving up if you default to HTTP APIs?
basics
~20 sHTTP APIs are cheaper and lower-latency, so default to them for plain Lambda or HTTP backends. Choose REST APIs when you need what HTTP APIs omit: API keys and usage plans, stage response caching, request validation, VTL mapping templates, private endpoints, or AWS WAF.
You are choosing how Amazon API Gateway will authorize callers of a new API. Compare IAM (SigV4) authorization, a Cognito user pool authorizer, an HTTP API JWT authorizer, and a Lambda authorizer — what decides which one you pick?
basics
~20 sPick by who the caller is. AWS principals get IAM SigV4. End users with tokens from an OIDC issuer get the built-in JWT authorizer on an HTTP API, or a Cognito user pool authorizer on a REST API. Anything else needs a Lambda authorizer.
A REST API in Amazon API Gateway is protected by a Lambda authorizer. What exactly must that function return for the request to be allowed, and how does the backend integration learn who the caller was?
basics
~20 sA REST API Lambda authorizer must return a principalId, plus a policyDocument — an IAM policy allowing execute-api:Invoke on the requested method ARN. Anything in the optional context map is passed through to the integration as caller identity.
Amazon API Gateway lets you set throttling limits at the account, stage, method or route, and usage-plan levels. Explain what the rate and burst values in a throttle pair actually control, and which of those layers decides whether a given request is rejected.
basics
~20 sEach API Gateway throttle is a token bucket: burst is the bucket's capacity, rate is how many tokens per second refill it. A request is rejected by the first empty bucket it meets, checked from the most specific usage-plan limit outward to the region-wide account limit.