skip to content

You created an interface VPC endpoint for an AWS service, but your application inside the VPC still resolves that service to a public address and its traffic keeps leaving through the NAT gateway. What would you check?

level: middleimportance: should knowfreq 50%

answer

  1. the ENI exists but nobody dials it
  2. a checkbox at creation time
  3. two VPC attributes gate it
  4. which resolver is the client asking
  5. endpoint-specific name versus the regional name

basics

~20 s

Check that private DNS is enabled on the interface endpoint and that the VPC has both enableDnsSupport and enableDnsHostnames turned on. Without private DNS, the service's normal hostname keeps resolving to public addresses and only the endpoint-specific name uses the endpoint.

solid answer

~50 s

An interface endpoint changes nothing on its own — it is just an ENI with a private IP. Traffic only moves to it when the name the SDK dials resolves to that IP, which is what the **private DNS** option does: AWS associates a managed private hosted zone with your VPC so the regional service hostname answers with the endpoint's private addresses. So I check three things in order. First, is `PrivateDnsEnabled` actually true on the endpoint — it is easy to miss at creation. Second, does the VPC have `enableDnsSupport` and `enableDnsHostnames` set, since private DNS silently cannot work without them. Third, is the client resolving through the VPC resolver at all, or has it been pinned to an external resolver, or is the application hardcoding a public endpoint URL or a different Region in its client config. I would confirm with a lookup from the instance: a private RFC 1918 answer means DNS is right and the problem is elsewhere.

code

bash · 7 lines
bash
# Does the regional name resolve to a private address inside the VPC?
dig +short secretsmanager.eu-west-1.amazonaws.com

# Turn private DNS on for an existing interface endpoint
aws ec2 modify-vpc-endpoint \
  --vpc-endpoint-id vpce-0a1b2c3d4e5f \
  --private-dns-enabled

go deeper

for a junior

Know that an interface endpoint only gets used when the service's hostname resolves to its private IP, and that this is controlled by the endpoint's private DNS option.

for a middle

Walk the resolution chain out loud: endpoint private DNS enabled, VPC enableDnsSupport and enableDnsHostnames on, the client using the VPC resolver, and no hardcoded endpoint or Region in the application's client configuration.

for a senior

Demonstrate a diagnosis order rather than guesses — resolve the name on the host first to split DNS problems from security-group or endpoint-policy problems, and confirm the new path with Flow Logs and a falling NAT byte count.

for a principal

Treat enabling private DNS as a VPC-wide cutover with blast radius. Define how endpoints are rolled out and validated, how on-premises and hub resolution is served, and what the rollback is when a shared name suddenly points somewhere new.

## Why the endpoint alone changes nothing Creating an interface VPC endpoint provisions an elastic network interface with a private IP address in each subnet you selected. That is all it does. No traffic is redirected, no route is installed, and no application is reconfigured. The endpoint becomes useful only when something causes the client to *send packets to that private IP*, and in practice that something is DNS. ## What private DNS actually does Every interface endpoint gets endpoint-specific DNS names of the form `vpce-<id>-<hash>.<service>.<region>.vpce.amazonaws.com`, which always resolve to the endpoint's private IPs. Using them means changing every client's configured endpoint URL, which is invasive. The **private DNS** option avoids that. When enabled, AWS associates a managed private hosted zone with your VPC for the service's regional domain name, so a lookup of the ordinary hostname — say `kms.eu-west-1.amazonaws.com` — returns the endpoint ENI addresses instead of AWS's public addresses. Application code, the AWS SDKs, and the CLI need no change at all: they dial the same name they always did and land on the private path. This is the single most common reason an endpoint appears to "do nothing": the box was left unticked, so the name still resolves publicly and the traffic still exits through the NAT gateway, still billed as NAT processing. ## The VPC attributes Private DNS depends on the VPC's own DNS settings. Two VPC attributes matter: - `enableDnsSupport` — whether the Amazon-provided DNS resolver answers at the VPC base + 2 address. - `enableDnsHostnames` — whether instances get DNS hostnames. Both must be enabled before private DNS on an endpoint can function; AWS will refuse to enable the option, or it will not resolve as expected. VPCs created long ago by hand, or by templates that disabled hostnames, hit this regularly. ``` aws ec2 describe-vpc-attribute --vpc-id vpc-0a1b2c3d --attribute enableDnsSupport aws ec2 describe-vpc-attribute --vpc-id vpc-0a1b2c3d --attribute enableDnsHostnames ``` ## Where else the resolution can go wrong Even with private DNS on, the client must be *asking the VPC resolver*. If the instance's resolver has been pointed at a corporate DNS server, a public resolver, or an on-premises forwarder, the private hosted zone is invisible and the answer comes back public. The same happens for callers outside the VPC: a peered VPC or an on-premises host does not automatically see the managed private zone, which is why hub designs pair the endpoint with an explicitly created private hosted zone and, for on-premises resolution, a Route 53 Resolver inbound endpoint. Application configuration is the other culprit. An SDK client constructed with an explicit endpoint override, an environment variable pointing at a public URL, or a Region that differs from the endpoint's Region will bypass the endpoint no matter how DNS is configured — the endpoint is regional, so a client talking to `us-east-1` gets nothing from an endpoint in `eu-west-1`. ## Diagnosing it in order Start on the instance itself and resolve the service hostname. A private address in your subnet's CIDR proves DNS is doing its job, and any remaining failure is network or authorization: check the endpoint ENI's security group allows inbound TCP 443 from the client's group, and check the endpoint policy is not denying the call. A public address proves DNS is the problem, and then the order is endpoint `PrivateDnsEnabled`, VPC DNS attributes, the instance's resolver configuration, and finally the application's endpoint or Region override. To confirm the traffic path rather than infer it, VPC Flow Logs on the subnet will show flows to the endpoint's private IP once it is working, and the NAT gateway's byte metrics should drop for that service. If you want a decisive test before touching DNS at all, point one client at the endpoint-specific DNS name: if that succeeds while the regional name fails, the endpoint and its security group are healthy and the fault is unambiguously in name resolution. ## The subtlety worth mentioning Enabling private DNS is a VPC-wide change to how a well-known name resolves. Every workload in that VPC begins using the endpoint at once, including ones you did not intend to move, and if the endpoint's security group or endpoint policy is wrong those workloads break together. On a busy VPC it is safer to create the endpoint with private DNS off, validate through the endpoint-specific name, then flip the option — an approach that also gives you a fast rollback, since disabling it restores public resolution immediately.

  • How would you let an on-premises host resolve the same interface endpoint privately?
    The managed private hosted zone is only visible to the VPC, so on-premises resolvers cannot see it. You deploy a Route 53 Resolver inbound endpoint in the VPC and forward the service's domain from the on-premises DNS server to it, so those queries are answered by the VPC resolver and return the endpoint's private IPs.
  • You enable private DNS on a busy VPC's endpoint and several unrelated services immediately fail. What happened?
    Private DNS is VPC-wide: every workload in the VPC now resolves that service to the endpoint, not just the one you were migrating. If the endpoint's security group does not admit their traffic, or its endpoint policy does not permit their calls, they all break at once. Roll back by disabling private DNS, then widen the security group and policy first.

saying these in an interview costs you the question

  • Assumes creating the endpoint reroutes traffic automatically
  • Never checks the VPC's DNS attributes
  • Thinks a peered VPC inherits the private hosted zone
  • Blames routing when the name still resolves publicly
  • Ignores an SDK endpoint or Region override in config

context