skip to content

Amazon Route 53's health checkers probe endpoints from the public internet, so they cannot reach an internal load balancer that has only private addresses. How do you still drive Route 53 failover from that endpoint's health?

level: seniorimportance: should knowfreq 40%

answer

  1. three types, only one probes
  2. checkers sit on the public internet
  3. alarm state instead of a request
  4. reverse the direction of observation
  5. decide what INSUFFICIENT_DATA means

basics

~20 s

Use a CloudWatch alarm health check instead of an endpoint health check: Route 53 reads the alarm's state rather than probing anything. Point the alarm at a metric that reflects the private endpoint, such as its unhealthy host count.

solid answer

~50 s

Route 53 offers three kinds of health check, and only one of them probes an endpoint. The second kind watches a CloudWatch alarm — Route 53 treats the check as unhealthy when the alarm is in ALARM state — and that is the standard answer for anything without a public address, because the metric is published from inside your account rather than fetched from outside. For an internal Application Load Balancer you would alarm on its `UnHealthyHostCount` or on healthy hosts dropping to zero, then attach that health check to the record. The third kind is a calculated health check, which combines child health checks and reports healthy only when at least a configured number of children are. One detail worth setting deliberately: `InsufficientDataHealthStatus` decides what the check reports when the alarm has no data, and defaulting that to unhealthy will fail you over during a metric gap.

go deeper

for a junior

Know that Route 53 health checks probe from the public internet and therefore cannot reach private endpoints, and that a CloudWatch alarm can stand in as the health signal instead.

for a middle

Name the three health check types — endpoint, CloudWatch alarm, calculated — and explain that the alarm variant reverses the direction: your account publishes a metric outward rather than Route 53 probing inward.

for a senior

Choose a metric that genuinely represents servability, set the insufficient-data behaviour deliberately, and be explicit about what end-to-end coverage you gave up by not probing the real path.

for a principal

Own what the failover signal means. Decide whether DNS should be driven by inferred component health at all, versus an explicit recovery control that a human or automation flips, and weigh the cost of false failovers against slow ones for a stateful multi-region system.

## The three health check types Route 53 health checks come in three forms, and knowing all three is what this question tests: 1. **Endpoint health checks** — Route 53's own checkers, spread across AWS regions, make requests to an IP address or domain name over HTTP, HTTPS or TCP. HTTP and HTTPS variants can additionally require a string to appear in the response body, so the check can assert the application is actually working rather than merely that a socket opened. 2. **CloudWatch alarm health checks** — the check has no probe at all. It watches a CloudWatch alarm you name and reports unhealthy while that alarm is in the `ALARM` state. 3. **Calculated health checks** — the check has no probe either. It aggregates other health checks (its children) and is healthy when at least a configured number of them are healthy. There is also a health check type tied to Route 53 Application Recovery Controller routing controls, used when you want an explicit, human- or automation-flipped switch rather than an inferred one. ## Why the endpoint check cannot help you here The checkers live on the public internet. They resolve and connect exactly as any outside client would, which is why an endpoint check is a genuine end-to-end test — it exercises DNS, routing, the firewall path, and TLS. That strength is also the constraint: an endpoint with only private addresses inside a VPC is unreachable from where the checkers sit, and no security group rule can change the fact that the address is not routable from the internet. So the direction of the flow has to be reversed. Rather than something outside reaching in to observe health, something inside publishes health outward, into CloudWatch, and Route 53 reads it there. ## Building the CloudWatch alarm check For an internal Application Load Balancer, the metrics are already there: the load balancer publishes healthy and unhealthy host counts per target group, so an alarm on unhealthy hosts rising, or on healthy hosts falling to zero, is a direct statement about whether the endpoint can serve. For an application with no such built-in metric, you publish your own — a small scheduled function inside the VPC that calls the private endpoint and emits a custom metric, or the application emitting a heartbeat metric itself. Then create the health check of the CloudWatch-metric type, referencing the alarm, and attach its id to the failover, weighted, or latency records you want it to control. The field that gets forgotten is `InsufficientDataHealthStatus`. A CloudWatch alarm can be in `INSUFFICIENT_DATA` — no data points arrived in the evaluation window — and you must tell Route 53 what to report in that case: healthy, unhealthy, or the last known status. Choosing unhealthy means a metric-publishing hiccup triggers a regional failover; choosing healthy means a genuinely dead endpoint that stops emitting anything looks fine. Last-known-status is often the least-bad default, and the deeper fix is to alarm on a metric that keeps reporting rather than one that simply stops. ## Where calculated checks fit Calculated checks let you express "this site is healthy when the API *and* the database layer are healthy", or "when at least two of three nodes are healthy", by combining children and setting the threshold. They can also be inverted, so a check reports the opposite of what its children say — occasionally useful for building a deliberate drain switch. Because they aggregate rather than probe, they combine freely with both endpoint and CloudWatch-alarm children, which is how a private-endpoint check gets folded into a larger picture of a region's health. ## Judgment: what the check should measure The substitution costs you something real. An endpoint check tests the path a user actually takes; a CloudWatch alarm tests what your own telemetry says about a component. A region can pass every internal metric while being unreachable from the internet — a broken route, an expired certificate, a misconfigured security group — and the alarm-based check will never notice. So when the public entry point *is* public, health-check it directly, and use alarm-based checks for the private components behind it. Where the whole path is private, consider a synthetic client inside the VPC that exercises the endpoint the way a user would and publishes the result as a metric, rather than alarming on a component-level counter. The goal is that the signal driving DNS failover means "users can be served here", not merely "this process is running".

  • What should a CloudWatch-alarm health check report when the alarm is in INSUFFICIENT_DATA?
    That is yours to set through InsufficientDataHealthStatus, and the choice has real consequences. Unhealthy means a gap in metric publishing fails your region over; healthy means a dead endpoint that stops emitting looks fine. Last known status is often the safest, and the stronger fix is alarming on a metric that keeps reporting a value rather than one that vanishes.
  • What do you lose by replacing an endpoint health check with a CloudWatch alarm health check?
    The end-to-end property. An endpoint check traverses the internet, DNS, routing, firewalls and TLS exactly as a user does, so it catches an expired certificate or a broken route. An alarm check only reflects what your own telemetry says about a component, so a region can look perfectly healthy internally while being unreachable from outside.
  • When would you use a calculated health check?
    When healthiness is a composite: a region counts as up only if the API and its data layer are both up, or if at least two of three nodes are serving. A calculated check aggregates child checks with a threshold, can mix endpoint and alarm children, and can be inverted — which is one way to build an explicit drain switch for a region.

saying these in an interview costs you the question

  • Suggests opening the internal load balancer to the internet for checks
  • Thinks a security group rule can let Route 53 checkers into a VPC
  • Unaware that Route 53 can watch a CloudWatch alarm
  • Leaves InsufficientDataHealthStatus at whatever it defaults to
  • Treats an internal component metric as proof users can reach the region

context