skip to content

Should the endpoint an ALB target group health-checks verify the service's downstream dependencies, such as its database or cache? Argue the tradeoff and say what you would standardise on.

level: principalimportance: should knowfreq 35%

answer

  1. routing decision versus alerting decision
  2. would every target fail at once
  3. shared dependency, correlated failure
  4. replacement storm on cold caches
  5. the load balancer fails open anyway

basics

~20 s

Generally no. A dependency check makes every target fail at the same instant when that dependency blips, turning a partial outage into a total one and potentially triggering fleet-wide instance replacement. Keep the load-balancer check shallow and expose dependency status on a separate endpoint for alarms.

solid answer

~50 s

Keep the target-group health check shallow — it should answer "can this process serve a request right now" — and put dependency verification on a separate endpoint that feeds CloudWatch alarms and dashboards rather than routing. A deep check creates correlated failure: every target queries the same database, so one blip marks the entire target group unhealthy simultaneously, taking down endpoints that never needed that dependency. It gets worse downstream: if the Auto Scaling group's health-check type is `ELB`, it will start terminating and replacing every instance, so a five-minute database hiccup becomes a fleet rebuild with cold caches. Note also that an ALB **fails open** — when every target in a target group is unhealthy it forwards to them anyway — so a fleet-wide deep-check failure does not even buy you a clean error. The middle ground is graceful degradation: report healthy, serve what you can, and alarm loudly on the dependency.

go deeper

for a junior

Know that a health-check endpoint should be cheap and unauthenticated, and that checking the database inside it can make every server go unhealthy at the same time.

for a middle

Explain the correlated-failure mechanism concretely — one shared dependency, one shared verdict — and where dependency status belongs instead: a separate endpoint feeding metrics and alarms.

for a senior

Show the operational chain you would prevent: deep check fails everywhere, Auto Scaling with health-check type ELB replaces the fleet, cold instances hammer the recovering dependency. Name the grace period and the degraded-mode response as the controls.

for a principal

Own the standard across services: a stated independence test for what may appear in a health check, degraded-mode behaviour as a design requirement, separation of routing signals from alerting signals, and the exceptions you allow for cell or shard architectures.

## Two different questions wear the same name "Health check" conflates two questions that want different answers: 1. **Should the load balancer send this process traffic?** A routing decision, made many times a minute, whose only useful comparison is *between targets*. 2. **Is this service able to do its job?** An alerting decision, whose audience is a human and whose right response is a page, a runbook, or a failover. The target group answers only the first. It has no capacity to answer the second usefully, because every target will give it the same answer at the same moment. ## Why depth creates correlated failure Suppose `/health` opens a connection to the primary database. The database becomes slow for ninety seconds. Every target's probe times out. Within one or two intervals the whole target group is unhealthy — and consider what you gained: nothing. There is no healthier target to shift traffic to, because the sick dependency is shared. What you lost is substantial: - **Endpoints that did not need the database now fail too.** Cached reads, static responses, and degraded modes all stop being served. - **Auto Scaling may start replacing the fleet.** With health-check type `ELB`, instances failing target-group checks past the grace period are terminated and replaced. Replacements boot with cold caches and empty connection pools, and hammer the recovering dependency — the textbook shape of a failure that outlives its trigger. - **The ALB fails open anyway.** When a target group contains only unhealthy targets, the load balancer routes requests to them regardless. So the deep check did not produce a clean, fast error for clients; it produced instance churn plus the same errors. The general principle: a health signal is useful in proportion to how *independently* targets can fail it. Process-local conditions — the listener is up, the thread pool is not wedged, the process finished booting — differ between targets and are exactly what routing should react to. Shared dependencies do not differ, so they belong in alarms, not in routing. ## The case for some depth The purist "return 200 unconditionally" position also fails, in a subtler way. A handler that answers from a dedicated thread and touches nothing will report healthy while the application's real worker pool is completely saturated — a target that accepts requests and never answers them, which the load balancer keeps feeding. Useful shallow checks therefore share fate with real request handling: served by the same server and thread pool, and reflecting local readiness state such as "finished warming", "not shutting down", "pool not exhausted". There is also a narrow legitimate case for dependency awareness: a **target-local** dependency that genuinely differs between instances — a broken local disk, a sidecar process that died, a per-instance connection pool that has been exhausted for minutes while other instances are fine. Those satisfy the independence test, and failing them out is correct. ## The pattern to standardise on - **Routing path** (the target group's health-check path): unauthenticated, cheap, no shared-dependency calls, served by the normal request path, and honest about local readiness and shutdown. During graceful shutdown it should start failing *before* the process stops accepting connections, so draining begins in the right order. - **Diagnostic path**: a separate, richer endpoint reporting per-dependency status, scraped on a slower cadence and turned into CloudWatch alarms. It never influences routing. - **Degraded mode in the application**: when a dependency is unavailable, serve what remains possible and surface the degradation in responses and metrics, instead of turning the whole target off. - **Decouple replacement from routing.** Set the Auto Scaling health-check grace period generously, and be deliberate about using health-check type `ELB` at all — reserve fleet replacement for signals that are truly per-instance. - **Alarm on the right metric.** `UnHealthyHostCount` per Availability Zone catches the localised case; end-user error rate and latency catch the shared-dependency case. Alarming only on health-check state is how a deep check becomes the alerting system by accident. ## Where the argument can go the other way If a target group fronts several independent stacks that each own their *own* dependency — cells or shards where instance A talks only to shard A — then dependency failure is per-target after all, and checking it is right. The test is always the same: **would this check fail on all targets at once?** If yes, it is an alarm. If no, it can be a health check.

  • What exactly does an ALB do when every target in a target group is unhealthy?
    It fails open and forwards requests to the unhealthy targets anyway rather than refusing traffic, on the reasoning that a possibly-degraded backend beats a guaranteed error. That is worth knowing before designing around health checks: a fleet-wide failure does not produce a clean load-balancer error, it produces whatever the targets do.
  • If the check must not call the database, how do you catch a target whose worker pool is wedged?
    By making the check share fate with real traffic: serve it from the same server and thread pool as normal requests rather than a privileged side channel, and reflect local readiness — warmed, not shutting down, pool not exhausted. A saturated target then times out on its own probe, which is a genuinely per-target signal.
  • How does the Auto Scaling health-check type change the stakes of this decision?
    With health-check type ELB, targets failing the load balancer's check past the grace period are terminated and replaced. A deep check therefore converts a shared dependency blip into a fleet rebuild with cold caches that hammers the recovering dependency. Reserve ELB-driven replacement for signals that genuinely differ per instance, and keep the grace period comfortably longer than a full boot.
  • Is there any architecture where a dependency check in the health endpoint is correct?
    Yes — when the dependency is not shared. In a cell or shard-per-instance design where each target talks only to its own datastore, failing that dependency is a per-target condition and routing around it is exactly right. The test is whether the check would fail on every target simultaneously; if it would, it is an alarm rather than a health check.

saying these in an interview costs you the question

  • Believing a deep check gives the load balancer somewhere better to route
  • Assuming an ALB returns a clean error once all targets are unhealthy
  • Returning 200 unconditionally from a privileged side channel
  • Coupling Auto Scaling replacement to a shared-dependency signal
  • Using health-check state as the primary alerting mechanism

context