Should the endpoint an ALB target group health-checks verify the service's downstream dependencies, such as its database or cache? Argue the tradeoff and say what you would standardise on.
answer
- routing decision versus alerting decision
- would every target fail at once
- shared dependency, correlated failure
- replacement storm on cold caches
- the load balancer fails open anyway
basics
~20 sGenerally no. A dependency check makes every target fail at the same instant when that dependency blips, turning a partial outage into a total one and potentially triggering fleet-wide instance replacement. Keep the load-balancer check shallow and expose dependency status on a separate endpoint for alarms.
solid answer
~50 sKeep the target-group health check shallow — it should answer "can this process serve a request right now" — and put dependency verification on a separate endpoint that feeds CloudWatch alarms and dashboards rather than routing. A deep check creates correlated failure: every target queries the same database, so one blip marks the entire target group unhealthy simultaneously, taking down endpoints that never needed that dependency. It gets worse downstream: if the Auto Scaling group's health-check type is `ELB`, it will start terminating and replacing every instance, so a five-minute database hiccup becomes a fleet rebuild with cold caches. Note also that an ALB **fails open** — when every target in a target group is unhealthy it forwards to them anyway — so a fleet-wide deep-check failure does not even buy you a clean error. The middle ground is graceful degradation: report healthy, serve what you can, and alarm loudly on the dependency.
go deeper
Know that a health-check endpoint should be cheap and unauthenticated, and that checking the database inside it can make every server go unhealthy at the same time.
Explain the correlated-failure mechanism concretely — one shared dependency, one shared verdict — and where dependency status belongs instead: a separate endpoint feeding metrics and alarms.
Show the operational chain you would prevent: deep check fails everywhere, Auto Scaling with health-check type ELB replaces the fleet, cold instances hammer the recovering dependency. Name the grace period and the degraded-mode response as the controls.
Own the standard across services: a stated independence test for what may appear in a health check, degraded-mode behaviour as a design requirement, separation of routing signals from alerting signals, and the exceptions you allow for cell or shard architectures.
## Two different questions wear the same name "Health check" conflates two questions that want different answers: 1. **Should the load balancer send this process traffic?** A routing decision, made many times a minute, whose only useful comparison is *between targets*. 2. **Is this service able to do its job?** An alerting decision, whose audience is a human and whose right response is a page, a runbook, or a failover. The target group answers only the first. It has no capacity to answer the second usefully, because every target will give it the same answer at the same moment. ## Why depth creates correlated failure Suppose `/health` opens a connection to the primary database. The database becomes slow for ninety seconds. Every target's probe times out. Within one or two intervals the whole target group is unhealthy — and consider what you gained: nothing. There is no healthier target to shift traffic to, because the sick dependency is shared. What you lost is substantial: - **Endpoints that did not need the database now fail too.** Cached reads, static responses, and degraded modes all stop being served. - **Auto Scaling may start replacing the fleet.** With health-check type `ELB`, instances failing target-group checks past the grace period are terminated and replaced. Replacements boot with cold caches and empty connection pools, and hammer the recovering dependency — the textbook shape of a failure that outlives its trigger. - **The ALB fails open anyway.** When a target group contains only unhealthy targets, the load balancer routes requests to them regardless. So the deep check did not produce a clean, fast error for clients; it produced instance churn plus the same errors. The general principle: a health signal is useful in proportion to how *independently* targets can fail it. Process-local conditions — the listener is up, the thread pool is not wedged, the process finished booting — differ between targets and are exactly what routing should react to. Shared dependencies do not differ, so they belong in alarms, not in routing. ## The case for some depth The purist "return 200 unconditionally" position also fails, in a subtler way. A handler that answers from a dedicated thread and touches nothing will report healthy while the application's real worker pool is completely saturated — a target that accepts requests and never answers them, which the load balancer keeps feeding. Useful shallow checks therefore share fate with real request handling: served by the same server and thread pool, and reflecting local readiness state such as "finished warming", "not shutting down", "pool not exhausted". There is also a narrow legitimate case for dependency awareness: a **target-local** dependency that genuinely differs between instances — a broken local disk, a sidecar process that died, a per-instance connection pool that has been exhausted for minutes while other instances are fine. Those satisfy the independence test, and failing them out is correct. ## The pattern to standardise on - **Routing path** (the target group's health-check path): unauthenticated, cheap, no shared-dependency calls, served by the normal request path, and honest about local readiness and shutdown. During graceful shutdown it should start failing *before* the process stops accepting connections, so draining begins in the right order. - **Diagnostic path**: a separate, richer endpoint reporting per-dependency status, scraped on a slower cadence and turned into CloudWatch alarms. It never influences routing. - **Degraded mode in the application**: when a dependency is unavailable, serve what remains possible and surface the degradation in responses and metrics, instead of turning the whole target off. - **Decouple replacement from routing.** Set the Auto Scaling health-check grace period generously, and be deliberate about using health-check type `ELB` at all — reserve fleet replacement for signals that are truly per-instance. - **Alarm on the right metric.** `UnHealthyHostCount` per Availability Zone catches the localised case; end-user error rate and latency catch the shared-dependency case. Alarming only on health-check state is how a deep check becomes the alerting system by accident. ## Where the argument can go the other way If a target group fronts several independent stacks that each own their *own* dependency — cells or shards where instance A talks only to shard A — then dependency failure is per-target after all, and checking it is right. The test is always the same: **would this check fail on all targets at once?** If yes, it is an alarm. If no, it can be a health check.
- What exactly does an ALB do when every target in a target group is unhealthy?It fails open and forwards requests to the unhealthy targets anyway rather than refusing traffic, on the reasoning that a possibly-degraded backend beats a guaranteed error. That is worth knowing before designing around health checks: a fleet-wide failure does not produce a clean load-balancer error, it produces whatever the targets do.
- If the check must not call the database, how do you catch a target whose worker pool is wedged?By making the check share fate with real traffic: serve it from the same server and thread pool as normal requests rather than a privileged side channel, and reflect local readiness — warmed, not shutting down, pool not exhausted. A saturated target then times out on its own probe, which is a genuinely per-target signal.
- How does the Auto Scaling health-check type change the stakes of this decision?With health-check type ELB, targets failing the load balancer's check past the grace period are terminated and replaced. A deep check therefore converts a shared dependency blip into a fleet rebuild with cold caches that hammers the recovering dependency. Reserve ELB-driven replacement for signals that genuinely differ per instance, and keep the grace period comfortably longer than a full boot.
- Is there any architecture where a dependency check in the health endpoint is correct?Yes — when the dependency is not shared. In a cell or shard-per-instance design where each target talks only to its own datastore, failing that dependency is a per-target condition and routing around it is exactly right. The test is whether the check would fail on every target simultaneously; if it would, it is an alarm rather than a health check.
saying these in an interview costs you the question
- Believing a deep check gives the load balancer somewhere better to route
- Assuming an ALB returns a clean error once all targets are unhealthy
- Returning 200 unconditionally from a privileged side channel
- Coupling Auto Scaling replacement to a shared-dependency signal
- Using health-check state as the primary alerting mechanism