An Application Load Balancer is failing requests and every target in its target group shows unhealthy, yet curling the health-check path directly against a target from inside the VPC returns 200. How do you find the cause?
answer
- read the reason code first
- timeout versus code mismatch
- the probe comes from the load balancer
- the probe carries no hostname you expect
- grep the access log for ELB-HealthChecker
basics
~20 sRead the reason code from describe-target-health first. Target.Timeout means the probe never arrived — usually the target's security group does not allow the health-check port from the load balancer. Target.ResponseCodeMismatch means it arrived but the status code fell outside the matcher.
solid answer
~50 sStart with `aws elbv2 describe-target-health` and read `TargetHealth.Reason`, because it splits the problem in two. `Target.Timeout` means the probe never got an answer: check that the target's security group allows the health-check port **from the load balancer's security group**, that network ACLs allow the return traffic, that the health-check port is the port the app actually listens on rather than a stale override, and that the health-check protocol matches (probing HTTPS against a plaintext listener fails silently). `Target.ResponseCodeMismatch` means the target answered but the code was outside the `Matcher` — a `/` path that 302s to a login page, a 404 because the app is name-based virtual-hosted and the probe carries no hostname you recognise, or a matcher of `200` against an app returning `204`. Your manual curl succeeded because it came from a different source and carried different headers.
code
bash · 3 linescurl -sS -o /dev/null -w 'status=%{http_code} time=%{time_total}\n' \
-H 'User-Agent: ELB-HealthChecker/2.0' \
http://10.0.12.34:8080/healthzgo deeper
Know where to look: describe-target-health shows a state and a reason per target, and the two common reasons are a timed-out probe and a status code outside the matcher.
Walk the two branches confidently — reachability (security group source, port, protocol, network ACLs) versus contract (path, redirect, matcher, hostname) — and say how you would reproduce the probe exactly.
Show that you diagnose from evidence rather than by changing settings: reason code, then the target's own access log for ELB-HealthChecker entries, then a targeted fix, and explain why ejecting targets under load makes an incident worse.
Own the standard: a dedicated unauthenticated health path, an explicit matcher, security groups referenced by group rather than by address range, and timeouts derived from measured endpoint latency, applied consistently so this class of incident stops recurring.
## Split the problem before touching anything "All targets unhealthy" has exactly two shapes, and the API tells you which one you have: ```bash aws elbv2 describe-target-health --target-group-arn "$TG_ARN" \ --query 'TargetHealthDescriptions[].[Target.Id,TargetHealth.State,TargetHealth.Reason,TargetHealth.Description]' ``` The reason codes are the diagnosis: | Reason | Meaning | Where to look | |---|---|---| | `Elb.InitialHealthChecking` | probes are still in flight | wait one interval | | `Target.Timeout` | no response within the timeout | reachability, wrong port, slow endpoint | | `Target.ResponseCodeMismatch` | answered, wrong status | path, matcher, redirects | | `Target.FailedHealthChecks` | generic failure, often connection refused | app not listening yet | | `Target.NotInUse` | not receiving traffic at all | target group not attached to a listener, or AZ not enabled | | `Target.DeregistrationInProgress` | draining | the deploy, not the app | Everything below hangs off that split. ## Branch one: the probe never arrives The overwhelmingly common cause is the security group. For an ALB the health check originates from the load balancer's own elastic network interfaces, so the **target's** security group must allow inbound on the health-check port with the load balancer's security group as the source. A rule that allows only a bastion's address, or that allows the application port while the health check probes a separate management port, produces exactly this symptom — and a curl from a bastion inside the VPC still works, because the bastion's source address is allowed and the load balancer's is not. Other members of this branch: - **Wrong port.** The health-check port is `traffic-port` by default, meaning the port each target was registered on. An explicit override left over from an earlier design probes a port nothing listens on. - **Protocol mismatch.** An HTTPS health check against a target that only speaks plaintext, or vice versa, times out rather than returning an error. - **Network ACLs.** Unlike security groups, network ACLs are stateless, so the reply from the target needs an explicit outbound allow on ephemeral ports. - **Subnet and zone wiring.** If a target sits in an Availability Zone the load balancer has not enabled, it is `unused`, not `unhealthy` — a different reason code and a different fix. - **Timeout too tight.** A health-check endpoint that occasionally takes longer than the health-check timeout will time out even though a leisurely manual curl always succeeds. ## Branch two: the probe arrives and is rejected Here the target answered, so reachability is fine and the failure is in the contract. The probe is a plain HTTP request with the user agent `ELB-HealthChecker/2.0`, sent to the target's address — it carries no cookies, no authentication and not your public hostname. That produces several classic traps: - The health-check path is `/`, and `/` redirects to `/login` with a **302**, which is outside a `200` matcher. - The app is virtual-hosted by name and returns **404** for any host it does not recognise. - The endpoint requires authentication and returns **401** or **403** to an unauthenticated prober. - The app returns **204 No Content** from a health endpoint while the matcher says `200`. The fastest confirmation is to reproduce the probe rather than your own curl — same path, same port, same headers — and then to look at the target's own access log for `ELB-HealthChecker/2.0` entries. If those entries are absent, you are in branch one; if they are present with a non-matching status, you are in branch two and the log line names the status. ## Widen the matcher or fix the endpoint Both are legitimate. `--matcher HttpCode=200-399` accommodates a redirect, but a redirect on a health path usually means you pointed the check at the app's front door instead of a purpose-built endpoint. The better outcome is a dedicated path that is unauthenticated, cheap, and host-agnostic. ## The symptom that is not this problem If targets are healthy and clients still get 503, the target group probably has no registered targets at all — a different failure. And note the ALB's fail-open behaviour: when a target group's targets are *all* unhealthy the load balancer keeps forwarding to them rather than refusing, so the client-visible symptom is 502s and timeouts from the app itself, not a clean load-balancer error.
- Why does a curl from a bastion host succeed while the load balancer's probe times out?Because the source differs. The target's security group most likely allows the bastion's address but not the load balancer's security group, so only the probe is dropped. The manual curl also carries your own hostname and path, which can hide a matcher or virtual-host problem that the probe runs straight into.
- The health check passes but clients still get 503 from the ALB. What would you check?A 503 from the load balancer itself typically means the target group has no registered targets, or the rule matched a target group that is empty. Check that the listener rule forwards where you think, that the Auto Scaling group or ECS service is actually registering targets, and that the load balancer has those targets' Availability Zones enabled.
- How would you make the health check less likely to eject targets during a load spike?Give the endpoint its own cheap path that does not queue behind expensive application work, raise the health-check timeout above the endpoint's realistic worst case, and increase the unhealthy threshold so one slow probe does not eject a target. Ejecting targets under load removes capacity precisely when you need it, which is how a spike turns into an outage.
- What does a reason code of Target.NotInUse tell you?That the target is registered but the load balancer has no reason to send it traffic — usually the target group is not referenced by any listener or rule, or the target sits in an Availability Zone the load balancer has not enabled. It is a wiring problem, so widening the matcher or fixing the app will change nothing.
saying these in an interview costs you the question
- Widening the matcher before reading the reason code
- Allowing a bastion address instead of the load balancer's security group
- Assuming a working manual curl proves reachability from the load balancer
- Forgetting the probe sends no hostname the app recognises
- Confusing unused with unhealthy