In AWS Elastic Load Balancing, what is a target group, and what happens to a registered target once it starts failing that target group's health check?
answer
- the listener's forwarding destination
- membership plus a liveness definition
- consecutive failures, not one
- interval times threshold equals detection time
- state is per target group, not per instance
basics
~20 sA target group is the named set of backends an Elastic Load Balancing listener forwards to, plus the health check run against each member. A target that fails enough consecutive checks is marked unhealthy and stops receiving new requests.
solid answer
~50 sA target group is the routing destination in Elastic Load Balancing: a listener (or a listener rule) forwards to a target group, and the target group holds the registered targets — EC2 instances, IP addresses, a Lambda function, or another ALB behind an NLB — along with the protocol and port used to reach them and a single health-check definition. ELB probes each registered target independently on the configured health-check protocol, port and path, and compares the response against the `Matcher`. A target that fails `UnhealthyThresholdCount` consecutive probes flips to `unhealthy` and the load balancer stops sending it new requests; once it passes `HealthyThresholdCount` probes in a row it returns to `healthy`. Health is tracked per target group, so the same instance can be healthy in one and unhealthy in another. `aws elbv2 describe-target-health` shows each target's state and reason code.
code
bash · 4 linesaws elbv2 describe-target-health \
--target-group-arn arn:aws:elasticloadbalancing:eu-west-1:111122223333:targetgroup/api/73e2d6bc24d8a067 \
--query 'TargetHealthDescriptions[].{Target:Target.Id,Port:Target.Port,State:TargetHealth.State,Reason:TargetHealth.Reason}' \
--output tablego deeper
Be able to say that a listener forwards to a target group, that the target group holds the registered backends plus one health check, and that failing that check consecutively stops new traffic to that target.
Explain the interval, timeout, thresholds and matcher, and do the arithmetic out loud: detection time is roughly interval times unhealthy threshold, which is how long a dead backend keeps taking requests.
Show that you read the reason code before touching config, and that you know health feeds Auto Scaling replacement and the HealthyHostCount alarm, so tightening thresholds has blast radius beyond routing.
Own the tradeoff between fast ejection and stability: aggressive thresholds turn a latency spike into a fleet-wide eviction, and coupling ASG replacement to load-balancer health can turn a dependency blip into a replacement storm.
## Where a target group sits An Application, Network or Gateway Load Balancer is made of three layers. The **load balancer** owns DNS name, subnets and security posture. A **listener** owns a protocol and port the clients connect to. A **target group** owns the backends and the definition of what "alive" means for them. Listener rules forward to target groups, so the target group is the unit you swap when you cut over a deployment, and it is the unit whose health you watch. A target group carries, at minimum: - a **target type** — `instance`, `ip`, `lambda`, or `alb` - a **protocol and port** used to reach targets (HTTP/HTTPS for an ALB, TCP/UDP/TLS for an NLB) - the **VPC** the targets live in (not applicable to `lambda`) - one **health-check** definition: protocol, port, path, interval, timeout, thresholds and matcher - **attributes** such as `deregistration_delay.timeout_seconds` and `slow_start.duration_seconds` Registering a target does not copy it anywhere; it adds an entry (instance ID plus port, or IP plus port) that the load balancer nodes start probing. ## What the health check actually does Each load balancer node runs the health check independently against every registered target, so a target sees several probes per interval — one per enabled Availability Zone. For an HTTP health check the node opens a connection to the health-check port (`traffic-port` by default, meaning the same port the target is registered on), issues a `GET` on the health-check path with the user agent `ELB-HealthChecker/2.0`, and compares the status code against the `Matcher` — `HttpCode` such as `200`, `200-299` or `200,202` for HTTP target groups, `GrpcCode` for gRPC ones. The state machine is deliberately hysteretic so a single blip does not eject a backend: - `initial` — registered, first probes in flight - `healthy` — passed `HealthyThresholdCount` consecutive probes - `unhealthy` — failed `UnhealthyThresholdCount` consecutive probes - `draining` — deregistering, finishing in-flight requests - `unused` — registered but not receiving traffic (for example the target group is not attached to a listener) As of 2025 an ALB target group defaults to a 30-second interval, a 5-second timeout, an unhealthy threshold of 2 and a healthy threshold of 5 — but treat those as values to confirm in the console or API, not to recite. What matters is the arithmetic: **detection time is roughly interval × unhealthy threshold**, so a 30-second interval with a threshold of 2 means up to a minute of requests still hitting a dead backend. Tightening either number shortens the outage window and raises the risk of ejecting a target that was merely slow for one probe. ## Consequences of the unhealthy state An unhealthy target receives **no new requests**. Requests already in flight are not killed. The target stays registered and keeps being probed, so recovery is automatic. Because health lives in the target group, a target registered in two target groups is judged twice, independently — a common setup when one ALB serves public traffic and an internal ALB serves service-to-service calls. Health also feeds other systems. An Auto Scaling group whose health-check type is set to `ELB` will terminate and replace an instance the target group reports unhealthy, after the health-check grace period expires. CloudWatch publishes `HealthyHostCount` and `UnHealthyHostCount` per target group and Availability Zone, which is the metric most teams alarm on. ```bash aws elbv2 describe-target-health --target-group-arn "$TG_ARN" ``` The `TargetHealth.Reason` field is the part beginners skip and seniors read first: `Elb.InitialHealthChecking`, `Target.Timeout`, `Target.ResponseCodeMismatch`, `Target.NotInUse` and `Target.DeregistrationInProgress` each point at a completely different cause. ## The point of the design Splitting membership ("who is behind this name") from liveness ("who can serve right now") is what makes rolling deploys, scale-out and instance replacement invisible to clients. The load balancer never needs to be reconfigured; targets are added and removed, and the health check decides when each one starts and stops carrying traffic.
- With a 30-second interval and an unhealthy threshold of 2, how long can a broken target keep receiving requests?Up to roughly a minute: the load balancer needs two consecutive failed probes 30 seconds apart, and the failure may occur just after a successful probe. Shortening the interval or the threshold cuts that window but makes the target group more sensitive to a single slow response, which can eject healthy but momentarily loaded targets.
- Can one EC2 instance be registered in more than one target group at the same time?Yes. The same instance can back several target groups — for example a public ALB and an internal one — and each target group evaluates its own health check independently, so the instance can be healthy in one and unhealthy in another. An Auto Scaling group can attach multiple target groups so replacements register everywhere automatically.
- What does the target state `unused` mean?The target is registered and reachable but the load balancer has no reason to send it traffic: typically the target group is not referenced by any listener or rule, the target is in an Availability Zone the load balancer has not enabled, or the target group is attached to a load balancer in a stopped state. It is a wiring problem, not a target problem.
saying these in an interview costs you the question
- Thinking one failed response immediately removes a target
- Believing health state is a property of the instance, not the target group
- Assuming an unhealthy target has its in-flight requests killed
- Confusing the target group with the listener that forwards to it
- Never looking at the reason code in describe-target-health