skip to content

Instances in an EC2 Auto Scaling group pass their EC2 status checks, but the application on them returns errors and the load balancer has marked them unhealthy. Why has Auto Scaling not replaced them, and what would you change?

level: seniorimportance: should knowfreq 56%

answer

  1. two systems, two opinions of health
  2. the default only sees the hypervisor
  3. running is not the same as serving
  4. grace period versus real boot time
  5. replacing does not fix a bad build

basics

~20 s

The group's health check type is EC2, which only reflects hypervisor and instance status checks — a running instance with a dead application passes. Setting the health check type to ELB makes the group consume the load balancer's verdict and replace failing instances.

solid answer

~50 s

An Auto Scaling group's `HealthCheckType` defaults to `EC2`, which means it only considers the EC2 system and instance status checks — is the host reachable, is the virtual machine running. An instance whose application has crashed, deadlocked or is returning 500s still passes both, so the group sees a perfectly healthy fleet while the load balancer quietly drains every target. The fix is to set the health check type to `ELB` so the group also consumes the target group's health status, and to make sure `HealthCheckGracePeriod` is longer than the real boot-plus-warmup time — otherwise you swap one failure mode for a launch-and-terminate loop. I would also stop before turning it on and ask *why* the application is failing: if the cause is a bad deploy rather than an instance-specific fault, replacement just launches more broken instances, and the loop can shrink healthy capacity faster than it restores it.

code

bash · 10 lines
bash
# Let the load balancer's verdict drive replacement, with a realistic grace period
aws autoscaling update-auto-scaling-group \
  --auto-scaling-group-name web-asg \
  --health-check-type ELB \
  --health-check-grace-period 420

# During an incident caused by a bad build, stop the churn loop deliberately
aws autoscaling suspend-processes \
  --auto-scaling-group-name web-asg \
  --scaling-processes HealthCheck ReplaceUnhealthy

go deeper

for a junior

Know that EC2 status checks only tell you the machine is up, and that the load balancer's health check is a separate, application-level opinion.

for a middle

Explain what HealthCheckType EC2 and ELB each cause the group to act on, that ELB is additive to the status checks, and what the health check grace period is for.

for a senior

Demonstrate the diagnosis: recognise a churn loop from changing instance ids with a flat error rate, decide when automatic replacement helps and when suspending it is the right call, and size the grace period from measured time-to-ready.

for a principal

Own the definition of healthy across the estate. Argue what a readiness probe should and should not depend on, so that a shared dependency's outage cannot convert into fleet-wide self-destruction across every group at once.

## Two independent opinions about health There are two systems forming a view of these instances, and by default they do not talk to each other. **EC2 status checks** are the platform's own. The *system status check* covers the underlying host and network — a failure means AWS infrastructure is broken and usually requires a stop/start onto new hardware. The *instance status check* covers the virtual machine: has it booted, is its network stack responding. Neither knows anything about your process, your port, or your dependencies. A machine sitting at a login prompt with your service crashed passes both. **The load balancer's health check** probes the application through its target group and is the thing that decides whether traffic is sent. When it fails, the target is drained. ## What HealthCheckType actually selects The group's `HealthCheckType` decides which of those opinions can cause a *replacement*: - `EC2` (the default) — platform status checks only. - `ELB` — platform status checks **plus** the health status reported by the attached load balancer target groups. `ELB` is additive, not a replacement: a failed system status check still gets the instance replaced. So the described symptom is exactly the default behaviour. The load balancer stops sending traffic to the broken instances, the group counts them as healthy and keeps desired capacity satisfied with them, and available serving capacity silently collapses while the fleet size looks correct on every dashboard. The classic version of this incident is a fleet of ten instances where eight are drained and two are absorbing all production traffic until they fall over too. ```bash aws autoscaling update-auto-scaling-group \ --auto-scaling-group-name web-asg \ --health-check-type ELB \ --health-check-grace-period 300 ``` ## The grace period is not optional detail `HealthCheckGracePeriod` is the window after an instance launches during which health checks cannot mark it unhealthy — it exists because a booting instance legitimately fails an application probe for a while. Its default is 300 seconds as of 2025. If your application needs longer than the grace period to become ready, turning on `ELB` health checks creates a **launch loop**: the instance boots, is judged unhealthy before it is ready, is terminated, and a replacement starts the same journey. You burn money, churn the fleet, and never reach capacity. When someone reports "my group keeps replacing instances every few minutes", an undersized grace period is the first thing to check — and the honest fix is usually to measure real time-to-ready and set the grace period above the worst case, not to disable the health check. ## Replacement order and available capacity By default the group terminates an unhealthy instance and then launches its replacement, meaning capacity dips during the swap. An instance maintenance policy lets you express the tolerance explicitly with a minimum and maximum healthy percentage, so you can require the replacement to be launched *before* the old one goes away. On a small group that difference is significant: replacing one instance out of three is a third of your capacity gone for the length of a boot. ## The judgment part The technically correct answer is `HealthCheckType ELB`. The senior answer adds a caveat, because automatic replacement is only the right response to an *instance-specific* fault. Ask what all the failing instances have in common. If they are failing because a shared dependency is down, or because the version they are running is broken, then replacement is actively harmful: each new instance inherits the same defect, fails, and is replaced again. You get a churn loop that consumes launch capacity, floods your logs with fresh instance ids that make correlation harder, and — because instances are being terminated mid-boot — can leave you with less healthy capacity than if you had done nothing. The signature is that instance ids keep changing while the error rate does not. In that situation the levers are to suspend the `ReplaceUnhealthy` and `HealthCheck` process on the group while you diagnose, or to roll back what is being launched. Suspending is a deliberate, temporary act with an owner and a plan to resume — not a way of making a red dashboard turn green. ## What sits underneath One boundary worth being explicit about in an interview: the *content* of the probe — which path, how many consecutive failures, what timeout — is configuration on the target group, not on the Auto Scaling group. The group only consumes the resulting verdict. Getting that separation right in your answer signals that you understand the layering rather than having memorised a checkbox.

  • After switching to ELB health checks, the group starts replacing every instance a few minutes after launch. What is the likely cause?
    The health check grace period is shorter than the application's real time to ready, so instances are judged unhealthy while still warming up and are terminated before they can serve. Measure the worst-case time from launch to a passing probe and set the grace period comfortably above it, rather than turning the health check back off.
  • During an incident every instance is failing its application health check because a downstream dependency is down. Should you leave automatic replacement enabled?
    No. Every replacement inherits the same failure, so you churn the fleet, lose capacity during each boot, and make correlation harder as instance ids keep changing. Suspend the HealthCheck and ReplaceUnhealthy processes while you work the dependency, with an explicit owner and a plan to resume them.
  • Which health signal still replaces an instance even when the health check type is left at EC2?
    The EC2 status checks themselves. A failed system status check means the underlying host or its network is broken, and a failed instance status check means the virtual machine is not responding. Either one marks the instance unhealthy and gets it replaced, regardless of what the load balancer thinks.

saying these in an interview costs you the question

  • Thinks a running instance is by definition a healthy instance
  • Assumes the group reads target group health by default
  • Enables ELB health checks without checking the grace period
  • Says replacing instances fixes an application-level failure
  • Confuses the target group's probe settings with the group's health check type

context