Every rolling deployment of an ECS service behind a load balancer produces a burst of 5xx errors, and some freshly started tasks are killed and replaced before they serve anything. Which ECS service settings explain that, and how do you make deployments clean?
answer
- two races: coming in, going out
- warm-up time versus first health check
- draining and SIGTERM run in parallel
- percentages bound the capacity dip
- let a bad deploy roll itself back
basics
~20 sNew tasks are being health-checked or killed before they are ready, and old tasks are stopped before their connections drain. Set healthCheckGracePeriodSeconds, align the container stopTimeout with the target group deregistration delay, and handle SIGTERM gracefully.
solid answer
~50 sTwo separate defects usually combine. On the **start** side, the ECS service begins honouring the load balancer health check as soon as a task registers; if the application needs 60 seconds to warm up, the target fails, ECS considers the task unhealthy and replaces it — a restart loop that also sends traffic to a not-yet-ready process. `healthCheckGracePeriodSeconds` on the service tells ECS to ignore ELB health status for that long after a task starts. On the **stop** side, ECS deregisters the old task and sends SIGTERM roughly together; the target group's deregistration delay keeps draining in-flight requests, but if the container exits immediately on SIGTERM — or is SIGKILLed at `stopTimeout` before draining completes — those requests become 5xx. Fix it by making the process stop accepting new connections, finish in-flight work, then exit, and by setting `stopTimeout` to cover the drain window. Then bound the blast radius with `minimumHealthyPercent` and `maximumPercent`, and enable the deployment circuit breaker with rollback.
go deeper
Know that ECS replaces tasks one batch at a time during a deployment, and that a task must pass its load balancer health check to stay in rotation.
Explain healthCheckGracePeriodSeconds, minimumHealthyPercent and maximumPercent, and describe the SIGTERM-then-SIGKILL stop sequence bounded by stopTimeout.
Reason about the three timers together — drain, stop timeout, deregistration delay — diagnose a deploy-time error burst from metrics, and enable the circuit breaker so bad deployments revert themselves.
Set the organisation's deployment contract: rolling versus blue/green, what capacity headroom is budgeted for deploys, and which signals must gate or revert a rollout automatically.
## Why deployments hurt A rolling ECS deployment is two concurrent races: bringing new tasks into rotation before they are ready, and taking old tasks out of rotation before they are finished. Every 5xx during a deploy is one of those two. ## The start side: the grace period When a service has a `loadBalancers` block, ECS watches the target group's health status for its tasks and will stop and replace a task the load balancer reports unhealthy. That is what you want in steady state and exactly wrong during startup of a slow-booting process — a JVM warming a connection pool, a service loading a cache. `healthCheckGracePeriodSeconds` on the service (up to 2,147,483,647, but sane values are tens to hundreds of seconds) tells ECS to ignore ELB health for that period after the task starts. Set it comfortably above your real p99 startup time. Without it, the classic symptom appears: tasks cycling forever, each one killed a minute in, the deployment never stabilising. The grace period does not stop the load balancer *routing* to the target once it passes its first checks, so the second half of readiness is the health check endpoint itself. It should return healthy only when the process can genuinely serve — dependencies connected, caches primed — and should not be a bare TCP accept or a route that returns 200 before initialization finishes. A container-level `healthCheck` in the task definition is a separate mechanism: it runs inside the task and, when it fails, marks the container unhealthy and the task is replaced. Using both is fine; using neither means ECS only knows a task died when the process exits. ## The stop side: draining versus SIGTERM When ECS decides to stop a task, it deregisters the target and the target group enters draining for `deregistration_delay.timeout_seconds` (default 300 seconds on an ALB), during which the load balancer sends no new requests but lets existing ones finish. In parallel, ECS sends **SIGTERM** to the containers and, after the container definition's `stopTimeout`, **SIGKILL**. So three timers must agree: - The application must catch SIGTERM, stop accepting new work, and drain in-flight requests. A process that exits on SIGTERM immediately drops whatever is in flight — instant 5xx. - `stopTimeout` must be at least as long as the drain your application actually needs, or SIGKILL truncates it. It defaults to 30 seconds and has a platform-imposed ceiling on Fargate. - The target group's deregistration delay should be tuned down from 300 seconds towards your real request duration; leaving it at the default just slows deployments without helping. An easy extra win is a short sleep between receiving SIGTERM and beginning shutdown, so that in-flight registration state in the load balancer settles before the process starts refusing work. Note also that connections held open by keep-alive can outlive deregistration — the drain covers requests, not the client's intent to send another. ## Bounding the rollout `deploymentConfiguration` carries two percentages of `desiredCount`: - `minimumHealthyPercent` — how far below desired ECS may go while replacing tasks. At 100 it never dips, so ECS must start new tasks before stopping old ones, which requires headroom. - `maximumPercent` — how far above desired ECS may go. At 200 it can double the service briefly, giving the fastest, safest rollout at the cost of transient capacity. A service that dips below capacity mid-deploy under load will produce 5xx for reasons that have nothing to do with health checks — the remaining tasks simply cannot absorb the traffic. `100/200` is the safe default for request-serving services; `50/100` is what you use when capacity is expensive and brief degradation is acceptable. ## The circuit breaker `deploymentCircuitBreaker` with `enable: true` and `rollback: true` makes ECS watch for a deployment whose tasks repeatedly fail to reach a steady state and automatically roll back to the last known-good task definition. It converts a bad deploy from a page into a graph blip. Combine it with a CloudWatch alarm on target group 5xx so a deployment that is technically healthy but functionally broken also gets caught. ## The clean recipe 1. Real readiness endpoint; `healthCheckGracePeriodSeconds` above p99 startup. 2. SIGTERM handler that drains; `stopTimeout` above the drain time. 3. Deregistration delay tuned to real request duration. 4. `minimumHealthyPercent` 100, `maximumPercent` 200 for anything user-facing. 5. Circuit breaker with rollback enabled, plus an alarm on error rate. When all five are right, a rolling deployment is invisible in the error graphs — and when a deployment still hurts after that, the cause is usually in the application's own startup or connection handling, not in ECS.
- Why can setting minimumHealthyPercent to 50 cause errors even when every task is healthy?Because ECS may stop half the tasks before starting replacements, and the remaining half must absorb all the traffic. If the service was sized near its limit, that halving produces latency and 5xx from saturation, not from health checks. Use 100 with maximumPercent 200 for user-facing services, and reserve 50 for cost-sensitive workloads that tolerate a brief dip.
- What does the ECS deployment circuit breaker actually detect?A deployment whose new tasks repeatedly fail to reach a steady state — tasks that keep stopping or never pass health checks. After enough consecutive failures it marks the deployment failed and, with rollback enabled, redeploys the last successful task definition. It does not detect a deployment that is healthy by ECS's definition but wrong by yours; that needs a CloudWatch alarm on error rate.
- How would you get a deployment with no in-rotation overlap between old and new versions at all?Use a blue/green deployment controller instead of the rolling one: ECS supports CODE_DEPLOY as the deployment controller, which stands up a replacement task set behind a second target group and shifts the listener over once it is healthy, with a bake period and automatic rollback. That buys clean cutover and instant rollback at the cost of running both versions' full capacity.
- Your tasks handle 30-second requests and the deregistration delay is 5 seconds. What breaks?Requests are cut off. Draining ends after 5 seconds and the load balancer stops waiting, so any request still running is terminated as the target leaves. The delay should exceed your realistic maximum request duration, and the container's stopTimeout should exceed that in turn so SIGKILL never arrives mid-drain.
saying these in an interview costs you the question
- Blames the load balancer rather than the shutdown path
- Leaves the container exiting immediately on SIGTERM
- Sets a health check that passes before the app can serve
- Ignores the capacity dip from minimumHealthyPercent
- Thinks deregistration delay alone guarantees a clean drain