After a batch of instances registered in Consul is terminated abruptly, their entries linger in the catalog in a critical state and never go away. Why does Consul keep them, and which registration and check settings make it remove them automatically?
answer
- a check result never deletes a registration
- name the setting with the word critical in it
- hooks on shutdown, not just on start
- leaving is not the same as failing
- catalog edits lose to the owning agent
basics
~20 sA failing check never deletes a registration — only an explicit deregistration does. Set deregister_critical_service_after on the check so the agent removes a service that stays critical, deregister in a shutdown hook, and stop agents with a graceful leave rather than a kill.
solid answer
~50 sConsul treats registration and health as independent, so a check going critical takes the instance out of DNS and out of `?passing` queries but leaves the registration in place — deliberately, so a temporary failure heals itself. Nothing reaps it unless you asked for that. The three levers are: set **`deregister_critical_service_after`** on the check definition, which makes the owning agent deregister the service once it has been critical for that duration (minimum one minute); deregister explicitly on shutdown with `PUT /v1/agent/service/deregister/<service_id>` from a shutdown hook or supervisor; and stop agents with **`consul leave`** rather than SIGKILL, since a graceful leave deregisters the node's services. If the whole node died, its `serfHealth` check goes critical from gossip and its services stop resolving immediately, but the failed node itself stays in the catalog until Consul reaps it much later. Deleting entries from the catalog while the owning agent is alive does not stick — anti-entropy re-asserts them.
code
json · 14 lines{
"service": {
"name": "web",
"id": "web-1",
"port": 8080,
"tags": ["primary"],
"check": {
"http": "http://localhost:8080/health",
"interval": "10s",
"timeout": "2s",
"deregister_critical_service_after": "5m"
}
}
}go deeper
Know that a critical instance stops receiving traffic but stays registered, and that something has to deregister it explicitly for it to disappear.
Explain the pulled check types versus the pushed TTL check, and name deregister_critical_service_after as the setting that lets the agent clean up on its own.
Diagnose which layer died — service gone with a live agent, versus the whole node gone — and pick the right lever for each, including the graceful-leave path in your shutdown tooling.
Set the fleet-wide convention: who is responsible for deregistration, what the timeout budget is, and how you keep a widespread incident from deregistering a whole pool that would otherwise have recovered.
## Why Consul keeps a broken registration The design decision is that a check result is a *statement about right now*, and a registration is a *statement that this thing exists*. A service that fails its check during a slow start, a GC pause or a dependency blip should stop receiving traffic and then resume receiving it, with no re-registration and no operator involvement. That only works if the registration survives the failure. The cost of that decision is exactly the symptom in the question: kill a workload without telling Consul, and its registration outlives it, sitting critical forever. ## Which layer actually died Diagnose this before reaching for a setting, because two different failures look similar in the UI: **The service died, the agent lives.** The agent keeps running its check against a process that no longer answers, so the check stays critical indefinitely. The node itself is healthy and still gossiping. This is the case `deregister_critical_service_after` is built for. **The whole node died.** The agent stopped gossiping, so `serfHealth` goes critical and every service on that node leaves discovery within seconds. But the node and its registrations remain in the catalog; Consul reaps a node that has stayed in the failed state only after a long grace period, measured in days, on the assumption that a node may come back. ## The fixes, in the order you should apply them **1. `deregister_critical_service_after` on the check.** This is the automatic cleanup, expressed in the check definition: ```json { "service": { "name": "web", "id": "web-1", "port": 8080, "check": { "http": "http://localhost:8080/health", "interval": "10s", "timeout": "2s", "deregister_critical_service_after": "5m" } } } ``` The owning agent removes the service once the check has been critical for that long. There is a minimum (one minute), and the value is a real trade: too short and a service that is slow to recover gets deregistered and must re-register itself; too long and dead entries accumulate. Match it to how long a legitimate recovery can take — usually a small multiple of your slowest startup. **2. Deregister on the way out.** For anything with a controlled shutdown, an explicit deregistration is faster and unambiguous: ``` PUT /v1/agent/service/deregister/web-1 ``` Run it from the application's shutdown hook or the supervisor's stop command. This also removes the window in which the instance is registered but no longer accepting connections. **3. Leave gracefully.** `consul leave` — or the agent handling a termination signal as a graceful leave — deregisters the node's services and removes the node from the membership as a *left* node rather than a *failed* one. The distinction matters: a left node is expected to be gone, a failed node is expected to return. ## Where check type changes the picture The check types differ in who initiates: - **HTTP, TCP, gRPC and script checks are pulled** by the agent on its interval. If the target is gone, the pull fails and the check is critical — forever, absent a deregister rule. - **TTL checks are pushed** by the application, which must periodically call `PUT /v1/agent/check/pass/<check_id>`. If the application dies, no update arrives and the check expires into critical on its own. TTL checks are useful for workloads the agent cannot probe from outside, but they push liveness responsibility into the application, which then has one more thing to get wrong — a hung process may keep heartbeating while serving nothing. ## The anti-pattern to name explicitly "Cleaning up" by deleting the entries through the catalog is not a fix while the owning agent is running: the agent is authoritative for its own node, and the next anti-entropy sync re-creates what you deleted. Removal has to happen at the agent that owns the registration — or, when the node is genuinely gone, by forcing the failed node out of membership at the server tier. ## The shape of a correct answer in an interview Say the principle first (a check never deregisters by itself), then name the three levers, then note the trade in choosing the timeout, then mention that catalog-side deletion does not stick. That sequence shows you have operated the thing rather than read the flag list.
- How would you choose the value for `deregister_critical_service_after`?Base it on the longest legitimate recovery you expect — a slow restart, a dependency outage the instance rides out — and add margin, because deregistering a live instance forces it to re-register and can empty a pool during a wide incident. Minutes, not seconds. The minimum Consul allows is one minute, and a value near that is only safe for workloads that re-register automatically at startup.
- When would you use a TTL check instead of an HTTP or TCP check?When the agent cannot meaningfully probe the workload from outside — a batch worker with no listening port, or a process whose real liveness is "I completed a loop iteration". The application then pushes `PUT /v1/agent/check/pass/<check_id>` on an interval. The cost is that liveness reporting lives in application code, so a partially hung process can keep heartbeating while doing no useful work.
- A node's agent was killed and the node will never come back. How do you get it out of the catalog now rather than waiting?Force it out of membership at the server tier rather than editing the catalog, since the node is gone and cannot re-assert itself but Consul still expects a failed node to return. Removing the node from membership drops its registrations with it. Doing this while the node is merely partitioned is what you must avoid — it would delete registrations that the agent would otherwise restore.
saying these in an interview costs you the question
- Expecting a failed check to remove the registration by itself
- Deleting entries from the catalog while the agent still runs
- Setting the deregister timeout to the minimum for everything
- Killing agents instead of letting them leave gracefully
- Assuming TTL heartbeats prove the process is doing useful work