Two instances of a service are registered in Consul and one instance's HTTP check has gone critical. Explain what a DNS lookup and the `/v1/health` and `/v1/catalog` HTTP endpoints each return now, and how that state differs from the instance having been deregistered.
answer
- registration and health are different facts
- one endpoint filters, one just lists
- warning is not the same as critical
- soft removal versus hard removal
- node gossip failure takes all its services out
basics
~20 sDNS and /v1/health/service/<name>?passing return only the healthy instance. /v1/catalog/service/<name> returns both, because the catalog lists registrations regardless of health. A critical instance is still registered and returns automatically on recovery; a deregistered one is gone until something registers it again.
solid answer
~50 sHealth filtering in Consul is applied by the *query*, not by the registration. A DNS lookup of `web.service.consul` builds its answer from instances whose checks are acceptable, so the critical instance is absent. `GET /v1/health/service/web?passing` does the same filtering and also returns the full check detail for what it does return. `GET /v1/catalog/service/web` is the trap: the catalog is the registration list, so it returns **both** instances with no health filtering at all — anyone writing a client-side balancer against the catalog endpoint routes straight into the broken instance. The difference from deregistration is who cleans up and what comes back: a critical instance is still a registration, still visible in the UI and the catalog, and it re-enters DNS by itself the moment its check passes again. A deregistered instance has been removed from the catalog entirely; nothing will bring it back except a new registration.
code
bash · 10 lines# Usable instances only: Consul applies the health filter
curl -s 'http://127.0.0.1:8500/v1/health/service/web?passing' \
| jq -r '.[] | "\(.Service.Address):\(.Service.Port)"'
# Every registration, healthy or not - do NOT build a pool from this
curl -s 'http://127.0.0.1:8500/v1/catalog/service/web' \
| jq -r '.[] | "\(.ServiceAddress):\(.ServicePort)"'
# Hard removal: deregister at the agent that owns the registration
curl -s -X PUT 'http://127.0.0.1:8500/v1/agent/service/deregister/web-1'go deeper
Be able to say that a failing check takes an instance out of DNS without removing its registration, and that it comes back on its own once the check passes.
Explain the difference between the health endpoint with ?passing and the unfiltered catalog endpoint, and why warning instances are still returned unless only_passing is set.
Demonstrate the operational read: why critical entries linger after a hard kill, why deleting them from the catalog does not stick while the agent lives, and how serfHealth pulls a whole node's services out at once.
Own the convention across the fleet — which query surface services are allowed to build pools from, whether warning counts as out of rotation, and who is responsible for deregistering on shutdown.
## Registration and health are two separate facts A Consul registration says "this node offers this service on this address and port". A check says "and right now it is passing, warning or critical". Nothing about a failing check removes the registration. Understanding that separation explains almost every surprising result in this area. ## The three check states Checks report **passing**, **warning** or **critical**. Warning is a real, distinct state, not a synonym for failing — a script check exiting `1` yields warning, exiting `2` or higher yields critical. It matters here because *by default, warning instances are still returned by DNS*. Only critical instances are excluded. If you want warning excluded too, you set `only_passing` in the agent's `dns_config`. Teams that assume warning is filtered are often routing traffic to instances that reported degradation. There is also a check nobody registers explicitly: **`serfHealth`**, the node-level check maintained from gossip. When a node stops gossiping, `serfHealth` goes critical and *every* service on that node drops out of discovery, regardless of how those services' own checks last reported. ## What each query surface does **DNS** (`web.service.consul`) returns addresses for instances that pass the health filter. No detail, no reason, no way to see the excluded ones. **`GET /v1/health/service/web`** returns every instance *with* its full check list, letting you decide yourself. Adding `?passing` makes Consul apply the filter server-side, which is what most client libraries do. **`GET /v1/catalog/service/web`** returns the registrations, unfiltered. This endpoint is about *what is registered*, not *what is usable*. Using it to build an upstream pool is a genuine production bug: the pool includes instances Consul already knows are down. ``` # usable instances only curl -s 'http://127.0.0.1:8500/v1/health/service/web?passing' # every registration, healthy or not curl -s 'http://127.0.0.1:8500/v1/catalog/service/web' ``` ## Critical versus deregistered | | Check critical | Deregistered | |---|---|---| | In the catalog | yes | no | | Returned by DNS | no | no | | Visible in the UI | yes, in red | no | | Returns on recovery | automatically, at the next passing check | only if something registers it again | | Who caused it | the check failing | an explicit deregistration, or an automatic one | The critical state is a *soft* removal from traffic. It is reversible, self-healing, and it preserves the information that this instance exists and is unwell — which is exactly what you want during a deploy, a slow restart or a dependency blip. Deregistration is a *hard* removal: the registration is gone and the instance vanishes from the catalog. ## How each one gets triggered Deregistration happens when: - The application or its supervisor calls `PUT /v1/agent/service/deregister/<service_id>` on the local agent — the correct thing to do in a shutdown hook. - The agent is stopped gracefully with `consul leave`, which deregisters the node's services on the way out. - A check has been critical long enough to trip `deregister_critical_service_after` on the check definition, if it was set. A critical state that *persists* usually means none of those happened: the process was killed, or the container was replaced while the agent kept running, and the check simply keeps failing against nothing. ## The consequence people meet in production A node's agent is killed abruptly. Gossip notices, `serfHealth` goes critical, and the node's services stop resolving — good, traffic moves away in seconds. But the node and its services stay in the catalog, red in the UI, for a long time. Nothing is broken; the catalog is telling you a registration exists whose owner is unreachable. Operators who "clean up" by deleting catalog entries while the agent is still alive discover the entries reappear, because the agent's anti-entropy sync re-asserts its own node's state. The removal has to happen at the agent that owns the registration.
- A check is in the warning state. Does Consul's DNS still return that instance?Yes, by default. Consul excludes critical instances from DNS answers but returns warning ones, on the reasoning that warning means degraded rather than down. If you want warning instances excluded as well, enable `only_passing` in the agent's `dns_config`. Teams that never set it and assume warning is filtered end up sending traffic to instances that explicitly reported trouble.
- An operator deletes a service from the catalog while its agent is still running. What happens?It comes back. The agent is authoritative for the services on its own node, so the next anti-entropy sync re-asserts the registration and the catalog entry reappears. Permanent removal has to be done at the owning agent — deregister the service through the local agent API, or remove its definition from the agent's config and reload.
- What is the `serfHealth` check, and why does it matter for discovery?It is the node-level check Consul maintains from gossip membership rather than from anything you configure. If a node stops gossiping, `serfHealth` goes critical and every service registered on that node drops out of DNS and out of health-filtered queries at once, no matter what those services' own checks last reported. Node liveness is effectively a check on all of them.
saying these in an interview costs you the question
- Building an upstream pool from /v1/catalog/service
- Assuming a failing check deregisters the instance
- Treating warning as excluded from DNS by default
- Thinking a critical instance must be re-registered to recover
- Deleting catalog entries while the owning agent is still running