Why is a NOT_FOUND status from a gRPC health Check not the same answer as a NOT_SERVING ServingStatus?
answer
- two different questions, two different noes
- one is a call failure, one is a response
- a fact about the name, not about health
- the enum only rides in a successful call
- an unknown name is not a down server
basics
~20 sThey answer different questions. NOT_SERVING is a value inside a successful HealthCheckResponse and means a registered service cannot take work. NOT_FOUND fails the call and means the server has nothing registered under the name that was sent.
solid answer
~40 sA `NOT_SERVING` outcome is a **successful** call: the server understood the name, looked up its state, and returned a `HealthCheckResponse` whose `ServingStatus` says it cannot take work right now. A `NOT_FOUND` outcome is a **failed** call with no response message at all, and the defined answer for a `service` name the server never registered — a typo, a stale configuration, or a service that has not registered yet. One is a fact about the target's health; the other is a fact about your request. They also demand different reactions: `NOT_SERVING` is an outage signal, while `NOT_FOUND` is a configuration fault on a server that is demonstrably alive, since only a live server serving health could have produced it.
go deeper
Know that one of these arrives inside a successful answer and the other means the call failed. That single distinction carries most of the point at this level.
Explain that the serving status describes a registered service's condition while the failure status describes the name in the request, and name the third case: a server that does not serve health at all.
Show the operational consequence. Describe a checker that maps every non-OK outcome to down, and what a mistyped service name then does to an on-call rotation, plus how a timed-out check differs from a negative answer.
The judgement is how much fidelity a health signal should carry across an estate. Rich outcomes let teams diagnose faster but only if every consumer handles them; a signal that half the consumers flatten is worse than a deliberately simple one.
## Three different ways to say no A health check can fail to say `SERVING` for reasons that have almost nothing in common. Collapsing them into a single "not healthy" bit is one of the most common mistakes made against this service. - **`NOT_SERVING`** — a value of `ServingStatus`, carried inside a `HealthCheckResponse`, delivered by a call that **succeeded**. The server knows the name in your request, and is telling you that service cannot take work at the moment. - **`NOT_FOUND`** — a gRPC status that **fails** the call. No `HealthCheckResponse` exists. The server does serve the health service, and has no service registered under the name you sent. - **`UNIMPLEMENTED`** — a gRPC status that fails the call for a different reason: this server does not serve the health service at all, so there is no health information to be had from it by this route. ## Where each one lives | Outcome | Delivered as | What it says about the target | Reasonable reaction | |---|---|---|---| | `SERVING` | Value in a successful response | That named service can take work | Use it | | `NOT_SERVING` | Value in a successful response | That named service cannot take work now | Treat as unhealthy | | `UNKNOWN` | Value in a successful response | Registered, no considered answer yet | Re-check; do not conclude | | `NOT_FOUND` | Failed call, no response | Nothing registered under that name | Fix the name, not the server | | `UNIMPLEMENTED` | Failed call, no response | No health service here | Use another signal | | `DEADLINE_EXCEEDED` | Failed call, no response | No answer within your budget | Corroborate before concluding | The column that matters is the third one. Two of these rows describe the **server's** condition. Two describe **your request**. One describes the **absence of information**. A caller that reacts identically to all six is not really checking health; it is checking whether a call returned `OK`. ## Why collapsing them causes false incidents The failure mode is concrete. Someone configures a checker with a service name of `catalogue` instead of the fully qualified `seedvault.catalogue.v1.Catalogue`. Every check fails with `NOT_FOUND`. A checker that maps every non-`OK` outcome to "down" now reports an outage on a server that is answering perfectly, and the longer the misconfiguration survives, the more confidently the dashboards lie. The inverse mistake costs just as much. A checker that treats `NOT_FOUND` as "unknown, ignore" will sail past a service that silently stopped registering itself, reporting nothing wrong precisely because it is asking about a name that no longer exists. ## The case where nothing comes back at all `DEADLINE_EXCEEDED` deserves separate handling, because it is the absence of information rather than a negative answer. A `NOT_SERVING` response proves the server is up, is reachable, and is honest enough to admit it cannot work — that is a lot of information. A check that simply never returns proves only that no answer arrived within the time you were willing to wait. The server may be wedged, saturated, or perfectly fine behind a slow path. This is also the outcome most affected by the caller's own choices. A check with no deadline never produces this signal at all; it just hangs, and a hung checker looks like a quiet one. ## What a caller should do with each 1. **Successful response** — act on the `ServingStatus` value. This is the only branch where the server has actually told you about its health. 2. **`NOT_FOUND`** — raise it as a configuration fault. Compare the name you sent against the fully qualified name the server registers, or fall back to the empty name for overall status. Do not mark the target down. 3. **`UNIMPLEMENTED`** — record that this target offers no health information, and decide deliberately whether that is acceptable. It is not an outage. 4. **`DEADLINE_EXCEEDED` or an unreachable target** — treat as a suspicion, not a verdict, and corroborate with a second signal before acting on it. The habit underneath all four is the one interviewers are probing: read the outcome for what it is a fact *about*. A health check that is only ever consulted as pass/fail throws away most of what the service was designed to tell you.
- What should a supervising caller do when a health Check returns NOT_FOUND?Treat it as a configuration fault, not an outage. The server answered, so it is alive and serving health; it simply has nothing registered under that name. Compare the name you sent against the fully qualified `package.Service` the server registers, or fall back to the empty name for overall status. Marking the target down here hides a typo behind a false incident.
- How does a health Check that times out differ from one that returns NOT_SERVING?`NOT_SERVING` is information: the server is up and telling you it cannot take work. `DEADLINE_EXCEEDED` is the absence of information — nothing arrived in the time you allowed, which is consistent with a wedged server, an overloaded one, or a healthy one behind a slow path. The first is a state you can act on; the second usually needs a second signal before you conclude anything.
saying these in an interview costs you the question
- Treats every non-OK health outcome as the server being down.
- Thinks NOT_FOUND means the server itself is unhealthy.
- Believes NOT_SERVING arrives as a call failure rather than in a successful response.
- Raises an outage for a mistyped service name.
- Cannot say what a caller should conclude when the health call itself times out.