Kubernetes probes support exec, httpGet, tcpSocket and grpc handlers. How does each one work, and what are the failure modes of choosing the wrong one?
answer
- kubelet -> Pod IP directly, never via a Service
- httpGet: 200-399 = success, HTTPS cert unverified
- tcpSocket: proves a listener, not a working app
- exec: forks per period, needs the binary in the image
- grpc: standard grpc.health.v1 Health/Check
basics
~20 shttpGet: the kubelet calls the container's IP and port; status 200-399 is success. tcpSocket: success if a TCP connection opens - it proves only that a listener exists. exec: runs a command inside the container, exit 0 is success, but forks a process every period. grpc: calls the standard gRPC health checking service.
solid answer
~60 sAll four are executed by the **kubelet on the node**, straight to the Pod IP - never through a Service. - **httpGet** - an HTTP request to the container's port and path; any status **200-399** is success, everything else fails. Cheap and expressive; the default choice. With `scheme: HTTPS` the certificate is not verified, and custom headers are supported for Host or a static token. - **tcpSocket** - success if a TCP connection can be opened. It proves a listener exists and nothing more: a deadlocked app whose accept queue still completes handshakes passes happily. Use only for non-HTTP servers with no better signal. - **exec** - runs a command inside the container; exit 0 is success. Most powerful, most expensive: it forks a process every period in every replica, and the binary must exist in the image, which breaks on distroless or scratch images. - **grpc** - the kubelet acts as a client of the standard `grpc.health.v1.Health/Check` service; SERVING is healthy. Generally available since 1.27, and it removes the need to ship a health-probe binary in the image. Rule of thumb: HTTP if you can, gRPC for gRPC servers, exec last.
code
yaml · 20 linesreadinessProbe:
httpGet:
path: /readyz
port: 8080
httpHeaders:
- name: Host
value: api.internal
---
livenessProbe:
tcpSocket:
port: 5432
---
livenessProbe:
exec:
command: ["/usr/local/bin/healthcheck", "--local"]
---
readinessProbe:
grpc:
port: 9090
service: api.v1.Ordersgo deeper
Name the four handlers and their success criteria: HTTP 200-399, TCP connect, exit code 0, gRPC SERVING.
Explain that the kubelet probes the Pod IP directly, and give the practical limits - tcpSocket proves too little, exec needs the binary and forks each period.
Choose deliberately per workload, account for probe cost at replica scale, and know the operational traps: distroless images, unverified HTTPS certs and probe-blocking NetworkPolicies.
Define a fleet standard for health endpoints and handler choice, including the gRPC health contract, and treat aggregate probe load as a capacity concern.
## Who runs the probe Every probe is executed by the **kubelet on the Pod's own node**, directly against the Pod IP. Three consequences people often miss: probes do not traverse a Service or an ingress, so they never test service discovery; probe traffic does not appear in load-balancer metrics; and a NetworkPolicy that blocks node-to-Pod traffic can break probes even though normal traffic works. ## httpGet The kubelet issues an HTTP GET to `http[s]://<podIP>:<port><path>`. **Any response code in the 200-399 range counts as success**; 4xx, 5xx, connection refused and timeouts count as failure. Fields: `path`, `port` (number or named port), `scheme` (HTTP/HTTPS), `host` (defaults to the Pod IP - overriding it is nearly always a mistake) and `httpHeaders` for a Host header or static token. With `scheme: HTTPS` the kubelet does **not** verify the server certificate, so self-signed certs are fine. Relying on a redirecting health endpoint is fragile - return the status directly. Best practice: a dedicated lightweight endpoint that does not authenticate, does not log at info level (probe traffic is high volume) and does not fan out to dependencies unless that is intentional for readiness. ## tcpSocket The kubelet opens a TCP connection to the port and closes it. Success means the kernel accepted a connection. The weakness is fundamental: the OS accept queue completes handshakes while the application thread pool is entirely blocked, so a wedged process still reports healthy. As a liveness probe that is close to useless; as a readiness probe it at least catches "the port is not open yet". Use it for protocols with no cheap health verb - a plain TCP server, some databases and brokers - and prefer an application-level check when one exists. Repeated connect-and-drop also litters server logs with connection-reset noise. ## exec The kubelet runs a command inside the container's namespaces; **exit code 0 is success**, anything else is failure, and a hang is cut off at `timeoutSeconds`. Advantages: it can inspect files, run a CLI health command, and check things not exposed over the network. Costs: - **Process churn.** Every period, in every container, in every replica, a process is forked. A one-second period across thousands of Pods is measurable node CPU, and shells are heavier than they look. - **Image dependency.** The binary must exist in the image. Distroless and scratch images have no shell or HTTP client, so a shell-based probe fails with an obscure exec error. - **Zombies and leaks.** Probe commands that spawn children and exit can leave defunct processes when nothing reaps them. - **Timeout semantics.** Modern kubelets enforce `timeoutSeconds` properly, but a slow command still consumes resources for its whole run. Use it when nothing else can express the check - verifying a local file, or querying replication state through a vendor CLI. ## grpc Instead of shipping a health-probe binary and using exec, the kubelet can be a native gRPC client calling `grpc.health.v1.Health/Check` with an optional `service` name. `SERVING` is success; any other status or a connection failure is failure. The only requirement is that the server implements the standard health service. Generally available since Kubernetes 1.27, and on by default as beta from 1.24. ## Choosing 1. HTTP server -> `httpGet` with a purpose-built endpoint. 2. gRPC server -> `grpc` with the standard health service. 3. Non-HTTP protocol with a real health command -> `exec`, kept cheap with a sane period. 4. Nothing better available -> `tcpSocket`, understanding that it proves very little. Remember probe cost scales with `replicas / periodSeconds`: an expensive check every second across a large deployment is a self-inflicted load problem, and it can hammer whatever the endpoint touches.
- Your image is distroless and the exec probe fails immediately. Why?Distroless images contain no shell and almost no binaries, so a command like /bin/sh -c '...' cannot be executed inside the container. Either switch to an httpGet or grpc probe, or ship a small static health binary in the image for the exec handler to invoke.
- Why is a tcpSocket liveness probe considered weak?The kernel completes TCP handshakes independently of whether the application is processing requests, so a deadlocked or thread-starved process still accepts connections and passes. It detects only a crashed or unbound listener, which the container runtime largely detects anyway.
saying these in an interview costs you the question
- Believing probes are routed through the Service rather than straight to the Pod IP
- Thinking only HTTP 200 is a success for httpGet, or that 3xx fails
- Assuming an HTTPS probe validates the server certificate
- Treating a tcpSocket probe as proof the application is functioning
- Ignoring the CPU cost of a frequent exec probe across many replicas