With externalTrafficPolicy: Local on a Kubernetes LoadBalancer Service, how does the external load balancer learn which nodes can actually serve traffic?
answer
- the load balancer probes a node port
- a second, dedicated node port
- kube-proxy answers per node
- 200 versus 503, localEndpoints
- probe interval is the black-hole window
basics
~20 sThrough spec.healthCheckNodePort. kube-proxy on every node answers HTTP there: 200 when the node has a ready local pod for the Service and kube-proxy is healthy, 503 otherwise. The load balancer health-checks that port and stops sending to 503 nodes.
solid answer
~50 sUnder `externalTrafficPolicy: Local` a node delivers external traffic only to its **own** ready endpoints and does not masquerade it, so the client IP survives. A node with no local endpoint **drops** the packet. To avoid that, Kubernetes allocates `spec.healthCheckNodePort` for `LoadBalancer` Services with `Local`. It comes from the NodePort range, or you can pick a free port, and it cannot be changed once set. kube-proxy on every node serves an HTTP health endpoint on it. The endpoint returns **200** when that node has at least one ready local endpoint and kube-proxy is healthy, and **503** otherwise. The JSON body carries `localEndpoints`. The load-balancer implementation points its health check at that port, so only nodes hosting pods stay in rotation. There is still a gap: while probes catch up after a pod moves, some connections can be dropped.
code
yaml · 15 linesapiVersion: v1
kind: Service
metadata:
name: telemetry-gateway
namespace: ingest
spec:
type: LoadBalancer
externalTrafficPolicy: Local
selector:
app: telemetry-gateway
ports:
- name: mqtts
port: 8883
targetPort: 8883
protocol: TCPgo deeper
Remember that Local sends traffic only to pods on the receiving node, so the load balancer must skip nodes without pods, and it does that by probing healthCheckNodePort.
Explain what kube-proxy returns on that port, when you get 200 and when 503, and why the port exists only for LoadBalancer Services with Local.
Talk through the black-hole windows during scale-down and node moves, how probe settings bound them, and how you would confirm a node's answer with curl and the no-local-endpoints metric.
Weigh the dependency Local creates on the load balancer's probe behaviour and support for weighting against simply keeping Cluster and carrying the client address another way.
## Why Local needs a health signal at all A Kubernetes **Service** exposed as `type: LoadBalancer` gets an external load balancer that sends traffic to cluster **nodes**. Each node's **kube-proxy** then forwards it to a pod. The field `spec.externalTrafficPolicy` decides what the node does next: - **`Cluster`** (the default): the node may forward to any ready endpoint in the cluster, and it **masquerades** (SNATs) the packet. The pod therefore sees a node IP, not the client's. - **`Local`**: the node forwards **only to endpoints running on itself** and does **not** masquerade, so the pod sees the real client address. The price of `Local` is spelled out in the API: traffic sent to a node with **no local endpoints is dropped**. An external load balancer that spreads connections across all 27 workers of **a 3-control-plane, 27-worker self-managed cluster** would send most of them into nothing, because **an IoT telemetry ingest gateway** runs on only a few nodes. The load balancer needs to know which nodes hold a pod. ## healthCheckNodePort For `type: LoadBalancer` with `externalTrafficPolicy: Local`, the API server sets `spec.healthCheckNodePort`: - It is allocated automatically from the NodePort range (by default `30000-32767`), or you may request a specific free port. - It is **only valid** for that combination. Setting it on a Service that does not need it fails validation. - It **cannot be updated** once set, and it is cleared if the Service stops needing it (for example, when the type changes). The load-balancer implementation reads the port and aims its health checks at it. The shared cloud-provider helper uses the path `/healthz`. ## What kube-proxy answers kube-proxy on **every** node listens on that port for each such Service: | Node state | HTTP status | Body `localEndpoints` | |---|---|---| | At least one **ready** local endpoint and kube-proxy healthy | `200` | the count, e.g. `2` | | No ready local endpoint | `503` | `0` | | kube-proxy itself reports unhealthy | `503` | the count, whatever it is | The response also carries a `serviceProxyHealthy` field and an `X-Load-Balancing-Endpoint-Weight` header set to the local count. Whether a given load balancer uses that header to weight nodes depends on the implementation. Many treat the check as plain pass/fail. ## The timing gaps The mechanism is eventually consistent, so there are windows where traffic is lost: 1. **A node's last pod goes away.** kube-proxy starts answering 503 and drops new external packets at once. The load balancer only notices after its probe interval times its unhealthy threshold. New connections sent in that window are dropped. 2. **A pod starts on a new node.** The node answers 200 once the pod is ready, but it receives no traffic until the load balancer's healthy threshold is met. That is capacity you pay for but do not use yet. 3. **kube-proxy restarts or is unhealthy.** Its 503 takes the node out of rotation even if the pods are fine. How kube-proxy handles terminating endpoints during a rollout, and the preStop timing that covers the gap, are a separate topic (connection draining). ## An in-cluster exception The API notes that traffic sent to the **load-balancer IP from inside the cluster** always gets `Cluster` semantics. kube-proxy short-circuits it to any endpoint, so a pod calling the external address is not dropped on a pod-less node. A pod calling a **NodePort** directly, however, is subject to the Local policy of the node it picks. ## Checking it on a node - `kubectl get svc telemetry-gateway -n ingest -o jsonpath='{.spec.healthCheckNodePort}'` gives the port. - `curl -i http://<node-ip>:<port>/healthz` on a node with a pod returns 200 and a non-zero `localEndpoints`. On a node without one it returns 503. - The kube-proxy metric `kubeproxy_sync_proxy_rules_no_local_endpoints_total` shows how many Local Services have no local endpoints on that node. The health port is how `Local` stays usable. It moves the question of where the pods are to the load balancer, which only learns the answer as fast as its probes run. ## Common mistakes - **Overriding the check to probe the Service's `nodePort`.** On a pod-less node that traffic is silently dropped, so the probe only times out. It cannot tell an empty node from a slow one, and it ignores kube-proxy's own health. Use the allocated health port, which answers explicitly with 200 or 503. - **Blocking the health port.** Node firewalls must admit the load balancer's probes on `healthCheckNodePort`, or every node looks unhealthy and the Service goes dark. - **Reading 200 as app health.** A 200 means the node has a **ready** local endpoint and kube-proxy is healthy. Application health comes from the pod's readiness probe, which feeds that count.
- Why does a node with healthy gateway pods still get pulled out of the load balancer's rotation?The health endpoint returns 503 not only when there are zero ready local endpoints but also when kube-proxy itself reports unhealthy. A kube-proxy that is restarting or failing to sync its rules takes the node out even though the pods are fine. Check kube-proxy's own health on its healthz address (default port 10256) and its logs on that node.
- What happens to connections that reach a pod-less node before the load balancer marks it unhealthy?They are dropped. Under externalTrafficPolicy: Local, kube-proxy installs a drop rule for external traffic on nodes with no local endpoints. How long that lasts is set by the load balancer's probe interval and unhealthy threshold, so tightening those shrinks the loss. Clients with retries and reconnect logic, which IoT devices usually have, get through on the next attempt to a healthy node.
- Can you change a Service's healthCheckNodePort after creation?No. The field cannot be updated once set. You may request a specific free port at creation. Otherwise the API server allocates one from the NodePort range. It is cleared automatically if the Service stops needing it, for example when the type changes away from LoadBalancer or the policy goes back to Cluster.
It is like a delivery firm that phones each depot before sending a truck. Depots with the item in stock say yes, and the rest say no, so trucks only go where the goods are.
saying these in an interview costs you the question
- The load balancer probes the application's own port to find pod-hosting nodes
- A pod-less node under Local forwards the packet to another node instead of dropping it
- healthCheckNodePort is allocated for every NodePort Service
- A 200 from the health port means the application inside the pod is healthy
- Local still SNATs the client address, but only once