An IoT telemetry ingest gateway behind a Kubernetes LoadBalancer Service with externalTrafficPolicy: Local has some pods saturated while others idle. Why, and how would you fix it?
answer
- two-tier split
- each healthy node, equal share
- count replicas per node
- spread on kubernetes.io/hostname
- long-lived connections keep the skew
basics
~20 sThe load balancer spreads load per node, and each node then splits its share only among its own local pods. With pods unevenly spread, pods on crowded nodes starve while lone pods overload. Spread the replicas one per node, or use a load balancer that weights nodes.
solid answer
~50 sUnder `Local`, the external load balancer usually gives every node that passes the `healthCheckNodePort` probe an equal share, whatever the number of pods on it. Each node's kube-proxy then splits its share only among its own local endpoints. On our 27-worker cluster, 11 gateway replicas landed on 7 nodes as 3, 3, 1, 1, 1, 1, 1. At 91% CPU allocation the scheduler places pods wherever they fit, which is how that happens. Each node gets about 14.3%, so a pod sharing a node with two others gets about 4.8% while a lone pod gets 14.3%. The ideal is 9.1% each. Long-lived device connections make this stick until devices reconnect. Fixes: spread replicas with a `topologySpreadConstraints` entry on `kubernetes.io/hostname` (or pod anti-affinity, or a DaemonSet on labelled ingest nodes); use a load balancer that honours the node weight kube-proxy reports; or switch back to `Cluster` and carry the client IP another way.
code
bash · 2 lineskubectl get pods -n ingest -l app=telemetry-gateway \
-o custom-columns=NODE:.spec.nodeName --no-headers | sort | uniq -cgo deeper
Recall that Local sends each node's traffic only to that node's own pods, so how many pods each node holds matters.
Work through the per-node then per-pod split with concrete numbers, and show how an uneven placement produces the hot pods.
Diagnose from replica-per-node counts and probe answers, then choose a spread rule that fits the cluster's allocation, including the Pending risk and how to rebalance long-lived connections.
Decide whether client-IP visibility is worth owning pod placement, compared with keeping Cluster and carrying the address another way, and document that cost in the platform's service guidance.
## The symptom An **IoT telemetry ingest gateway** runs as 11 replicas behind a `type: LoadBalancer` Service with `externalTrafficPolicy: Local`. The team chose `Local` so the gateway sees each device's real source address. On **a 3-control-plane, 27-worker self-managed cluster** running at **91% node CPU allocation**, dashboards show a few gateway pods pinned at their CPU limit while others sit near idle, even though every pod is identical. ## Why Local balances per node, not per pod `externalTrafficPolicy: Local` changes two things on each node. The node forwards external traffic **only to endpoints on itself**, and it does **not** masquerade, which is why the client address survives. The external load balancer learns which nodes have endpoints from the `healthCheckNodePort` probe. Unless it weights nodes, it gives **each healthy node an equal share**. That makes the load two-tier: equal per node, then split per pod within each node. The arithmetic for the placement found here: | Node group | Nodes | Pods per node | Share per node | Share per pod | |---|---|---|---|---| | Crowded | 2 | 3 | 1/7 = 14.3% | about 4.8% | | Lone pod | 5 | 1 | 1/7 = 14.3% | 14.3% | | **Ideal** | — | — | — | 1/11 = **9.1%** | Each lone pod carries three times the load of a pod on a crowded node, and about 57% more than it should. ## Why the placement is uneven - At **91% CPU allocation**, the scheduler places pods wherever requests still fit. Nothing asks it to spread gateway pods across nodes. - Without a spread rule, rollouts and evictions pile replicas onto whichever nodes had room at that moment. - Under `Cluster` this would not matter much, because any node can forward to any pod. `Local` exposes the placement directly. ## Why it persists - **Long-lived connections.** Devices hold MQTT-over-TLS connections for hours. Balancing happens when a connection opens, so moving a pod does not move existing connections until devices reconnect. - **`sessionAffinity: ClientIP`**, if set, pins each source address to one pod for `sessionAffinityConfig.clientIP.timeoutSeconds` (default **10800**, at most **86400**). Thousands of devices behind a few carrier NAT addresses become a few very heavy clients. The affinity state is also **per node**, so it cannot even out load across nodes. ## Diagnosing it 1. Count replicas per node: `kubectl get pods -n ingest -l app=telemetry-gateway -o custom-columns=NODE:.spec.nodeName --no-headers | sort | uniq -c`. 2. Compare per-pod CPU (`kubectl top pods -n ingest`) with that count. Hot pods should line up with lone-pod nodes. 3. Probe each node's `healthCheckNodePort` to confirm which nodes the load balancer considers healthy, and read `localEndpoints`. 4. Check whether the load balancer uses the `X-Load-Balancing-Endpoint-Weight` header kube-proxy returns. Most implementations treat the probe as pass/fail. ## Fixing it 1. **Spread replicas one per node.** Add a `topologySpreadConstraints` entry with `topologyKey: kubernetes.io/hostname` and `maxSkew: 1`. With `whenUnsatisfiable: DoNotSchedule` and nodes 91% allocated, some replicas may stay **Pending**, so reserve headroom or use `ScheduleAnyway` and accept some residual skew. Required pod anti-affinity gives a strict one-per-node rule with the same capacity caveat. 2. **Give the gateway its own nodes.** A DaemonSet with a `nodeSelector` on labelled ingest nodes guarantees exactly one pod per node, which is the shape Local rewards. 3. **Use node weighting**, if your load-balancer implementation honours the endpoint-count weight. Then per-node share follows pod count. 4. **Rebalance connections.** After fixing placement, a rolling restart makes devices reconnect onto the new layout. Stagger it to avoid a reconnect storm. 5. **Go back to `Cluster`** if the source address can reach the gateway another way, for example a PROXY-protocol header from the load balancer. That is a separate mechanism with its own trade-offs, owned elsewhere. Pod placement then stops mattering for balance. ## The judgement call `Local` makes you own pod placement to get the client address. If the gateway's scaling is volatile and the cluster runs hot, spreading constraints and headroom are part of the price. Name that price in the design, rather than finding it on a CPU dashboard. ## Common mistakes - **Scaling out to fix skew.** New replicas may land on already-crowded nodes, and the imbalance stays. - **Blaming the application.** Identical pods with very different CPU usually point to placement, not a code path. - **Restarting only the hot pods.** Their devices reconnect, but the placement that caused the skew is unchanged, so the pattern returns. - **Forgetting that affinity is per node.** `sessionAffinity: ClientIP` cannot move a client to a different node, because the load balancer, not kube-proxy, picks the node.
- The gateway Service also sets sessionAffinity: ClientIP. How does that interact with the skew?It makes it worse. Each source address is pinned to one pod for `timeoutSeconds` (default 10800, maximum 86400), so devices behind a few carrier NAT addresses become a few very heavy clients stuck to single pods. kube-proxy keeps the affinity state per node, so under Local it only applies within the node the load balancer chose, and it cannot redistribute across nodes. Prefer device-level balancing in the application, or drop affinity.
- Why might a DoNotSchedule spread constraint make things worse on this cluster?At 91% CPU allocation, strict spreading may find no node that satisfies both the skew and the 750m request, so new replicas stay Pending during a scale-up or rollout. Total capacity then falls exactly when it is needed. Reserve headroom, use ScheduleAnyway and accept bounded skew, or dedicate labelled nodes to the gateway.
- When would you abandon Local altogether for this gateway?When placement cannot be controlled cheaply, for example with bursty autoscaling on a hot shared cluster, and the client address can travel another way, such as a PROXY-protocol header added by the load balancer. With Cluster, any node forwards to any pod, so balance no longer depends on placement. The costs are the extra hop and the SNAT, which the application-level address mechanism has to make up for.
saying these in an interview costs you the question
- Local makes the load balancer weight nodes by pod count automatically
- kube-proxy on a busy node forwards overflow to pods on other nodes
- Restarting the hot pods rebalances load permanently
- Adding replicas always fixes skew under Local
- sessionAffinity: ClientIP spreads NATed devices across pods