A Kubernetes Service exposed outside the cluster has externalTrafficPolicy set to Cluster by default, and it can be set to Local instead. What changes when you switch it, and what are the tradeoffs?
answer
- Cluster = may forward + SNAT -> node IP as source
- Local = local pods only, real client IP, no extra hop
- Local: pod-less nodes fail the LB health check
- Local risk: per-node balancing -> uneven pod load
- internalTrafficPolicy = same idea for in-cluster traffic
basics
~20 sCluster forwards external traffic to a pod on any node, adding a hop and rewriting the source IP to the node's. Local only serves traffic from pods on the receiving node, preserving the real client IP but dropping traffic on pods-less nodes and risking uneven load.
solid answer
~60 s`externalTrafficPolicy` controls what a node does with traffic that arrives from outside the cluster on a node port or load balancer. **`Cluster`** (default): any node accepts the traffic and may forward it to a backing pod on **another** node. That second hop requires source NAT, so the pod sees the *node's* IP, not the client's. Upside: every node is a valid entry point, so load spreads evenly and no node needs a local pod. **`Local`**: a node only serves the request from pods running on **itself**, with no SNAT, so the pod sees the **real client IP**. Nodes with no backing pod fail their load balancer health check and are pulled out of the pool. Downsides: traffic is distributed across *nodes* rather than pods, so an uneven pod spread produces uneven load; and during a rollout a node can briefly have zero ready pods. Use `Local` when you genuinely need client IPs (rate limiting, geo, audit, allow-lists) or want to cut a hop, and pair it with topology spread so every load-balancer node has a pod. Otherwise keep `Cluster`. `internalTrafficPolicy` is the analogous setting for in-cluster traffic.
code
yaml · 37 linesapiVersion: v1
kind: Service
metadata:
name: edge
spec:
type: LoadBalancer
externalTrafficPolicy: Local
selector:
app: edge
ports:
- port: 443
targetPort: 8443
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: edge
spec:
replicas: 6
selector:
matchLabels:
app: edge
template:
metadata:
labels:
app: edge
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: edge
containers:
- name: edge
image: example/edge:2.1go deeper
Recognise that the default forwards anywhere and hides the client IP, and that Local keeps the client IP.
Explain the SNAT-versus-extra-hop mechanics and that pod-less nodes fail the health check under Local.
Own the tradeoff: uneven per-node distribution, rollout windows, topology spread as the mitigation, and where in the chain to preserve client identity for L7 headers.
Decide it at platform level: which tier terminates external connections, whether the edge is DaemonSet-shaped, cross-zone traffic cost, and what client-identity guarantees downstream services are allowed to assume.
## The problem it solves External traffic enters through a node - a node port, or a cloud load balancer whose backend pool is the nodes. The node that receives the packet is not necessarily running a pod that can serve it. What it does next is what `externalTrafficPolicy` decides. ## Cluster (the default) Every node accepts traffic for the Service and load-balances over **all** ready backing pods cluster-wide. If the chosen pod is on another node, the packet is forwarded there. That forwarding requires **source NAT**: the packet's source address is rewritten to the forwarding node's IP so the reply comes back the same way. The consequence is that the application sees the node IP as the client, and every request appears to originate from a handful of node addresses. The benefits are real: load spreads evenly across pods regardless of how they are placed, every node is a healthy entry point, and adding or draining nodes is uneventful. The costs: lost client identity, and an extra network hop (extra latency, and cross-zone traffic that may be billed). ## Local With `Local`, a node serves external traffic **only** from pods running on that node, and never forwards. Because there is no second hop, no SNAT is needed and the pod sees the original client address. To make this workable, the node also answers the load balancer's health check based on whether it has a ready local pod: nodes without one report unhealthy and the cloud load balancer stops sending them traffic. For raw NodePort access there is no such protection - hitting a node port on a pod-less node with `Local` simply drops the packet, which is a classic "it works from some nodes, not others" bug. **The load-distribution catch.** The cloud load balancer spreads across *nodes*, and then each node spreads across only its *local* pods. With 3 nodes where node A has 1 pod and node B has 5, the single pod on A receives roughly the same share as all five on B combined. The fix is placement: `topologySpreadConstraints` (or pod anti-affinity) that spreads replicas evenly, plus enough replicas relative to nodes. A DaemonSet-shaped workload (ingress controllers, edge proxies) sidesteps the problem entirely, which is why ingress controllers are the canonical `Local` user. **The rollout catch.** During an update, a node can transiently have zero ready pods. Its health check goes unhealthy and the load balancer withdraws it - usually fine, but the withdrawal is not instant, so there is a window where traffic arrives at a node with nothing to serve it. Adequate replica counts, a PodDisruptionBudget, and surge-based rolling updates shrink the window. ## Why client IP matters Rate limiting, abuse blocking, IP allow-lists, geolocation, and audit logging all become useless when every request appears to come from ten node IPs. `Local` is the L4 answer. The L7 answer is a proxy that sets `X-Forwarded-For` and terminates connections at the edge; note that if that proxy itself runs in the cluster behind a `Cluster`-policy Service, the header will carry a node IP, which is exactly the trap. Get the policy right at the outermost hop and let the header carry it inward - and only trust the header from proxies you control. ## internalTrafficPolicy A separate field, `spec.internalTrafficPolicy`, applies the same idea to traffic originating **inside** the cluster: `Cluster` (default) balances over all pods; `Local` restricts to pods on the calling pod's own node, and if none exist, the connection fails. Its use case is node-local agents (logging or metrics sidecar daemons) where crossing the network would be wasteful. Do not use it for general services: a node with no local pod gets hard failures, not fallback. ## Choosing Default to `Cluster` and change deliberately. Choose `Local` when client IP is a functional requirement, or when you run a DaemonSet-style edge tier where every node has a pod anyway and the saved hop matters. When you do choose it, verify: is the replica count comfortably above the node count in the load-balancer pool, are replicas spread across zones and nodes, and does the health-check behaviour match what the cloud provider actually implements? Some providers expose additional annotations for health-check ports and intervals that materially change failover speed.
- With externalTrafficPolicy: Local, some requests to a NodePort succeed and others are dropped depending on which node you hit. Why?Under Local a node serves external traffic only from pods running on itself, and drops it when there is none. A cloud load balancer hides this because pod-less nodes fail the health check and are withdrawn from the pool, but raw NodePort access has no such filter, so hitting a pod-less node directly drops the packet. Either spread pods across all nodes, front the Service with a load balancer that honours the health checks, or use Cluster policy.
- You need real client IPs for rate limiting, but your ingress controller runs behind a Service with the default Cluster policy. What does the application see, and how do you fix it?It sees the X-Forwarded-For header populated with the node IP that source-NATed the connection, because the ingress controller itself never received the true client address. The fix is to set externalTrafficPolicy: Local on the ingress controller's Service - the outermost in-cluster hop - so the controller sees the real client and writes a correct header that inner services can trust.
Cluster policy is a call centre where any office answers and transfers you internally - convenient, but the transfer hides your caller ID. Local policy means each office only handles walk-ins it can serve, so your identity survives, but an understaffed office turns people away.
saying these in an interview costs you the question
- Claiming Local preserves client IP for in-cluster traffic too (that is internalTrafficPolicy).
- Enabling Local without checking that replicas are spread across the load-balancer's nodes.
- Assuming Local guarantees even load because the cloud load balancer is 'balanced' - it balances nodes, not pods.
- Believing X-Forwarded-For is automatically correct regardless of the Service policy.
- Thinking Cluster policy drops traffic on nodes without a local pod.