skip to content

In Kubernetes, how does a Service's internalTrafficPolicy: Local differ from trafficDistribution: PreferSameNode when keeping calls on the caller's node?

level: middleimportance: nice to knowfreq 30%

answer

  1. hard rule versus hint
  2. what happens with no local endpoint
  3. EndpointSlice hints, forNodes
  4. PreferClose is an old name
  5. DaemonSet agent rollout window

basics

~20 s

internalTrafficPolicy: Local is a hard rule: ClusterIP traffic goes only to endpoints on the caller's node, and is dropped if there are none. trafficDistribution: PreferSameNode is a preference: it uses a same-node endpoint when one is ready and otherwise falls back to any ready endpoint.

solid answer

~50 s

`internalTrafficPolicy` governs only in-cluster traffic to a Service's ClusterIP. `internalTrafficPolicy: Local` tells kube-proxy to use **only** endpoints on the calling pod's node. If that node has none, the traffic is **dropped** and callers time out. That suits a node-local agent whose data must not leave the node. `trafficDistribution: PreferSameNode` is a **hint**. The EndpointSlice controller marks endpoints with node hints, kube-proxy prefers a ready endpoint on the same node, and when none is ready it falls back to other ready endpoints: same-zone ones if zone hints exist, otherwise any in the cluster. `PreferSameZone` works the same way per zone, and `PreferClose` is its deprecated older name. The older `service.kubernetes.io/topology-mode: Auto` annotation instead has the controller assign zone hints in proportion to capacity. Choose Local when correctness needs same-node traffic, and a preference when you want locality without the risk of dropped traffic.

code

yaml · 13 lines
yaml
apiVersion: v1
kind: Service
metadata:
  name: telemetry-buffer
  namespace: ingest
spec:
  selector:
    app: telemetry-buffer
  ports:
    - name: grpc
      port: 7443
      targetPort: 7443
  trafficDistribution: PreferSameNode

go deeper

for a junior

Remember that internalTrafficPolicy: Local only ever uses pods on the caller's node, while trafficDistribution only expresses a preference.

for a middle

Explain what each does when the caller's node has no ready endpoint, where the hints come from, and why PreferClose is just an old name for PreferSameZone.

for a senior

Show how a DaemonSet rollout turns Local into dropped writes, and how you would detect that and design callers to cope, or choose a preference instead.

for a principal

Set platform guidance on when a hard node-local rule is justified compared with preference-based locality, weighing correctness, cross-zone cost and the risk of concentrating load.

## The situation On **a 3-control-plane, 27-worker self-managed cluster**, **an IoT telemetry ingest gateway** writes each batch to a buffering agent that runs as a **DaemonSet**, one pod per node, behind a ClusterIP Service. The team wants each gateway pod to talk to the agent **on its own node**, to avoid a network hop and keep the buffer local. Kubernetes offers two Service fields for this, and they fail in opposite ways. ## The two knobs - **`spec.internalTrafficPolicy`** accepts `Cluster` (the default) or `Local`. With `Local`, kube-proxy routes ClusterIP traffic **only to endpoints on the same node as the client pod**. The API is explicit: traffic is **dropped** if there are no local endpoints. - **`spec.trafficDistribution`** is a **preference** that implementations may use as a hint. The values are `PreferSameZone`, `PreferSameNode`, and `PreferClose`, a **deprecated** older name that means exactly `PreferSameZone`. The EndpointSlice controller writes **hints** into each endpoint (`hints.forZones`, and for same-node, `hints.forNodes`), and kube-proxy reads them. | Aspect | `internalTrafficPolicy: Local` | `trafficDistribution: PreferSameNode` | |---|---|---| | Nature | Hard constraint | Preference | | Ready endpoint on caller's node | Used | Used | | No ready endpoint on caller's node | **Dropped** | **Falls back** to all ready endpoints | | Applies to | ClusterIP traffic only | Any traffic routed with Cluster semantics | | Good for | Data that must stay on the node | Locality as an optimisation | ## Where the difference bites While the agent DaemonSet rolls out, a node's agent pod is briefly not ready. With `Local`, every gateway pod on that node loses its writes for that window. kube-proxy installs a drop rule, and the metric `kubeproxy_sync_proxy_rules_no_local_endpoints_total` counts such Services. With `PreferSameNode`, those writes go to an agent on another node until the local one is ready again. kube-proxy's fallback check counts only **ready** endpoints. If no ready endpoint is hinted for this node, or not every ready endpoint carries a node hint, it ignores the node preference and tries the zone level, then plain cluster-wide routing. ## Zones and the older annotation - `PreferSameZone` keeps traffic in the caller's **zone** when a ready endpoint there is hinted for it. The API warns you to use it only when clients and endpoints are spread so that the preference will not overload endpoints in one zone. - The annotation `service.kubernetes.io/topology-mode: Auto` is the older topology-aware routing. There, the EndpointSlice controller assigns zone hints **in proportion to** capacity and may leave hints out when a safe allocation is not possible. It is more conservative than `PreferSameZone`, which simply prefers the caller's own zone. ## Scope and interplay 1. `internalTrafficPolicy` never changes **external** traffic, which is `externalTrafficPolicy`'s job. `trafficDistribution` shapes whichever traffic kube-proxy routes with Cluster semantics, so it can also affect external traffic under `externalTrafficPolicy: Cluster`. 2. When `internalTrafficPolicy: Local` applies, a topology preference adds nothing, because the candidate set is already the local node. 3. Traffic from inside the cluster to a load-balancer IP gets Cluster semantics, not the external Local rule. ## Choosing - Pick **`Local`** when routing elsewhere would be *wrong*: a node-level log or metrics agent reading host data, or a buffer whose contents belong to that node. Then make sure a local pod always exists, for example with a DaemonSet and tolerations matching every node, and make callers tolerate brief outages. - Pick **`PreferSameNode`** or **`PreferSameZone`** when locality only saves latency or cross-zone cost, and a cross-node call is acceptable. - Watch endpoint load either way. Preferences can concentrate traffic when callers are unevenly placed. ## Common mistakes - **Using `Local` for a Deployment with fewer replicas than nodes.** Callers on nodes without a replica get every call dropped. - **Expecting a preference to be a guarantee.** `PreferSameNode` sends traffic off the node whenever no local endpoint is ready, which matters if the data must stay local. - **Expecting `internalTrafficPolicy` to affect the load balancer.** External traffic follows `externalTrafficPolicy`. - **Setting `PreferClose` in new manifests.** It still works, but it is a deprecated alias; write `PreferSameZone` instead.

  • How would you confirm a node-local buffer Service is dropping traffic on a node?
    Check whether that node has a ready agent pod (`kubectl get pods -o wide` filtered to the node) and read kube-proxy's `kubeproxy_sync_proxy_rules_no_local_endpoints_total` metric on that node, which counts Local-policy Services with no local endpoints. Callers typically see timeouts rather than refusals, because the packets are dropped.
  • Why would you use trafficDistribution: PreferSameZone instead of the topology-mode: Auto annotation?
    PreferSameZone is a simple rule: use a same-zone endpoint when one is ready. It is predictable and works for small endpoint counts. The Auto annotation has the EndpointSlice controller assign hints in proportion to zone capacity and withhold them when that is unsafe, so traffic may cross zones in ways that are hard to predict. Pick PreferSameZone when callers and endpoints are evenly spread across zones.

saying these in an interview costs you the question

  • internalTrafficPolicy: Local falls back to other nodes when no local pod is ready
  • trafficDistribution guarantees traffic never leaves the caller's node
  • internalTrafficPolicy: Local also preserves the external client's source IP
  • PreferClose selects the nearest node rather than the same zone