skip to content

How do you make a Kubernetes Service send repeated requests from the same client to the same backing pod, and what are the limitations of that mechanism?

level: middleimportance: nice to knowfreq 34%

answer

  1. sessionAffinity: None | ClientIP - that is all there is
  2. timeoutSeconds default 10800 (3h), max 86400
  3. L4 IP hash: no cookies, no headers
  4. SNAT or a fronting proxy collapses affinity onto one pod
  5. per-node, non-durable: externalise real session state

basics

~20 s

Set spec.sessionAffinity: ClientIP, optionally tuning sessionAffinityConfig.clientIP.timeoutSeconds (default 10800). It hashes the client source IP to a pod. It is L4 only, breaks behind NAT or a proxy, and gives no stickiness when the pod dies or the pod set changes.

solid answer

~60 s

A Service's only affinity knob is `spec.sessionAffinity`, with values `None` (default) and `ClientIP`. With `ClientIP`, the node dataplane maps a client's source IP to a chosen pod and keeps that mapping for `spec.sessionAffinityConfig.clientIP.timeoutSeconds` (default 10800, i.e. 3 hours). The limitations matter more than the feature: - **L4 only.** It hashes an IP; there is no cookie, header or JWT-based affinity at the Service layer. Cookie affinity belongs to an L7 layer. - **Source IP must be meaningful.** If external traffic is source-NATed (`externalTrafficPolicy: Cluster`), all requests appear to come from a few node IPs and affinity collapses onto a few pods. If clients sit behind corporate NAT or a shared proxy, they all share one IP. - **Not durable.** It is per-node state. The mapping is lost when the pod goes away, when the pod set changes, or when a different node handles the connection. Treat it as a best-effort optimisation (cache locality), never as a correctness mechanism for session state - keep sessions in a shared store instead.

code

yaml · 16 lines
yaml
apiVersion: v1
kind: Service
metadata:
  name: web
spec:
  type: LoadBalancer
  externalTrafficPolicy: Local
  sessionAffinity: ClientIP
  sessionAffinityConfig:
    clientIP:
      timeoutSeconds: 3600
  selector:
    app: web
  ports:
    - port: 80
      targetPort: 8080

go deeper

for a junior

Know the field name and value: sessionAffinity: ClientIP, which sticks a client IP to one pod.

for a middle

Add the timeout, the fact that it is L4 IP hashing only, and that it breaks under NAT or a fronting proxy.

for a senior

Lead with why affinity is the wrong fix for session correctness, then explain the SNAT interaction, the per-node non-durable state, and where L7 cookie affinity belongs.

for a principal

Frame it as a statefulness decision: whether services may hold user state at all, what the platform offers instead (shared session store, tokens), and the cost of hotspots and uneven drain during scaling.

## The mechanism ```yaml spec: sessionAffinity: ClientIP sessionAffinityConfig: clientIP: timeoutSeconds: 10800 ``` With `sessionAffinity: None` (the default) each new connection to the Service's virtual IP is load-balanced independently. With `ClientIP`, the per-node dataplane records that source IP X was sent to pod P, and sends subsequent connections from X to P until the mapping idles out after `timeoutSeconds` (default 10800 seconds = 3 hours; the maximum is 86400). That is the entire feature: there is no other affinity mode on a Service. ## Why it usually disappoints **It hashes the address it sees, not the user.** Several layers can make that address meaningless: - *Source NAT on external traffic.* With the default `externalTrafficPolicy: Cluster`, packets that a node forwards to a pod on another node are source-NATed to the node's IP. Every client behind that node hashes to the same value, so affinity degenerates: one pod receives everything routed through that node. Setting `externalTrafficPolicy: Local` preserves the client IP and makes affinity meaningful again - at the cost of the load-distribution tradeoffs that policy brings. - *Fronting proxies.* If an ingress or gateway terminates connections and then calls the Service, the Service sees the *proxy's* pod IP. Every user shares one source, so all traffic pins to one backend. In that topology, affinity has to be configured on the proxy (typically cookie-based), not on the Service. - *Client-side NAT.* Mobile carriers, corporate egress and CDNs put thousands of users behind one address; affinity then hands them all to a single pod, creating a hotspot. **It is per-node, in-memory state.** Two nodes handling traffic for the same Service maintain independent mappings, so which pod you get depends on where you entered. Restart the dataplane, reschedule the pod, or scale the Deployment, and the mapping is discarded - the client silently lands somewhere else. That is fatal if your application stored the session only in the pod's memory. **It does nothing for multiplexed protocols.** With HTTP/2 or gRPC, one connection already carries all requests to one pod; affinity adds nothing. And for a single long-lived WebSocket connection there is nothing to re-balance in the first place. ## What it is legitimately good for Affinity is a **performance** hint, not a correctness guarantee. Real uses: - Warm per-pod caches: repeat requests from the same client hit a pod that already has their data loaded. - Expensive per-client setup (a connection pool, a compiled model) that you would rather not repeat per pod. - Reducing cross-pod chatter for workloads that shard state by client. In every case the system must remain correct when the mapping breaks, because it will. ## The right answers for real session state If you actually need a user's session to survive, choose one of these instead: 1. **Externalise the session** into Redis, a database, or a signed cookie/token. Then any pod can serve any request and affinity becomes irrelevant. This is the default correct answer in an interview. 2. **L7 cookie affinity at the ingress or mesh layer**, where the proxy sets its own cookie identifying the chosen backend. It survives NAT because it identifies the *session* rather than the network path, and typically has better draining behaviour on scale-down. It is still a performance tool, not a durability guarantee. 3. **Stable identity via a StatefulSet plus a headless Service**, where the client addresses a specific pod by name. That is sharding, not affinity, and it is the right shape when a particular pod genuinely owns particular state. ## Interview framing The strong answer names `sessionAffinity: ClientIP` and its timeout, immediately qualifies it as L4 IP hashing with per-node non-durable state, points out that SNAT and fronting proxies commonly render it useless, and finishes by saying that sticky sessions are a cache optimisation while session correctness belongs in shared storage.

  • A Service has sessionAffinity: ClientIP but almost all traffic ends up on one pod. What is the likely cause?
    The Service is seeing a single source address for many clients. That happens when external traffic is source-NATed under externalTrafficPolicy: Cluster, when an ingress or gateway proxies the requests so the Service only ever sees the proxy pod's IP, or when many users share one NAT egress address. Affinity then hashes one value and pins everything to one backend. Either preserve the real client IP at the outermost hop, or move affinity to the L7 proxy.
  • Is ClientIP affinity enough to run an application that keeps user sessions in pod memory?
    No. The mapping is per-node, in-memory and best-effort: it is lost when the pod is rescheduled, when the pod set changes during a rollout or scale event, when the affinity timeout expires, or when the client's traffic enters through a different node. The user is then dropped onto a pod that has never seen their session. Store sessions in a shared store or a signed token and treat affinity purely as a cache optimisation.

saying these in an interview costs you the question

  • Claiming Kubernetes Services support cookie-based sticky sessions.
  • Relying on ClientIP affinity to keep in-memory sessions correct.
  • Ignoring that source NAT under the default externalTrafficPolicy destroys affinity.
  • Believing the mapping survives pod rescheduling or a rollout.
  • Enabling affinity behind an ingress controller and expecting per-user stickiness.

context