skip to content

Calls through a Kubernetes Service succeed most of the time, but a small percentage fail with connection resets, and the rate spikes during deployments. What causes this intermittent pattern and how would you fix it?

level: seniorimportance: should knowfreq 52%

answer

  1. Partial failures = per-request variable: which backend, when, which connection
  2. SIGTERM and endpoint propagation race — not synchronised
  3. preStop sleep + graceful drain + longer grace period
  4. Failure rate ≈ 1/N replicas → one bad backend
  5. keep-alive / HTTP/2 pin clients to one pod

basics

~20 s

Endpoint updates are eventually consistent: a terminating pod stops accepting connections before every node's kube-proxy has removed it, so a fraction of requests are sent to a dead backend. Fix with a preStop sleep, graceful shutdown, and readiness flipping before termination.

solid answer

~60 s

Two families of cause, and the deployment correlation points strongly at the first. **Endpoint propagation race.** When a pod is deleted, two things happen in parallel: the kubelet sends SIGTERM, and the EndpointSlice controller marks the address not-ready, after which every node's kube-proxy must reprogram its rules. Those are not synchronised. If the application shuts its listener immediately on SIGTERM, there is a window — usually a second or two, longer on large clusters — where nodes still forward to a socket that is gone, producing resets. Fix: a `preStop` hook that sleeps a few seconds (so the pod keeps serving while endpoints propagate), an application that finishes in-flight requests before closing, and `terminationGracePeriodSeconds` comfortably longer than preStop plus drain. **Per-backend faults.** One replica in a bad state that still passes readiness: with N replicas you see roughly 1/N failures. Isolate by calling pod IPs directly. Also check connection reuse — keep-alive and HTTP/2 pin a client to one backend, so one bad pod hurts a fixed subset of clients disproportionately.

code

yaml · 12 lines
yaml
spec:
  terminationGracePeriodSeconds: 45
  containers:
    - name: api
      lifecycle:
        preStop:
          exec:
            command: ["sh", "-c", "sleep 10"]   # keep serving while endpoints propagate
      readinessProbe:
        httpGet: {path: /ready, port: 8080}
        periodSeconds: 2
        failureThreshold: 2

go deeper

for a junior

Know that pod termination and endpoint updates are not instantaneous, and that a preStop delay plus graceful shutdown reduces errors during rollouts.

for a middle

Describe the termination sequence and the propagation path through EndpointSlices and kube-proxy, and set preStop and grace periods coherently.

for a senior

Use the failure statistics to discriminate causes, bypass the Service to isolate a bad backend, and account for connection reuse under HTTP/2 and gRPC.

for a principal

Treat graceful shutdown as a platform default — templated preStop and grace periods, readiness semantics, retry and idempotency policy — and weigh a mesh or per-request proxy against the cost of connection-level balancing.

## Why "a small percentage" is the important clue Total failure has few causes and they are easy to find. Partial failure means the request path is *usually* correct, so the fault is in something that varies per request: which backend was chosen, when the request arrived relative to a state change, or which connection it reused. That narrows the search dramatically. ## Cause 1: the endpoint removal race (the deployment correlation) Deleting a pod triggers several independent sequences: 1. The API server sets `deletionTimestamp`. 2. The kubelet sees it and runs any `preStop` hook, then sends SIGTERM to the container. 3. The EndpointSlice controller sees it and marks the pod's address `ready: false` / `terminating: true`. 4. kube-proxy on **every node** watches EndpointSlices and rewrites iptables or IPVS rules. 5. Any external load balancer or ingress controller updates its own backend list. Steps 2 and 3–5 race. On a small idle cluster the propagation is milliseconds; with thousands of nodes, a busy API server, or a controller under load, it can be seconds. If the application closes its listening socket the instant it receives SIGTERM, every request that kube-proxy sends in that window hits a closed port and the client sees a reset — which is exactly the pattern of a low, deployment-correlated error rate. The standard remedy is counter-intuitive: make the pod keep serving *after* it has been told to die. ```yaml lifecycle: preStop: exec: {command: ["sh", "-c", "sleep 10"]} terminationGracePeriodSeconds: 45 ``` The preStop hook delays SIGTERM, so the pod continues to accept traffic while the not-ready state propagates. The application then handles SIGTERM by refusing *new* connections, finishing in-flight requests, and exiting — and the grace period must exceed preStop plus the longest legitimate request, otherwise the kubelet SIGKILLs mid-request (exit 137 with reason Error) and you have traded one truncation for another. A complementary tactic is flipping readiness to failing *before* deleting the pod (some frameworks expose a shutdown hook that does this), so removal from endpoints begins earlier in the sequence. ## Cause 2: one unhealthy backend kube-proxy spreads new connections across ready endpoints roughly evenly. So one bad replica out of eight gives you about 12.5% failures — a stable, suspiciously round number. "Bad" here means still passing readiness but not actually serving: an exhausted connection pool, a wedged thread pool, a stale config, a half-failed migration, a node with a broken route. Isolate by taking the Service out of the loop: ``` kubectl get endpointslices -l kubernetes.io/service-name=api -o jsonpath='{.items[*].endpoints[*].addresses[*]}' # then curl each pod IP directly in a loop and compare ``` If exactly one address fails consistently, you have your pod, and the next question is why readiness did not catch it — usually because the probe is shallower than the failure. ## Cause 3: connection reuse defeats load balancing kube-proxy balances **connections**, not requests. A client with HTTP keep-alive, a connection pool, or HTTP/2 and gRPC (which multiplex everything over a single long-lived connection) picks a backend once and stays there. Consequences: - A newly scaled-up replica receives no traffic from existing clients. - One bad backend affects a fixed subset of clients completely rather than everyone slightly. - During a rollout, clients holding connections to terminating pods experience the reset at the moment of shutdown — GOAWAY handling and connection-max-age settings matter. Mitigations: bound connection lifetime on the client, use a headless Service with client-side load balancing for gRPC, or put a proxy/mesh in the path that balances per request. ## Cause 4: infrastructure-level intermittency Worth keeping in the differential: conntrack table exhaustion on busy nodes (`nf_conntrack: table full, dropping packet` in kernel logs) causes seemingly random drops; source-NAT port exhaustion for high-fan-out egress does the same; and for `externalTrafficPolicy: Cluster` on NodePort/LoadBalancer Services, an extra hop and SNAT can hide the true client IP and add a failure surface. These are less common than the shutdown race but produce a similar statistical signature. ## How to investigate systematically 1. **Quantify.** What percentage, and is it a round fraction of the replica count? Does it track deployments, scaling events, or nothing in particular? 2. **Correlate with pod lifecycle.** Overlay error timestamps with pod deletion and creation times. Errors clustered in the seconds after a delete are the race; errors evenly spread are a bad backend or infrastructure. 3. **Bypass the Service.** Curl each pod IP directly, in a loop, and compare. This separates "one backend is bad" from "the routing is racy". 4. **Check the shutdown path.** Does the app handle SIGTERM? Is there a preStop hook? Is the grace period long enough? Read the timings in `kubectl describe pod` for a terminating replica. 5. **Look at the client.** Connection pooling and HTTP/2 change the failure distribution; a client that retries idempotent requests on reset masks the problem entirely — which is also a legitimate mitigation. ## The design lesson Graceful shutdown in Kubernetes is a *distributed* operation: no single component knows the moment a pod stops being routable everywhere. The only reliable strategy is overlap — keep serving after you are told to stop, long enough for everyone else to find out.

  • Why does adding a preStop sleep reduce errors during rollouts, when it seems to just delay shutdown?
    Because it creates deliberate overlap. Endpoint removal and SIGTERM are dispatched independently, and kube-proxy on every node needs time to reprogram its rules. During the preStop sleep the pod is already marked not-ready but is still serving, so requests that were routed before the update arrive at a live listener instead of a closed socket. Once the sleep ends and SIGTERM is delivered, essentially no traffic is still being directed at the pod.
  • Your failure rate sits at almost exactly 25% with four replicas. What does that suggest and how do you confirm it?
    That one of the four backends is bad while still counted as ready, since kube-proxy spreads new connections roughly evenly. Confirm by reading the Service's EndpointSlice addresses and calling each pod IP directly in a loop — a single address failing consistently identifies the pod. The follow-up question is why readiness did not exclude it, which usually means the probe checks something shallower than the actual failure.
  • How does HTTP/2 or gRPC change the picture for Service load balancing?
    kube-proxy balances at connection setup, and HTTP/2 and gRPC multiplex many requests over one long-lived connection, so a client effectively picks a backend once and keeps it. New replicas receive no traffic from existing clients, and a single bad backend fails all requests from the clients pinned to it rather than a small share of everyone's. Typical remedies are bounding connection age on the client, using a headless Service with client-side load balancing, or introducing a proxy or mesh that balances per request.

Closing a shop the second head office marks it 'closed' still leaves customers arriving with last week's directions in hand. The fix is to keep the lights on for a few minutes after the sign changes, until every map has been reprinted.

saying these in an interview costs you the question

  • Assuming the SIGTERM and endpoint-removal sequences are synchronised
  • Setting terminationGracePeriodSeconds shorter than the preStop hook plus in-flight request time
  • Concluding the Service is misconfigured when only a fraction of requests fail
  • Ignoring client connection reuse and expecting per-request balancing from a ClusterIP Service
  • Treating client-side retries as a complete fix rather than a mitigation for a shutdown race

context