skip to content

Kubernetes replaced the older Endpoints API with EndpointSlices for tracking the backends of a Service. What problem did that solve, and how does the newer object differ in structure?

level: middleimportance: should knowfreq 48%

answer

  1. One giant object -> sharded slices of 100
  2. Write and watch amplification was the bug
  3. Conditions: ready / serving / terminating
  4. zone plus nodeName = topology hints
  5. Label kubernetes.io/service-name

basics

~20 s

The old Endpoints object crammed every backend of a Service into one resource, so one pod change rewrote the whole object and pushed it to every watcher - painful at thousands of endpoints. EndpointSlices shard the same data into capped chunks (100 endpoints by default), so a change updates one small slice, and they add per-endpoint conditions and topology fields.

solid answer

~50 s

The legacy **Endpoints** object holds all backends of a Service in a single resource. Every readiness flip or pod churn rewrites that whole object, and the API server sends the full new version to every watcher - kube-proxy on every node, ingress controllers, mesh control planes. For a Service with thousands of endpoints the object gets large (and can approach etcd value limits), so update traffic scales as endpoints times watchers. **EndpointSlices** shard the same information: the EndpointSlice controller creates multiple slices per Service, each holding up to 100 endpoints by default, labelled `kubernetes.io/service-name`. A single pod change now rewrites one small slice instead of one giant object, which cuts control-plane traffic dramatically. The schema also got richer. Each endpoint carries **conditions** - `ready`, `serving`, and `terminating` - rather than the old ready/notReady bucketing, plus `zone` and `nodeName` for topology-aware routing and a per-slice `addressType` (IPv4, IPv6, FQDN) which makes dual-stack clean. EndpointSlices are GA since v1.21 and are what kube-proxy actually consumes; the Endpoints API is deprecated.

code

bash · 2 lines
bash
kubectl get endpointslices -l kubernetes.io/service-name=web
kubectl get endpointslices -l kubernetes.io/service-name=web -o yaml

go deeper

for a junior

Know that EndpointSlices list the pod IPs backing a Service, that there can be several per Service, and that they replaced the older Endpoints object.

for a middle

Explain the scalability motivation - full-object rewrites broadcast to every watcher - and name the new fields: conditions, zone, nodeName, addressType.

for a senior

Connect it to observed control-plane behaviour during large rollouts, and to consumers you own (custom controllers, ingress) that must merge across slices.

for a principal

Treat it as a case study in API design for scale: sharding to bound update fan-out, richer per-item state to enable draining and topology routing, and a mirroring path to keep one consumer contract.

## What these objects are for A Service is a selector plus a virtual IP. Something must turn that selector into the concrete list of pod IPs currently able to serve traffic, and publish it so dataplane components can program themselves. That list is the Endpoints / EndpointSlice objects, written by a controller in kube-controller-manager and consumed by kube-proxy, ingress controllers, service meshes, and DNS. ## The problem with Endpoints The original `Endpoints` object is one resource per Service containing every address. Three problems emerged at scale: 1. **Write amplification.** Objects in Kubernetes are replaced wholesale, not patched field-by-field on the wire. One pod becoming ready rewrites the entire object, even if it holds 5,000 addresses. 2. **Watch amplification.** Every watcher receives the full object on every change. With kube-proxy on 500 nodes plus ingress controllers, one pod flip fans out into hundreds of large payloads. Rolling a big Deployment produces a storm of these, and this became a documented control-plane bottleneck - API server and etcd CPU spikes, watch lag, slow endpoint convergence. 3. **Object size limits.** etcd caps values (default around 1.5 MB); very large Services could approach or hit that. ## How EndpointSlices fix it The EndpointSlice controller writes multiple objects per Service, each capped (default 100 endpoints, tunable up to 1000 via the controller's `--max-endpoints-per-slice`). Slices are ordinary namespaced objects labelled `kubernetes.io/service-name: <svc>` so consumers list them by label. Now a single pod change rewrites a roughly 100-endpoint slice, so both the payload and the fan-out shrink by about the sharding factor. The controller also tries to pack changes into existing slices rather than constantly rebalancing, trading a little fragmentation for far fewer writes. ## Structural differences worth naming - **Per-endpoint conditions.** Endpoints had two flat lists: `addresses` (ready) and `notReadyAddresses`. EndpointSlice gives each endpoint a conditions object with `ready`, `serving`, and `terminating`. This separation matters: a pod that is terminating but still passing its readiness probe reports `ready: false, serving: true, terminating: true`, which lets a proxy keep draining in-flight work to it without sending new traffic - impossible to express in the old API. - **Topology fields.** `zone` and `nodeName` per endpoint enable topology-aware routing hints, so a proxy can prefer same-zone backends and cut cross-AZ traffic cost. - **addressType per slice.** Each slice is `IPv4`, `IPv6`, or `FQDN`, which makes dual-stack Services a clean pair of slice sets instead of a mixed list. - **Multiple slices per Service is normal**, and slices may be uneven in size. Do not assume one slice, and never assume ordering. ## Practical implications - **Debugging** shifts from `kubectl get endpoints <svc>` to `kubectl get endpointslices -l kubernetes.io/service-name=<svc> -o yaml`. The former still works via a compatibility path in current versions, but it is the deprecated view. - **Anything you write that watches Service backends** - a custom ingress controller, an operator that reconciles external load balancers - should consume EndpointSlices and merge across slices, not read Endpoints. - **Endpoints without a selector** (manually managed, pointing at an external database, say) still works; the EndpointSlice mirroring controller copies such Endpoints into slices so consumers only need one code path. - **The 100-per-slice default is a tradeoff**: larger slices mean fewer objects but bigger updates. Raising it helps clusters with many mid-sized Services and hurts ones with a few huge Services. ## Timeline EndpointSlices went beta in v1.17 and GA in v1.21, at which point kube-proxy switched to consuming them by default. The Endpoints API was formally deprecated in v1.33 - it still exists for compatibility but new work should treat EndpointSlice as the real API.

  • Why does the default cap of 100 endpoints per slice matter, and when would you change it?
    The cap sets the blast radius of a single endpoint change: with 100 per slice, one pod flip rewrites and re-broadcasts about 100 endpoints' worth of data. Raising it (up to 1000) reduces object count and controller bookkeeping, which helps clusters with very many small Services, but it increases the payload of every update, which hurts Services with huge, churning backend sets. It is a cluster-wide controller flag, so tune it to the dominant workload shape.
  • Does a Service with no selector still produce EndpointSlices?
    Yes. If you manage Endpoints manually - the usual pattern for pointing a Service at an external database - the EndpointSlice mirroring controller copies those Endpoints into corresponding slices. That keeps consumers on a single API surface, and kube-proxy programs the DNAT rules the same way as for selector-based Services.

Endpoints was a single shared document that had to be re-sent in full to every reader whenever one line changed; EndpointSlices split it into numbered pages so a one-line edit only re-sends one page.

saying these in an interview costs you the question

  • Saying EndpointSlices changed Service semantics or how traffic is routed - it is a scalability change to how backends are published
  • Assuming exactly one EndpointSlice per Service
  • Treating ready and serving as synonyms, so terminating-but-draining endpoints cannot be described
  • Claiming Endpoints was removed - it is deprecated but still present for compatibility
  • Thinking the sharding is per node rather than a fixed-size chunking of the endpoint list

context