skip to content

In Kubernetes, how do you run a canary with two Deployments behind one Service using replica counts, and where does that approach break down?

level: middleimportance: should knowfreq 55%

answer

  1. share follows the pod count
  2. a track label beside app
  3. Service selects the common label
  4. one pod is the smallest step
  5. balanced per connection, not request

basics

~20 s

Put stable and canary Deployments behind one Service that selects only their shared app label; the canary gets roughly its fraction of ready pods. That share moves in one-pod steps, is counted per connection, and drifts when either side scales.

solid answer

~40 s

You label both pod templates `app: autocomplete` and add `track: stable` or `track: canary`. Each Deployment's selector includes its `track` label, while the Service selects only `app: autocomplete`, so its backends are both sets of pods. With 47 stable pods and 1 canary pod, the canary gets roughly 1/48, about 2%, of new connections. The limits are all about precision. The smallest step is one pod, so 1% needs 100 pods in total and 0.5% needs 200. kube-proxy picks a backend per **connection**, not per request, so a few keep-alive clients can push the real share far from the pod ratio. If an HPA scales the stable Deployment, the ratio changes without anyone deciding to change it. When you need a precise or very small share, you move the split into a weighted route.

code

yaml · 10 lines
yaml
apiVersion: v1
kind: Service
metadata:
  name: autocomplete
spec:
  selector:
    app: autocomplete
  ports:
    - port: 80
      targetPort: 7311

go deeper

for a junior

Recall the layout: two Deployments with a shared app label and different track labels, and a Service that selects only the app label.

for a middle

Explain the share as canary pods over all Ready pods, why kube-proxy balances per connection, and why the Deployments need distinct selectors.

for a senior

Diagnose a canary whose real share is far from its pod ratio, and account for HPA drift, session affinity and the pod count a small share would need.

for a principal

Decide when coarse replica ratios are acceptable and when the platform should offer weighted routing so share and capacity are separate controls.

## The idea A **canary** sends a small slice of real traffic to a new version before everyone gets it. Kubernetes has no canary object, but a Service spreads connections across every Ready pod its selector matches. If most of those pods run the old version and a few run the new one, the few get a small share. The pod ratio becomes the traffic ratio, roughly. ## The layout The pattern, described in the Kubernetes documentation on managing resources, uses a `track` label: - Deployment `autocomplete-stable`: selector and pod labels `app: autocomplete`, `track: stable`, 47 replicas, current image. - Deployment `autocomplete-canary`: selector and pod labels `app: autocomplete`, `track: canary`, 1 replica, new image. - Service `autocomplete`: selector `app: autocomplete` only, so it matches both sets of pods. The `track` label does two jobs. It keeps the two Deployments' selectors apart, and the documentation warns against controllers whose selectors overlap. It also lets your metrics tell canary pods from stable ones, which you need to judge the canary at all. ## Doing the arithmetic kube-proxy in iptables mode chooses a backend for each new connection with a random rule per endpoint, weighted so that every Ready endpoint is equally likely. So the expected canary share of new connections is canary pods divided by all Ready pods. | Stable pods | Canary pods | Expected canary share | |---|---|---| | 47 | 1 | 1/48, about 2.08% | | 99 | 1 | 1/100 = 1% | | 199 | 1 | 1/200 = 0.5% | | 47 | 5 | 5/52, about 9.6% | The steps follow from that: 1. Deploy `autocomplete-canary` with 1 replica and the new image. 2. Wait for it to be available, then watch the canary's metrics, filtered by `track: canary`, for the bake period. 3. To widen exposure, scale the canary up (and optionally stable down). 4. To promote, set the new image on `autocomplete-stable`, which rolls it out, then scale the canary to zero. 5. To abort, scale the canary to zero. Its pods leave the Service's endpoints and traffic returns to stable. ## Where it breaks down - **Granularity.** The share can only move in steps of one pod. The search namespace runs 1,180 pods, but autocomplete is only 48 of them, so 2% is the smallest step. Getting to 0.5% would mean running 200 autocomplete pods, far beyond what the load needs. - **Per-connection balancing.** Clients using HTTP keep-alive or gRPC open a few long-lived connections and send many requests over each. If a busy client's connection lands on the canary, the canary's request share can be many times 2%. If none does, it can be close to zero. - **Autoscaling drift.** A HorizontalPodAutoscaler on `autocomplete-stable` that goes from 47 to 61 replicas quietly cuts the canary's share from about 2.1% to about 1.6%. An HPA on the canary itself can do the opposite. - **Session affinity.** With `sessionAffinity: ClientIP` on the Service, each client IP sticks to one pod for a while (10800 seconds by default), so exposure follows client IPs, not requests. - **No targeting.** You cannot send only internal users or a particular header to the canary; the Service has no idea what is inside a request. - **Capacity coupling.** Share and capacity are the same knob. A canary at a tiny share still needs enough pods to be redundant, and adding pods raises its share. ## When it is still good enough Replica-ratio canaries are fine when a few percent is an acceptable first step, the service has enough replicas for that step, and clients open short connections. They need nothing beyond core Kubernetes, which matters on a small bare-metal cluster. When any of the limits above bites, the fix is to split traffic by weight at a proxy (a Gateway API route, an Ingress controller's own feature or a service mesh), so share no longer depends on pod count. ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: autocomplete-canary spec: replicas: 1 selector: matchLabels: app: autocomplete track: canary template: metadata: labels: app: autocomplete track: canary spec: containers: - name: autocomplete image: registry.example.com/search/autocomplete:4.13.0 ```

  • Your canary has 1 of 48 pods but serves about 20% of requests. What is the likely cause?
    kube-proxy balances per connection. A few high-volume clients hold long-lived keep-alive or gRPC connections, and one of them landed on the canary pod, so it carries far more than its 1/48 share of requests. Check per-pod request rates and client connection reuse, and move the split to a proxy that balances per request if the share has to be predictable.
  • Why should each Deployment's selector include the track label even though the Service ignores it?
    Without it, both Deployments would have the same selector, `app: autocomplete`, and each would match the other's pods. The Kubernetes documentation warns against overlapping controller selectors because the results are confusing. The `track` label keeps each Deployment's pods its own and also lets metrics separate canary from stable.

saying these in an interview costs you the question

  • One canary pod out of 48 gives exactly 2% of requests
  • kube-proxy weights endpoints by their CPU requests
  • Both Deployments can use the same selector as the Service
  • A 0.5% canary just needs one canary pod
  • Autoscaling the stable Deployment leaves the canary share unchanged