skip to content

As the owner of a shared Kubernetes platform, when do hand-built canary and blue-green releases from Deployments, Services and HTTPRoute weights stop being enough, and what replaces them?

level: principalimportance: should knowfreq 34%

answer

  1. traffic moves, nothing judges it
  2. rolled out means readiness passed
  3. who aborts at 3am
  4. scripts drift across teams
  5. a controller is another control plane

basics

~20 s

Hand-built releases suffice while few services release rarely under a watching person. When many teams need automatic bake-and-abort, precise splits and an audit trail, a progressive-delivery controller or mesh should own the steps; core Kubernetes never judges a release.

solid answer

~50 s

Core Kubernetes checks only that pods are available: `kubectl rollout status` saying 'successfully rolled out' means readiness passed, not that errors stayed flat. With primitives, the bake-and-abort loop is a person or a script comparing canary metrics, split by a `track` label, and then zeroing a weight or flipping a selector. That works for a handful of watched services. It stops being enough when releases are frequent, when many teams each build their own scripts, when you need very small or per-request splits, header targeting or east-west coverage, or when rollback must happen at 3am without anyone awake. Then I would offer one supported path: a progressive-delivery controller (the Argo Rollouts or Flagger class) or a mesh, with a shared definition of the abort signal. The cost is another control plane and new objects for GitOps and runbooks, so plain Deployments stay for services that do not need it.

go deeper

for a junior

Recall that Kubernetes reports a rollout as done when pods are available, and that it does not check error rates.

for a middle

Explain the hand-built loop: a track label, metrics split by it, and aborting by zeroing a weight, scaling to zero or flipping a selector.

for a senior

Show you have run releases: know the signals that should abort one, why unattended rollback fails without automation, and where gateway splits miss traffic.

for a principal

Own the platform call: which services get a controller, what the shared abort signal is, and what running another control plane will cost the teams.

## What Kubernetes does and does not judge Everything built from primitives, whether a selector flip, a replica ratio or HTTPRoute weights, moves traffic. None of it decides whether the new version is **good**. Core Kubernetes only knows: - whether containers pass their **readiness** probes, which gates Service endpoints; - whether a Deployment has enough **available** replicas, which is what `kubectl rollout status` waits for before printing 'successfully rolled out'; - whether a rollout has stopped progressing: past `progressDeadlineSeconds` the Deployment controller sets the `Progressing` condition to `False` with reason `ProgressDeadlineExceeded`, and it does **not** roll anything back. A search-autocomplete canary can pass all three while returning errors on a few percent of queries or doubling tail latency. Deciding when to advance and when to abort is **bake time** plus a **rollback signal**, and neither exists in core Kubernetes. ## The hand-built loop With primitives, a release looks like this: 1. Start the canary with its own `track: canary` label and a small weight or pod ratio. 2. Watch a metrics backend, split by the `track` label, for the bake period. 3. Decide, by a person or a script, whether error rate or latency is acceptable. 4. Advance by raising the weight or pod count, or abort by setting the weight to 0, scaling the canary to 0, or flipping the selector back. 5. Record what happened somewhere, since nothing in the cluster keeps that history. Each step is simple. The problem is who does it, how consistently, and how fast when it matters. ## Signs the primitives are no longer enough | Pressure | Why primitives struggle | |---|---| | Many teams, frequent releases | Every team writes its own scripts; abort rules drift apart and nobody can audit them. | | Unattended rollback | A person has to see the signal and act; an overnight release has no one watching. | | Very small or precise shares | Replica ratios cannot go below one pod; weights help only at the gateway. | | Header or cohort targeting | Services cannot look at requests; you need route rules and a way to manage them per release. | | East-west callers | Gateway weights miss pod-to-pod traffic; covering it needs a mesh. | | Evidence for change review | Nothing records which step failed and why. | On a 9-node bare-metal cluster running a 1,180-pod search namespace, the capacity question also changes the answer. Blue-green doubles a service's pods on machines you cannot add quickly, so a weighted canary with a small pod count is often the only release style that fits. ## What replaces them, and at what cost The usual replacement is a **progressive-delivery controller** (Argo Rollouts and Flagger are the common ones) or a **service mesh**, driven by a shared definition of the abort signal. Their mechanics belong to their own topics. What matters at platform level is the trade: - **Gains:** one tested release loop, automatic abort on a defined signal, consistent step sizes, and a record of each release on the release object. - **New control plane:** another controller to upgrade, monitor and keep available, and its failure modes become release failures. - **New objects:** some controllers replace the Deployment with their own resource, others manage extra Deployments beside yours. Either way, GitOps diffs, dashboards and runbooks must learn them. - **Signal quality becomes the bottleneck:** an automated abort is only as good as its metrics. A noisy query aborts good releases, and a missing label hides bad ones. - **Two models in the cluster:** teams that stay on plain Deployments and teams on the controller need different runbooks. ## How I would decide - Keep plain rolling updates for internal, low-risk or rarely changed services. - Allow hand-built canaries with a documented script and a required `track` label where a team releases rarely and watches each step. - Offer the controller as the supported path for high-traffic, user-facing services such as autocomplete, with a shared, reviewed abort signal. - Decide east-west coverage explicitly, rather than assuming a gateway split covers it. - Revisit when the number of teams, release frequency or incident history changes.

  • A team says their Deployment's progressDeadlineSeconds protects the canary. Does it?
    No. When a rollout stops progressing past that deadline, the Deployment controller sets the `Progressing` condition to `False` with reason `ProgressDeadlineExceeded` and takes no other action. It also only notices pods that fail to become available, not pods that are available and returning errors. Aborting on errors still needs an external signal and something that acts on it.
  • Why insist on a track label even for teams that keep hand-built canaries?
    Without a label that separates canary pods from stable ones, metrics mix both versions and a 2% canary's errors vanish in the average. The label is also what an abort script or a later controller uses to find the canary, so requiring it keeps today's scripts and tomorrow's automation looking at the same thing.

saying these in an interview costs you the question

  • 'successfully rolled out' means the new version is healthy
  • progressDeadlineSeconds rolls a failing Deployment back automatically
  • A liveness probe will catch a version that returns wrong results
  • Every service should go through the progressive-delivery controller
  • Gateway weights give full canary coverage, including internal callers