skip to content

In Kubernetes, how do you build a blue-green deployment from two Deployments and one Service, and how do cutover and rollback work?

level: juniorimportance: must knowfreq 68%

answer

  1. two Deployments, one Service
  2. one label tells the colours apart
  3. the switch lives in spec.selector
  4. new connections move, open ones stay
  5. keep blue warm to roll back

basics

~20 s

Run blue and green as two Deployments whose pods differ by one label, such as slot, and point one Service at blue. Cutover patches the Service selector to green; rollback patches it back while blue is still running.

solid answer

~40 s

Kubernetes Deployments only offer `RollingUpdate` and `Recreate`, so blue-green is built by hand. You run two Deployments, `autocomplete-blue` and `autocomplete-green`, whose pods share `app: autocomplete` and differ by a label such as `slot: blue` or `slot: green`. The production Service selects `app: autocomplete, slot: blue`. You roll out green, wait until `kubectl rollout status` reports it available, smoke-test it through a separate preview Service, then run `kubectl patch service` to change the selector to `slot: green`. The EndpointSlice controller swaps the backends and kube-proxy reprograms every node, so **new** connections go to green within seconds. TCP connections that are already open stay on blue until they close. Rollback is the same patch in reverse, which is only instant if blue is still scaled up, so you keep it running for a rollback window.

code

yaml · 33 lines
yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: autocomplete-green
spec:
  replicas: 38
  selector:
    matchLabels:
      app: autocomplete
      slot: green
  template:
    metadata:
      labels:
        app: autocomplete
        slot: green
    spec:
      containers:
        - name: autocomplete
          image: registry.example.com/search/autocomplete:4.12.3
          ports:
            - containerPort: 7311
---
apiVersion: v1
kind: Service
metadata:
  name: autocomplete
spec:
  selector:
    app: autocomplete
    slot: blue
  ports:
    - port: 80
      targetPort: 7311

go deeper

for a junior

Recall the parts: two Deployments with a colour label, one Service, and a selector patch for cutover and rollback. Be able to write the patch command.

for a middle

Explain the path from the selector patch to the EndpointSlice update and kube-proxy reprogramming, and why open TCP connections stay on blue.

for a senior

Show the operating discipline: preview Service, rollback window with blue kept warm, draining long-lived connections, and schema changes that work with both colours.

for a principal

Weigh double capacity on fixed bare-metal nodes and all-or-nothing exposure against a canary, and decide which services deserve blue-green at all.

## Why blue-green has to be assembled A Kubernetes **Deployment** has exactly two built-in `strategy.type` values: `RollingUpdate` (the default) and `Recreate`. Neither keeps two complete versions side by side and switches traffic between them in one step. **Blue-green** is that pattern: a live environment (blue) and a fully deployed but idle environment (green), with one switch deciding which one serves. Kubernetes has no blue-green object, but it has every piece you need to build one: - a **Deployment** per colour, each managing its own ReplicaSet and pods; - **labels** on the pod template that say which colour a pod belongs to; - a **Service** whose `spec.selector` decides which pods receive traffic; - the **EndpointSlice controller**, which keeps the list of backends behind a Service in step with its selector. ## The topology Take a search-autocomplete service running 38 replicas on a 9-node bare-metal cluster. The layout is: | Object | Labels or selector | Role | |---|---|---| | Deployment `autocomplete-blue` | pods carry `app: autocomplete`, `slot: blue` | current version | | Deployment `autocomplete-green` | pods carry `app: autocomplete`, `slot: green` | new version | | Service `autocomplete` | selects `app: autocomplete`, `slot: blue` | production traffic | | Service `autocomplete-preview` | selects `app: autocomplete`, `slot: green` | smoke tests only | Each Deployment's own `spec.selector` includes the `slot` label, so the two Deployments never select each other's pods. A Deployment's selector must match its pod template's labels and cannot be changed after creation, so the colour label has to be there from the start. ## Cutover, step by step 1. Update `autocomplete-green` to the new image (or create it) and wait for `kubectl rollout status deployment/autocomplete-green` to report success. 2. Test green through `autocomplete-preview`, which production clients never use. 3. Patch the production Service's selector from `slot: blue` to `slot: green`. 4. The EndpointSlice controller rewrites the EndpointSlices labelled `kubernetes.io/service-name=autocomplete` to list green pod IPs. 5. kube-proxy on every node sees the change and reprograms its rules, so new connections to the ClusterIP land on green pods. 6. Watch error rate and latency for green. Keep blue at full size until you are confident. The switch is a single API write, so all **new** connections move together. There is no stretch of time where the Service's backends are a mix of old and new pods, which is the main thing a rolling update cannot give you. ## Rollback Rollback is step 3 in reverse: patch the selector back to `slot: blue`. It is only fast because blue's pods are still running and Ready. If you scaled blue to zero right after the flip, rolling back means waiting for 38 pods to be scheduled, pull images and pass readiness, and that is not blue-green any more. Teams set a rollback window, then scale blue to zero or reuse it as the next release's idle colour. ## What the flip does not do | Assumption | Reality | |---|---| | Every request moves to green at once | Only new connections move. In iptables mode kube-proxy leaves existing TCP conntrack entries alone, so keep-alive and gRPC clients keep using blue pods until those connections close. | | The flip checks green is healthy | The API server accepts any selector. Pointing it at pods that are not Ready leaves the Service with no ready endpoints, and connections fail. | | Clients must re-resolve DNS | The Service's ClusterIP and DNS name do not change; only its backend list does. | | Data is switched too | Both colours usually share one database, so a schema change must work with both versions. | ## What it costs - **Double capacity** during the window: 38 more autocomplete pods in a namespace that already runs 1,180 pods, and their resource requests have to fit on nine fixed machines, which cannot add nodes on demand. - **All-or-nothing exposure**: every new user hits green at once, so a defect reaches everyone. A canary limits that. - **Manual bookkeeping**: nothing in Kubernetes records which colour is live other than the selector itself, so the procedure belongs in a script or a delivery tool. ```bash kubectl rollout status deployment/autocomplete-green kubectl patch service autocomplete -p '{"spec":{"selector":{"app":"autocomplete","slot":"green"}}}' kubectl get endpointslices -l kubernetes.io/service-name=autocomplete ```

  • After the selector flip, some clients still hit blue pods ten minutes later. Why, and what do you do?
    They hold long-lived TCP connections, such as HTTP keep-alive or gRPC channels. kube-proxy picks a backend only when a connection opens, and in iptables mode it leaves established TCP flows alone, so those clients stay on blue until the connection closes. You keep blue running until its traffic drains, or make blue close idle connections on a timer or send `Connection: close`, so clients reconnect and land on green.
  • Why not just edit the image in one Deployment and call that blue-green?
    Changing the image starts a rolling update. The Deployment creates a new ReplicaSet and shifts pods over gradually, so old and new versions serve side by side for a while, and there is no single switch to flip back. Blue-green needs both versions fully up at the same time, which takes two Deployments and a selector deciding which one is live.
  • How do you test green before any production client reaches it?
    Create a second Service, such as `autocomplete-preview`, that selects `slot: green`. Smoke tests and internal checks call that Service while production clients keep using the main Service, which still selects blue. Because both Services find pods by label, green's pods serve the preview traffic with no other change.

It is like a railway switch in front of two finished platforms: moving the lever sends every new train to the other platform, but a train already in the station stays where it is.

saying these in an interview costs you the question

  • Blue-green is a built-in Deployment strategy type in Kubernetes
  • Changing the Service selector moves every open connection to green at once
  • The API server refuses a selector that matches no Ready pods
  • Scaling blue to zero right after the flip keeps rollback instant
  • The Service gets a new ClusterIP, so clients must re-resolve DNS