skip to content

Deployment Strategies

Blue-green, rolling, canary and feature-flagged releases, all aimed at limiting blast radius and making rollback boring. You will learn how progressive delivery lets you validate a release against real traffic before everyone gets it.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a blue-green deployment, two identical production environments (blue = live, green = idle) exist side by side. Explain how traffic cutover works and what happens if the new version needs to be rolled back.

level: juniorimportance: must knowfreq 75%

answer

  1. blue=live green=idle
  2. instant router flip
  3. expand-contract schema
  4. 2x infra cost
  5. instant rollback

basics

~20 s

You run two copies of the app - one live, one idle with the new version. Once the idle one is tested and ready, you flip a switch (router/load balancer) so all traffic goes to it. If something breaks, you flip back instantly.

solid answer

~30 s

Blue-green keeps two full environments: blue currently serving traffic, green idle. You deploy the new version to green, run smoke/health checks against it while it's isolated from users, then cut over traffic at the router/load-balancer/DNS layer - either instantly or via a brief drain. Rollback is just flipping the router back to blue, which still has the old version warm and running, giving near-instant recovery. The cost is running 2x infrastructure during the switch window and needing to handle any state/DB schema shared between both versions (backward-compatible migrations).

go deeper

for a junior

Can describe the two-environment idea and that rollback is fast because the old version is still running.

for a middle

Explains the cutover mechanism concretely (LB/router/service selector) and the double-infra cost.

for a senior

Discusses shared-state/schema compatibility (expand-contract), connection draining, and cold-start effects post-cutover.

for a principal

Weighs blue-green against canary/rolling for cost-at-scale across many independently-deployed microservices, and designs org-wide migration/rollback conventions.

## How the cutover works Blue-green deployment maintains two complete, independently running production environments. **'Blue'** is currently serving all live traffic; **'green'** is an idle but fully provisioned copy where the new version is deployed. The new version is tested against green in isolation — smoke tests, health checks, sometimes synthetic traffic — without any real users touching it. Once confidence is high, traffic is cut over at a routing layer: - a **load balancer** swapping target groups - a **Kubernetes Service** selector changing which Deployment's pods it points at - a **DNS record** update (slower, due to TTL caching) The switch is typically near-instantaneous from the router's perspective, and in-flight requests to blue are usually allowed to drain via a short connection-draining window before blue is fully removed from rotation. ## Why the technique exists The technique exists to solve two problems at once: 1. **Eliminating deploy-time downtime.** 2. **Making rollback essentially free.** Because blue keeps running, unmodified, through the entire cutover, reverting a bad release is just flipping the router back — no redeploy, no rebuild, no waiting for old code to come back up cold. This is fundamentally different from deployment strategies that update instances in place, where 'rollback' means running the deployment process again in reverse. ## The trade-off The core trade-off is cost and realism. - **Cost.** Running two full environments means paying for roughly double the infrastructure for the duration of the switch window (or needing the ability to rapidly provision a second environment on demand, which many cloud platforms support). - **Realism.** Blue-green also doesn't test the new version under a genuine mix of real production traffic before the cutover — unlike canary, there's no gradual, partial exposure to catch load-dependent or data-dependent bugs before all users are affected; the switch is all-or-nothing. - **Shared state.** Because both blue and green commonly point at the same shared database, any schema change has to remain compatible with whichever version might be running — which is not guaranteed by the blue-green pattern itself and has to be engineered deliberately. ## Failure modes 1. The most consequential failure mode is exactly that **shared-state** problem: if a deploy to green includes a destructive schema migration (dropping a column, renaming a table) and that migration runs against the shared database, then rolling back to blue doesn't actually restore a working system, because blue's code still expects the old schema that no longer exists. This is why teams pair blue-green with the **expand-contract (parallel change)** pattern — ship a backward-compatible schema change first, deploy the new code, and only remove old columns/tables in a later, separate release once rollback is no longer a live possibility. 2. A second failure mode is a **post-cutover latency or error spike** even though the switch itself was instant: green's instances are cold — JIT warmup, empty connection pools, cold caches, autoscaler-provisioned nodes not yet at steady state — so the receiving fleet can be under-warmed even though routing changed immediately. 3. A third is **DNS-based cutovers** appearing to 'not take effect' for some users due to client-side or resolver caching beyond the configured TTL. ## Where it shows up A concrete real-world instance of this pattern: **AWS Elastic Beanstalk** and **CodeDeploy** both offer built-in blue-green deployment modes that provision a parallel environment and swap either a load balancer target group or an environment CNAME on cutover; teams running **Kubernetes** commonly implement the same idea manually with two Deployments and a Service whose selector is updated to point from one to the other. Because of its double-infrastructure cost, blue-green tends to be reserved for services where instant, guaranteed-clean rollback matters more than infrastructure efficiency — critical customer-facing services, or releases considered risky enough that the team wants zero exposure window rather than a gradually widening one.

  • If the two environments share a single database, what pattern lets you roll back safely after a schema change?
    Use the expand-contract (parallel change) pattern: first deploy a backward-compatible schema change that both old and new code can read/write, then deploy the new code, and only drop old columns/tables in a later, separate release once you're confident you won't roll back. This way blue can keep running against the same schema green uses.
  • How do you handle in-flight requests during the cutover?
    Drain connections from blue before fully redirecting: stop sending new requests to blue, let existing requests finish (a short drain timeout), then decommission or keep it warm as the rollback target. Load balancers and service meshes support connection draining natively.
  • Why might a 'blue-green' deployment on Kubernetes still cause a brief latency spike?
    Because the green pods start cold - JIT warmup, connection pools, caches, and autoscaler-provisioned nodes all need to reach steady state, so the instant traffic switch can hit an underwarmed fleet even though the switch itself is instantaneous.

Like keeping a fully-stocked spare store next door and swapping which one's front door sign says 'open' - the backup was already running before the switch.

saying these in an interview costs you the question

  • says it eliminates all downtime with no mention of shared state
  • doesn't mention rollback speed/mechanism
  • conflates blue-green with rolling deployment
  • ignores double infrastructure cost
  • assumes stateless services only, doesn't consider DB migrations

context

open as a page

Describe how a canary deployment progressively shifts traffic to a new service version, and what signals should automatically trigger a rollback.

level: middleimportance: must knowfreq 85%

basics

~20 s

You send a new version to a small slice of real users first (like 5%), watch error rates and latency, and only widen the slice if it looks healthy. If it looks bad, you send everyone back to the old version.

open as a page

How do feature flags let a team decouple 'deploying code' from 'releasing a feature,' and what operational costs does relying heavily on flags introduce?

level: seniorimportance: must knowfreq 80%

basics

~20 s

Feature flags are on/off switches in the code. You can ship new code to production turned off, then flip it on for some or all users later without a new deployment - and flip it back off fast if it breaks.

open as a page

In a rolling deployment, instances of a service are updated a few at a time while old and new versions both serve traffic simultaneously. What risks does this mixed-version window create, and how do you mitigate them?

level: middleimportance: should knowfreq 70%

basics

~20 s

You replace old servers with new ones in small batches instead of all at once, so there's no downtime, but for a while some users hit the old version and some hit the new one - they both need to work together correctly.

open as a page

Design an automated rollback system for a service deployment pipeline: what should trigger it, and what can cause an automated rollback itself to fail or make things worse?

level: seniorimportance: should knowfreq 65%

basics

~10 s

The pipeline watches health metrics right after a deploy, and if things look bad (errors spike, requests fail), it automatically switches back to the previous version without waiting for a human to notice.

open as a page

'Progressive delivery' combines canary-style traffic control, feature flags, and automated analysis into one release process. As a principal engineer choosing a deployment strategy per service, when would you deliberately NOT use canary or progressive delivery, even though the tooling is available?

level: principalimportance: nice to knowfreq 45%

basics

~20 s

Sometimes a fancy gradual rollout isn't worth it - for tiny low-traffic services, batch jobs, or changes where a database migration already commits you either way, a simpler rolling deploy with good tests is safer and cheaper than building a whole canary pipeline.

open as a page