In a blue-green deployment, two identical production environments (blue = live, green = idle) exist side by side. Explain how traffic cutover works and what happens if the new version needs to be rolled back.
answer
- blue=live green=idle
- instant router flip
- expand-contract schema
- 2x infra cost
- instant rollback
basics
~20 sYou run two copies of the app - one live, one idle with the new version. Once the idle one is tested and ready, you flip a switch (router/load balancer) so all traffic goes to it. If something breaks, you flip back instantly.
solid answer
~30 sBlue-green keeps two full environments: blue currently serving traffic, green idle. You deploy the new version to green, run smoke/health checks against it while it's isolated from users, then cut over traffic at the router/load-balancer/DNS layer - either instantly or via a brief drain. Rollback is just flipping the router back to blue, which still has the old version warm and running, giving near-instant recovery. The cost is running 2x infrastructure during the switch window and needing to handle any state/DB schema shared between both versions (backward-compatible migrations).
go deeper
Can describe the two-environment idea and that rollback is fast because the old version is still running.
Explains the cutover mechanism concretely (LB/router/service selector) and the double-infra cost.
Discusses shared-state/schema compatibility (expand-contract), connection draining, and cold-start effects post-cutover.
Weighs blue-green against canary/rolling for cost-at-scale across many independently-deployed microservices, and designs org-wide migration/rollback conventions.
## How the cutover works Blue-green deployment maintains two complete, independently running production environments. **'Blue'** is currently serving all live traffic; **'green'** is an idle but fully provisioned copy where the new version is deployed. The new version is tested against green in isolation — smoke tests, health checks, sometimes synthetic traffic — without any real users touching it. Once confidence is high, traffic is cut over at a routing layer: - a **load balancer** swapping target groups - a **Kubernetes Service** selector changing which Deployment's pods it points at - a **DNS record** update (slower, due to TTL caching) The switch is typically near-instantaneous from the router's perspective, and in-flight requests to blue are usually allowed to drain via a short connection-draining window before blue is fully removed from rotation. ## Why the technique exists The technique exists to solve two problems at once: 1. **Eliminating deploy-time downtime.** 2. **Making rollback essentially free.** Because blue keeps running, unmodified, through the entire cutover, reverting a bad release is just flipping the router back — no redeploy, no rebuild, no waiting for old code to come back up cold. This is fundamentally different from deployment strategies that update instances in place, where 'rollback' means running the deployment process again in reverse. ## The trade-off The core trade-off is cost and realism. - **Cost.** Running two full environments means paying for roughly double the infrastructure for the duration of the switch window (or needing the ability to rapidly provision a second environment on demand, which many cloud platforms support). - **Realism.** Blue-green also doesn't test the new version under a genuine mix of real production traffic before the cutover — unlike canary, there's no gradual, partial exposure to catch load-dependent or data-dependent bugs before all users are affected; the switch is all-or-nothing. - **Shared state.** Because both blue and green commonly point at the same shared database, any schema change has to remain compatible with whichever version might be running — which is not guaranteed by the blue-green pattern itself and has to be engineered deliberately. ## Failure modes 1. The most consequential failure mode is exactly that **shared-state** problem: if a deploy to green includes a destructive schema migration (dropping a column, renaming a table) and that migration runs against the shared database, then rolling back to blue doesn't actually restore a working system, because blue's code still expects the old schema that no longer exists. This is why teams pair blue-green with the **expand-contract (parallel change)** pattern — ship a backward-compatible schema change first, deploy the new code, and only remove old columns/tables in a later, separate release once rollback is no longer a live possibility. 2. A second failure mode is a **post-cutover latency or error spike** even though the switch itself was instant: green's instances are cold — JIT warmup, empty connection pools, cold caches, autoscaler-provisioned nodes not yet at steady state — so the receiving fleet can be under-warmed even though routing changed immediately. 3. A third is **DNS-based cutovers** appearing to 'not take effect' for some users due to client-side or resolver caching beyond the configured TTL. ## Where it shows up A concrete real-world instance of this pattern: **AWS Elastic Beanstalk** and **CodeDeploy** both offer built-in blue-green deployment modes that provision a parallel environment and swap either a load balancer target group or an environment CNAME on cutover; teams running **Kubernetes** commonly implement the same idea manually with two Deployments and a Service whose selector is updated to point from one to the other. Because of its double-infrastructure cost, blue-green tends to be reserved for services where instant, guaranteed-clean rollback matters more than infrastructure efficiency — critical customer-facing services, or releases considered risky enough that the team wants zero exposure window rather than a gradually widening one.
- If the two environments share a single database, what pattern lets you roll back safely after a schema change?Use the expand-contract (parallel change) pattern: first deploy a backward-compatible schema change that both old and new code can read/write, then deploy the new code, and only drop old columns/tables in a later, separate release once you're confident you won't roll back. This way blue can keep running against the same schema green uses.
- How do you handle in-flight requests during the cutover?Drain connections from blue before fully redirecting: stop sending new requests to blue, let existing requests finish (a short drain timeout), then decommission or keep it warm as the rollback target. Load balancers and service meshes support connection draining natively.
- Why might a 'blue-green' deployment on Kubernetes still cause a brief latency spike?Because the green pods start cold - JIT warmup, connection pools, caches, and autoscaler-provisioned nodes all need to reach steady state, so the instant traffic switch can hit an underwarmed fleet even though the switch itself is instantaneous.
Like keeping a fully-stocked spare store next door and swapping which one's front door sign says 'open' - the backup was already running before the switch.
saying these in an interview costs you the question
- says it eliminates all downtime with no mention of shared state
- doesn't mention rollback speed/mechanism
- conflates blue-green with rolling deployment
- ignores double infrastructure cost
- assumes stateless services only, doesn't consider DB migrations