How does gateway routing make blue-green deployments and canary releases possible, mechanically?
answer
- two live targets, one routing rule
- blue-green = atomic flip + instant rollback
- canary = weighted % + sticky routing
- watch metrics on the slice before widening it
basics
~20 sThe gateway can send a slice of traffic, or all of it after a flip, to a new version instead of the old one, just by changing routing rules. Clients never need to change since they call the same address.
solid answer
~40 sBoth rely on the gateway holding two live backend targets for one logical service, an old ('blue'/stable) version and a new ('green'/canary) version, and controlling which requests reach which target purely through routing rules. In blue-green, traffic stays 100% on blue while green is deployed and smoke-tested out of band, then cutover is a single atomic routing-table flip to green, with blue kept warm as an instant rollback target. In canary, the gateway splits traffic by percentage (weighted routing) or a targeting signal like a header or cohort, sending a small fraction to the new version while watching error rates and latency, then progressively widening that fraction to 100%. Both work because the routing table, not the client, is the single control point deciding which version handles a request.
go deeper
Should describe, roughly, that a new version can get a small slice of traffic before going fully live, and that the old version stays available as a fallback.
Should explain both mechanisms concretely: atomic flip for blue-green, weighted/percentage split for canary, and why routing (not redeploying) is the control point.
Should discuss sticky routing, metric-driven widening/halting of a canary, and the resource cost trade-off between the two strategies.
Should reason about combining the two (canary within a blue-green target), automated rollback tied to SLO burn or alarm thresholds, and the org-level process (who owns the widening decision, what gates a promotion) needed to run this safely at scale.
## One capability, two strategies Blue-green deployments and canary releases are both, mechanically, applications of the same underlying capability: the gateway's routing table can point a given path or host at more than one backend target, and which target actually serves a given request can be changed without touching any client. That single fact is what turns 'deploy a new version' from an all-or-nothing, client-visible event into a controlled, reversible, gateway-side operation. ## Blue-green: the atomic flip In a blue-green rollout, two full, independently running deployments of a service coexist: - **blue** is the currently live version handling all production traffic; - **green** is the newly deployed version, fully started up and warm but receiving no real user traffic yet. The team runs smoke tests and health checks directly against green (bypassing the gateway, or via a separate internal route), and once satisfied, performs the cutover by updating a single routing rule so the path that used to resolve to blue's backend pool now resolves to green's. Because this is one atomic change to the routing table rather than a rolling redeploy of instances, the cutover is effectively instantaneous from the client's perspective and, critically, **reversible**: if green misbehaves under real traffic, flipping the same rule back to blue is just as fast as the forward flip was, and blue's instances are still running and warm, not torn down. The main resource cost is running two full production-capacity deployments simultaneously for the duration of the rollout. ## Canary: the weighted split Canary releases use the same routing infrastructure but split traffic gradually rather than flipping it wholesale. The gateway is configured with **weighted routing**: for example, 95% of requests to the target's original rule, 5% to the new version, using either a random-weighted selection per request or a deterministic signal like a header, cookie, or hashed user ID so that the same user consistently lands on the same variant (**sticky canary membership**). The team watches error rates, latency percentiles, and business metrics on the 5% slice; if the canary looks healthy, the weight is progressively increased (`5% -> 25% -> 50% -> 100%`), and if it looks unhealthy, the weight is dropped back to zero, again as a routing-table change rather than a redeploy. The benefit over a blue-green flip is **blast-radius control**: a bug only affects a bounded fraction of real users at a time, and it surfaces on live traffic and live data rather than only in a smoke test, which catches classes of bugs (concurrency issues, data-dependent edge cases, cache interactions) that a pre-cutover smoke test typically misses. ## The trade-off The trade-off between the two is what each optimizes for. | Strategy | Optimizes for | At the cost of | |---|---|---| | **Blue-green** | rollback speed and simplicity: it's a single flip in either direction, easy to reason about | it exposes 100% of users to the new version the instant it's live, so any bug the smoke tests missed is immediately a full-scale incident | | **Canary** | blast-radius containment | complexity: you need percentage-based or cohort-based routing support in the gateway, metrics granular enough to distinguish canary-slice health from overall health, and usually sticky routing so a given user doesn't flip between old and new behavior mid-session, which is itself an additional piece of state the gateway or an associated layer has to track | ## Failure modes Failure modes specific to this use of gateway routing include: - **session inconsistency**, where a canary lacks sticky routing and a user's requests bounce between old and new backend versions within one session, producing confusing partial-state bugs that are hard to reproduce; - **metric dilution**, where the canary's 5% slice is too small relative to the monitoring system's granularity to detect a real regression before it's rolled out further; - and **stale routing-table propagation**, where a weighted rule change takes longer to reach all gateway instances than expected (in a horizontally scaled gateway fleet with eventually-consistent config distribution), so some fraction of gateway nodes are still serving the old weights well after the rollout dashboard reports the change as live. ## Where you see it A concrete real-world pattern: a service mesh like Istio expresses this directly as VirtualService traffic-splitting rules with weight fields per destination subset, and AWS CodeDeploy's blue-green ECS/Lambda deployment strategies manipulate ALB target-group weights the same way, shifting traffic in configurable increments (linear or canary) with automatic rollback tied to CloudWatch alarms if error-rate thresholds are breached during the shift.
- Why is sticky routing important for a canary release but not usually necessary for a blue-green cutover?In a canary, users are randomly split between two live versions, so without stickiness a single user's requests can bounce between old and new backend behavior within one session, which can corrupt session state or produce inconsistent UI. Blue-green flips everyone to the same version atomically, so there's no simultaneous split to be inconsistent about, every request after the flip goes to the same target.
- What's the cost of running a blue-green deployment compared to a canary, in terms of infrastructure?Blue-green requires running two full production-capacity deployments simultaneously for the rollout window, since 100% of traffic can hit either one at any time, which roughly doubles compute cost during that window. A canary only needs the new version scaled to handle its current traffic slice (5%, then 25%, etc.), so it can start much smaller and scale up as its weight increases.
- How would you detect that a canary rollout is unhealthy before widening it to more traffic?Compare error rate, p99 latency, and key business metrics on the canary slice against the same metrics on the stable slice over the same time window, rather than looking at the canary's absolute numbers alone. A statistically significant regression on the canary slice relative to the baseline, even a small one, is the trigger to halt or roll back the weight rather than proceed.
Blue-green is like switching a train onto a second, fully built parallel track with one lever pull, instantly reversible. Canary is like letting a small trickle of water down a new pipe before opening the main valve, watching for leaks while only a little water is at risk.
saying these in an interview costs you the question
- Thinks blue-green and canary are the same technique with different names
- Doesn't mention that both rely on the gateway holding two live backend targets simultaneously
- Assumes canary rollout requires redeploying the client
- No awareness of sticky routing / session consistency issues in canary splits
- Can't explain how rollback works for either strategy