In a rolling deployment, instances of a service are updated a few at a time while old and new versions both serve traffic simultaneously. What risks does this mixed-version window create, and how do you mitigate them?
answer
- batch replace old with new
- maxSurge/maxUnavailable
- readiness probe gates next batch
- mixed-version window guaranteed
- contract must span N and N+1
basics
~20 sYou replace old servers with new ones in small batches instead of all at once, so there's no downtime, but for a while some users hit the old version and some hit the new one - they both need to work together correctly.
solid answer
~30 sRolling deployment replaces instances/pods in batches, each new instance passing a readiness/health check before the next batch starts and the old one is terminated. It avoids blue-green's double-infrastructure cost and has no hard traffic cutover, but during the rollout both old and new versions are live simultaneously, so the new version's API/message contracts and shared database schema must be backward- and forward-compatible - otherwise you get intermittent errors depending on which instance a request lands on. Kubernetes' default Deployment strategy is rolling update, controlled by maxSurge/maxUnavailable, and it pauses if new pods don't become ready.
go deeper
Knows rolling deployment replaces instances gradually instead of all at once, avoiding downtime.
Can name the batch/readiness-probe mechanism and Kubernetes-style controls (maxSurge/maxUnavailable).
Explains why contract compatibility across exactly one version step is mandatory, and diagnoses mixed-version bugs from their intermittent, instance-correlated symptoms.
Weighs rolling vs canary/blue-green per service risk profile, sets org contract-testing/expand-contract policy, and designs rollback-time SLAs given rolling rollback's slower, doubly-exposed nature.
## How the rollout proceeds A rolling deployment updates a fleet of service instances incrementally rather than all at once. 1. Given N running instances of the old version, the orchestrator spins up a batch of new-version instances — the batch size controlled by parameters like Kubernetes' **maxSurge** and **maxUnavailable**. 2. Each new instance must pass a **readiness probe** before the orchestrator registers it with the load balancer and begins routing traffic to it. 3. Only once a batch of new instances is healthy does the orchestrator terminate an equivalent batch of old instances, then repeats until every old instance has been replaced. | Parameter | What it controls | |---|---| | maxSurge | how many extra instances beyond N can exist temporarily | | maxUnavailable | how many can be offline at once | There's no dedicated idle environment and no explicit traffic-percentage control the way canary has — the split between old and new at any moment is just whatever fraction of instances have been updated so far. ## Why the technique exists Rolling deployment exists to achieve zero-downtime deploys without paying for a full duplicate environment (blue-green) and without operating a separate traffic-splitting/metrics-gating system (canary) — it's the default, low-ceremony option built into most orchestrators, which is why it's Kubernetes' out-of-the-box Deployment strategy. The trade-off is exactly what the question asks about: for the entire rollout window, old and new code versions are both live and both receiving real production traffic, and any request might land on either version depending on which instance the load balancer picks. ## The key mitigation The key mitigation is API and data contract compatibility across exactly one version step: the new version's request/response schema, message formats, and any shared-database schema changes must be readable and writable by the old version too, for the duration of the rollout. This is the same **expand-contract** discipline used in blue-green and canary, but here it's non-negotiable rather than a rollback-safety nicety, because mixed-version traffic is guaranteed on every rolling deploy, not an edge case. Teams enforce this by: - **never removing** a field or endpoint in the same release that stops using it - running **consumer-driven contract tests** between service pairs - treating database migrations as **strictly additive** within a single rolling deploy ## Where it breaks **Failure modes:** - A broken contract manifests as intermittent, hard-to-reproduce errors that correlate with which instance served the request — e.g., roughly one in four requests failing during a rollout where one in four instances is on the new version, a strong diagnostic signal. - Client-side connection pooling or sticky sessions can mask the problem in testing (a client pinned to one instance never sees the mismatch) while it's very real in aggregate production traffic. - A stuck rollout, where new pods repeatedly fail readiness checks, can leave a fleet in a half-and-half state for a long time unless the orchestrator is configured to auto-pause or auto-rollback on a deadline. - Unlike blue-green, rollback from a rolling deployment is itself another rolling deployment back to the old image — not instant, and it re-enters a mixed-version window, doubling exposure time to compatibility risk. ## Where it shows up A concrete real-world instance: **Kubernetes' default RollingUpdate strategy** for Deployments is exactly this mechanism, tunable via spec.strategy.rollingUpdate.maxSurge and maxUnavailable, and it's the strategy most teams default to for internal, non-customer-facing, or lower-risk services precisely because it needs no extra tooling beyond what the orchestrator already provides.
- How is a rolling deployment's guarantee about mixed versions different from canary's?Canary deliberately controls and limits the mixed-version window as an explicit, monitored percentage that you can hold or roll back at will. Rolling deployment doesn't give you that control - mixed versions happen automatically and briefly as an unavoidable side effect of the batch-by-batch replacement, with no dedicated comparison or gating step.
- What Kubernetes settings control how aggressive a rolling update is, and what's the trade-off between them?maxSurge controls how many extra pods beyond the desired count can be created temporarily, and maxUnavailable controls how many pods can be offline at once. A higher maxSurge speeds up the rollout and reduces capacity loss but costs more resources temporarily; a lower maxUnavailable protects capacity/availability during the rollout but can slow it down since fewer old pods can be retired per step.
- If a rolling deployment needs to be rolled back, is that faster or slower than a blue-green rollback?Slower - a rolling rollback is itself another rolling deployment, batch by batch, back to the previous image, so it takes real time and passes through another mixed-version window. Blue-green rollback is close to instant because the old environment is still fully running and just needs a routing flip.
Like resurfacing a highway one lane at a time while keeping traffic flowing - drivers are simultaneously on old and new pavement, so the two surfaces better line up at the seams.
saying these in an interview costs you the question
- claims rolling deployment has zero risk since it's 'zero downtime'
- doesn't recognize that old and new versions run simultaneously
- thinks rollback is instant like blue-green
- no mention of readiness/health checks gating batches
- ignores backward-compatibility requirement on shared DB/API contracts