skip to content

'Progressive delivery' combines canary-style traffic control, feature flags, and automated analysis into one release process. As a principal engineer choosing a deployment strategy per service, when would you deliberately NOT use canary or progressive delivery, even though the tooling is available?

level: principalimportance: nice to knowfreq 45%

answer

  1. signal needs traffic volume
  2. irreversible side effects aren't proportional to %
  3. batch jobs have no traffic to split
  4. cost of tooling vs actual blast radius
  5. tier strategy by service risk

basics

~20 s

Sometimes a fancy gradual rollout isn't worth it - for tiny low-traffic services, batch jobs, or changes where a database migration already commits you either way, a simpler rolling deploy with good tests is safer and cheaper than building a whole canary pipeline.

solid answer

~40 s

Progressive delivery is the right default for high-traffic, customer-facing, independently-deployable services where traffic-splitting and metrics-comparison machinery has enough signal to work and the tooling/observation-window cost is worth it. It's a poor fit for: low-QPS services where a canary slice can't reach statistical significance in a reasonable window; changes gated by an irreversible side effect (a destructive migration, a one-time backfill, a non-idempotent external API call) where traffic percentage is irrelevant because damage isn't proportional to exposure; batch/offline/cron workloads with no live traffic to split; and small teams/services where maintaining canary analysis templates and dashboards costs more than the actual risk being mitigated, in which case a simple rolling deploy plus solid automated tests and fast manual rollback is the pragmatic choice.

go deeper

for a junior

Aware that not every deploy needs the fanciest strategy available; can give a basic reason like 'it's more work.'

for a middle

Can name at least one concrete case (e.g., low traffic) where canary doesn't work well.

for a senior

Explains why irreversible side effects break canary's proportional-risk assumption, and why batch workloads don't fit traffic-splitting techniques at all.

for a principal

Designs a per-service tiering policy balancing blast radius, traffic volume, reversibility, and organizational tooling cost, rather than mandating one strategy universally.

## What progressive delivery combines Progressive delivery is the umbrella term for combining several previously separate techniques into one continuous release pipeline, rather than treating 'deploy the code,' 'shift traffic,' and 'turn the feature on' as disconnected manual steps: - canary-style percentage-based traffic shifting - feature flags for fine-grained release targeting - automated metric analysis for promotion/rollback decisions Tools like **Argo Rollouts** or **Flagger** implement this by managing both the traffic split and, optionally, flag state, while running an analysis step against a metrics backend at each stage and automatically promoting, holding, or aborting — wiring the canary and automated-rollback mechanisms together into one declarative pipeline rather than several loosely coupled tools operated by hand. ## Why the deliberate no is the harder call A principal engineer needs a considered opinion on when NOT to use this, rather than defaulting to 'always use the most sophisticated available strategy,' because progressive delivery has real fixed costs — building and maintaining analysis templates, instrumenting per-version metrics, tuning thresholds, running the extra service-mesh/ingress infrastructure a traffic split requires — and those costs are worth paying only when the technique's core assumption holds: that exposing a fraction of traffic and measuring the result produces a useful, statistically meaningful signal proportional to risk. Several common situations break that assumption. ## Where the assumption breaks 1. **First, low-traffic services.** A service handling 20 requests per minute gets roughly one request per minute at a 5% canary stage, nowhere near enough volume to distinguish a real regression from ordinary variance within a reasonable observation window — you'd need to hold each stage impractically long, eroding the speed benefit progressive delivery is supposed to provide. For such services, a straightforward rolling deployment backed by strong pre-production tests and a fast, well-rehearsed manual rollback is often genuinely safer than a canary pipeline whose statistical power is too weak to trust. 2. **Second, changes gated by an irreversible or non-proportional side effect.** A destructive database migration, a one-time backfill, or a call to a non-idempotent external API (charging a customer, sending an email, firing a partner webhook) doesn't get safer by exposing it to 5% of traffic instead of 100% — the damage from a single bad execution isn't proportional to how many requests hit the new code, unlike a stateless bug whose blast radius genuinely scales with exposure percentage. For these, the right tool is different gating entirely: a manual approval step, a dry-run/shadow-mode execution that doesn't commit real side effects, or thorough pre-production testing plus structural reversibility (expand-contract, idempotency keys), not a traffic-percentage dial. 3. **Third, batch, offline, or cron-triggered workloads.** These have no live request traffic to split — percentage-based traffic shifting is a category error for a nightly ETL job; the equivalent safety mechanism is closer to running the new version against a shadow dataset and diffing output, or a blue-green-style swap of which job definition is active. 4. **Fourth, organizational cost versus actual risk** — and this is the genuinely principal-level judgment call. A small team maintaining internal, low-blast-radius services can spend disproportionate ongoing effort building and keeping healthy a full progressive-delivery pipeline — dashboards, analysis templates, on-call runbooks for when the automation itself misbehaves — for services where the realistic worst case of a bad deploy is a handful of internal users seeing an error for ten minutes until someone reverts. In that setting, the standardization and tooling burden isn't justified by the risk being mitigated, and a principal engineer's job is to right-size deployment strategy per service based on blast radius, traffic volume, and reversibility, not to mandate the most sophisticated tool uniformly across a whole fleet. ## Where it shows up A well-known real-world articulation of this: large platform teams running **Argo Rollouts** or **Spinnaker** at scale typically tier services — critical, customer-facing, high-traffic services get full canary/progressive-delivery treatment with automated analysis, while internal tools and low-traffic services are explicitly allowed plain rolling deployments, because applying the heavy machinery everywhere would be wasted engineering effort without a proportional safety return.

  • Why doesn't canary's percentage-based exposure help when a release includes a non-idempotent external API call, like charging a customer's card?
    Because the harm from a single bad execution (double-charging one customer) isn't reduced by only sending 5% of traffic through the new code - that 5% still fully executes the flawed charge for whoever hits it, so the safety benefit of canary (proportional blast radius) doesn't apply. The right control is a dry-run/shadow mode or idempotency key, not a traffic percentage.
  • What's the statistical problem with canarying a service that only receives 20 requests per minute?
    At a 5% canary stage that's roughly one request per minute going to the new version, which is far too small a sample to reliably distinguish a genuine regression from ordinary random variance within a reasonable observation window. You'd need to either hold the canary stage for an impractically long time or accept a much weaker confidence in the result.
  • How would you decide, as a principal engineer, which of a company's 50 microservices get full progressive-delivery treatment versus plain rolling deploys?
    Tier services by blast radius (customer-facing/revenue-critical vs. internal), traffic volume (enough for statistical signal), and reversibility of their typical changes (stateless vs. involving irreversible side effects), then apply progressive delivery where the combination of risk and adequate signal justifies the tooling cost, and default lower-risk or low-signal services to simpler rolling deploys plus strong tests.

Like not bothering with a full clinical-trial-style staged rollout for a one-time, irreversible surgery - percentage exposure only helps when the risk genuinely scales with how many people you expose.

saying these in an interview costs you the question

  • insists every service should always use canary/progressive delivery regardless of traffic or risk
  • doesn't recognize that irreversible side effects break the proportional-blast-radius assumption
  • no awareness that low traffic breaks canary's statistical validity
  • ignores the ongoing tooling/maintenance cost of progressive delivery pipelines
  • treats 'more sophisticated deployment strategy' as strictly better with no cost side

context