skip to content

In Argo Rollouts, how does an AnalysisTemplate attached to a canary step decide whether to promote or abort, and what does Argo CD show for a Rollout that is paused mid-canary?

level: seniorimportance: should knowfreq 38%

answer

  1. a separate controller, not Argo CD
  2. steps: weight, pause, analysis
  3. measurements graded against a condition
  4. too many failures aborts to stable
  5. paused is Suspended, not Healthy

basics

~20 s

An AnalysisRun queries a metrics provider on an interval and evaluates each metric's successCondition; exceeding failureLimit aborts the Rollout, shifting traffic back to the stable ReplicaSet. Argo CD's bundled health check reports a paused Rollout as Suspended, and an aborted one as Degraded.

solid answer

~50 s

Argo Rollouts replaces the Deployment with a `Rollout` resource whose `strategy.canary.steps` list interleaves `setWeight`, `pause` and `analysis` steps. When an `analysis` step runs, the controller creates an `AnalysisRun` from the referenced `AnalysisTemplate` (or `ClusterAnalysisTemplate`). Each metric in it has a provider — Prometheus, Datadog, a Job, a plain web request — an `interval` and `count`, and a `successCondition` or `failureCondition` evaluated against `result`. Measurements that fail accumulate against `failureLimit`; once exceeded, the AnalysisRun fails and the Rollout is aborted, meaning traffic weight returns to the stable ReplicaSet and the canary is scaled down. If measurements keep passing to the end of the run, the step completes and the Rollout proceeds to the next step. On the Argo CD side, the bundled health check for `Rollout` maps a rollout sitting on an indefinite `pause: {}` to `Suspended` and an aborted rollout to `Degraded`, so the Application's health reflects progressive delivery rather than showing a false green.

code

yaml · 30 lines
yaml
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: checkout
spec:
  replicas: 10
  selector:
    matchLabels:
      app: checkout
  template:
    metadata:
      labels:
        app: checkout
    spec:
      containers:
        - name: checkout
          image: registry.example.com/checkout:1.4.2
  strategy:
    canary:
      steps:
        - setWeight: 20
        - pause: { duration: 5m }
        - analysis:
            templates:
              - templateName: success-rate
            args:
              - name: service-name
                value: checkout
        - setWeight: 50
        - pause: {}

go deeper

for a junior

Know that Argo Rollouts is a separate controller providing canary and blue-green behaviour, and that Argo CD syncs its resources like any other manifest.

for a middle

Explain the step list — setWeight, pause, analysis — and that an analysis step creates an AnalysisRun which queries a metrics provider and grades each measurement against a condition.

for a senior

Bring the outcomes and their limits: exceeding failureLimit aborts back to the stable ReplicaSet, no-data measurements are inconclusive rather than passing, and Argo CD reports a paused rollout as Suspended so pipelines do not sail past the gate.

for a principal

Own the separation of mechanism from policy: the controller can abort automatically, but which signal, what threshold and how long to bake are reliability decisions, and a canary too small or too short to gather evidence gives you the ceremony of safety without the substance.

## Where Argo Rollouts sits Argo CD applies manifests and assesses health; it does not shift traffic. Argo Rollouts is a separate controller that introduces a `Rollout` resource — a drop-in replacement for a Deployment's pod template that manages ReplicaSets itself and adds step-driven canary and blue-green behaviour. Argo CD's part in this pairing is twofold: it syncs the `Rollout` and its `AnalysisTemplate`s from Git like any other manifest, and it reads the Rollout's status through a bundled health check so the Application's health reflects what the rollout controller is doing. ## The canary step list ```yaml apiVersion: argoproj.io/v1alpha1 kind: Rollout metadata: name: checkout spec: replicas: 10 strategy: canary: steps: - setWeight: 20 - pause: { duration: 5m } - analysis: templates: - templateName: success-rate args: - name: service-name value: checkout - setWeight: 50 - pause: {} ``` The controller walks this list one step at a time. `setWeight` moves the share of traffic sent to the canary ReplicaSet. `pause` with a `duration` waits and continues; `pause: {}` with no duration waits **indefinitely** for a human to promote. `analysis` starts a background evaluation while the rollout continues (or blocks, depending on placement). ## What an AnalysisRun actually does An `AnalysisTemplate` is a reusable definition of "what good looks like"; when a step references it, the controller instantiates an `AnalysisRun`. Each metric inside carries: - a **provider** — Prometheus, Datadog, New Relic, CloudWatch, a `job` that runs a Kubernetes Job, or `web` for an arbitrary HTTP call; - `interval` and `count` — how often to measure and how many measurements to take; - `successCondition` and/or `failureCondition` — expressions evaluated against `result`, the value the provider returned; - `failureLimit` (and `inconclusiveLimit`) — how many failed measurements are tolerated before the run itself fails. ```yaml apiVersion: argoproj.io/v1alpha1 kind: AnalysisTemplate metadata: name: success-rate spec: args: - name: service-name metrics: - name: success-rate interval: 1m count: 5 successCondition: result[0] >= 0.99 failureLimit: 1 provider: prometheus: address: http://prometheus.monitoring.svc.cluster.local:9090 query: >- sum(rate(http_requests_total{service="{{args.service-name}}",status!~"5.."}[2m])) / sum(rate(http_requests_total{service="{{args.service-name}}"}[2m])) ``` Each measurement is graded independently. Passing measurements advance the count; failing ones accrue against `failureLimit`. Exceed it and the AnalysisRun is `Failed`, which **aborts** the Rollout: weight returns to the stable ReplicaSet, the canary scales down, and the Rollout reports as degraded. Complete all `count` measurements without exceeding the limit and the run is `Successful`, letting the step list continue. A third outcome matters in practice: `Inconclusive`. If the provider returns nothing — a query with no data because no traffic reached the canary — the measurement is inconclusive rather than successful, and depending on `inconclusiveLimit` the rollout pauses for a human instead of silently promoting on absent evidence. ## Blue-green analysis The blue-green strategy expresses the same idea with different fields: `activeService` and `previewService`, `autoPromotionEnabled`, and `prePromotionAnalysis` / `postPromotionAnalysis`. Pre-promotion analysis runs against the preview service *before* the active service is cut over, so a failure means the switch never happens; post-promotion analysis runs after, with `scaleDownDelaySeconds` keeping the old ReplicaSet around long enough to switch back. ## How Argo CD reads all this Argo CD bundles a health check for `Rollout`, so the Application status is not blind to the rollout state: - rollout still progressing through steps → `Progressing`; - sitting on an indefinite `pause: {}` awaiting manual promotion → `Suspended`; - aborted or otherwise failed → `Degraded`; - fully promoted and available → `Healthy`. `Suspended` is the one candidates get wrong. A rollout waiting for a human is not a failure and not a success, and if Argo CD reported it Healthy, an automated pipeline waiting on Application health would sail past a gate that exists precisely to stop it. Note that the Application can be `Synced` throughout — Git declared the Rollout, the Rollout exists, the diff is clean — while health carries the entire progressive-delivery story. ## The judgment interviewers want That analysis is only as good as the metric behind it: a success-rate query with a two-minute window and a five-minute canary can promote on almost no data, and a canary receiving 5% of traffic may never accumulate enough events for a meaningful signal. State the mechanism, then state that limitation — and keep the *policy* question (what SLI, what threshold, how long a bake) separate from the mechanics, because that is a reliability-engineering decision rather than a rollout-controller feature.

  • What does aborting a Rollout actually do to running traffic?
    The controller returns the traffic weight to the stable ReplicaSet and scales the canary down, so the previously-running version serves everything again. The Rollout reports as degraded and stays on the old ReplicaSet until someone retries or a new revision is pushed. Nothing is reverted in Git — the Rollout manifest still names the new image, which is why the Application can be Synced and Degraded at once.
  • What happens when an analysis metric returns no data at all?
    The measurement is Inconclusive rather than successful, and inconclusive measurements accrue against inconclusiveLimit. That distinction exists so a canary with no traffic cannot be promoted on absent evidence: instead of passing, the rollout pauses for a human decision. It is the reason a low canary weight plus a short bake often produces stalled rollouts rather than confident ones.
  • How does blue-green pre-promotion analysis differ from a canary analysis step?
    Pre-promotion analysis runs against the previewService while the activeService still points entirely at the old ReplicaSet, so a failure means the cutover never happens and no user traffic touched the new version. Canary analysis grades a version that is already serving a slice of real traffic. Pre-promotion is safer but tests with synthetic load; canary tests with reality at limited blast radius.

saying these in an interview costs you the question

  • Thinks Argo CD itself shifts traffic between versions
  • Says a paused Rollout shows as Healthy in Argo CD
  • Believes a failed analysis reverts the Git revision too
  • Assumes a query returning no data counts as success
  • Confuses Rollout steps with Deployment rolling-update settings

context