skip to content

In a Jenkins Pipeline, what exactly does wrapping steps in retry(3) re-run, and why is that dangerous around a deployment step?

level: middleimportance: nice to knowfreq 32%

answer

  1. the whole closure, from the top
  2. the count includes the first attempt
  3. no delay and no cleanup between attempts
  4. workspace state carries over unchanged
  5. idempotency is the real question

basics

~20 s

retry(3) re-executes the entire enclosed block from the top when it fails, for up to three attempts in total. Nothing is undone between attempts, so any non-idempotent work inside — a partial upload, an applied migration — is simply done again.

solid answer

~50 s

`retry(n)` takes a block and runs it up to `n` times *in total* — the first attempt counts, so `retry(3)` means the body runs at most three times, not four. It re-runs the whole block from the first step, not from the step that failed, and there is no built-in delay between attempts, so add a `sleep` if you want backoff. It does not reset the workspace, roll anything back, or distinguish an infrastructure hiccup from a real defect: a genuine assertion failure is retried exactly like a dropped TCP connection, which is how a broken build gets three chances to look flaky. Around deployment that matters more, because the body has side effects outside the workspace. Wrap the smallest possible block, make it idempotent, and prefer retrying the network call rather than the whole deploy.

code

groovy · 19 lines
groovy
stage('Deploy') {
  steps {
    // Narrow retry: only the fetch is retried, with backoff
    retry(3) {
      script {
        try {
          sh 'curl --fail -sSo release.tgz https://artifacts.example.com/release.tgz'
        } catch (err) {
          sleep time: 15, unit: 'SECONDS'
          throw err
        }
      }
    }
    // Not retried: side effects outside the workspace
    timeout(time: 10, unit: 'MINUTES') {
      sh './deploy.sh release.tgz'
    }
  }
}

go deeper

for a junior

Know that retry(n) re-runs the whole enclosed block on failure up to n attempts in total, and that timeout is the separate wrapper for a step that hangs.

for a middle

Explain the mechanics: no resume mid-block, no delay, no state reset, and no ability to tell a flake from a real failure — then say what that implies about where you place the wrapper.

for a senior

Show the production judgment: retry the narrow transport-level operation, require idempotency before retrying anything with external side effects, and treat a retried test suite as hidden signal rather than a fix.

for a principal

Own the policy question — where retries are permitted at all, how flake rates are measured rather than papered over, and how retry budgets interact with overall pipeline latency and executor capacity.

## What the step does `retry` is a Pipeline wrapper step: you give it a count and a block, and it runs the block, catching failures and running it again. ```groovy retry(3) { sh 'curl --fail -o artifact.tgz https://artifacts.example.com/build-42.tgz' } ``` Two details trip people up immediately. First, the count is the number of *attempts*, not the number of *retries*: `retry(3)` runs the body a maximum of three times. Second, the unit of retry is the whole block. If the block contains five steps and the fourth fails, all five run again from the top — Jenkins has no notion of resuming mid-block. Declarative also exposes it as a stage option, where the unit is the entire stage body: ```groovy stage('Flaky UI tests') { options { retry(3) } steps { sh './run-ui-tests.sh' } } ``` ## What it does not do It does not wait between attempts. There is no backoff parameter; if you need one, put a `sleep` inside the block before the operation, or accept that three attempts against a service that is down will all fail within seconds and buy you nothing. It does not reset state. The workspace is exactly as the failed attempt left it: half-written files, a populated cache, a partially extracted archive. Some tools cope; a build that fails on a corrupt partial download and then retries against the same corrupt file will fail three times for the same reason. It does not classify failures. A connection reset and a failing unit test are both "the body threw", and both are retried. This is the quality problem: wrapping a test stage in `retry` converts a reproducible failure into an intermittent one, hides a real regression behind a green build, and removes the pressure to fix the underlying flake. If you retry tests at all, retry them narrowly and record how often the retry fired, so the flakiness stays visible. ## Why deployment is the dangerous case Inside a test stage, a repeated attempt mostly costs time. Inside a deploy stage, the block's side effects are outside the workspace and outside Jenkins entirely. Consider a deploy block that pushes an image, applies a database migration, and flips a traffic weight, and that fails on the last step. The retry re-pushes the image (harmless), re-applies the migration (possibly not harmless at all), and flips the weight again. Nothing rolled the migration back, because `retry` has no rollback concept — it only knows how to call your block again. The discipline is therefore: - **Wrap the smallest thing.** Retry the flaky HTTP call, not the seven-step deploy. - **Make the body idempotent.** Applying the same declarative manifest twice is safe; appending a row twice is not. - **Retry the transport, not the decision.** Network calls, registry pulls and agent provisioning are legitimately retryable; "did this migration succeed" is not. ## Pairing with timeout `retry` handles a step that *fails*; it does nothing for a step that *hangs*, because a hung step never throws. The companion wrapper is `timeout`: ```groovy timeout(time: 10, unit: 'MINUTES') { sh './publish.sh' } ``` When the budget expires, Jenkins interrupts the body and the run ends up aborted rather than hanging until someone notices — and `post { aborted { ... } }` gives you a place to react. Every long-running or network-facing block in a pipeline deserves a timeout, whether or not it deserves a retry; a job with no timeout can hold an executor indefinitely. ## What a strong answer sounds like "`retry(3)` re-runs the whole block up to three times total, with no delay and no cleanup between attempts, and it cannot tell a flake from a real failure. That is fine around a download and dangerous around a deploy, because the deploy's side effects are external and are simply repeated. I wrap the narrowest retryable operation, make it idempotent, and use `timeout` for the separate problem of steps that hang."

  • How many times does the body of retry(3) run if it fails every time?
    Three times in total. The count is attempts, not retries after the first failure, so `retry(3)` gives you the original attempt plus two more. After the third failure the exception propagates and the stage fails normally.
  • retry handles failures — what handles a step that hangs instead of failing?
    `timeout`. A hung step never throws, so retry never engages; `timeout(time: 10, unit: 'MINUTES') { ... }` bounds the block and interrupts it when the budget expires, ending the run as aborted rather than holding an executor indefinitely. Long-running and network-facing blocks want a timeout whether or not they want a retry.
  • Your team wraps the whole test stage in retry(2) to stop flaky failures. What is wrong with that?
    It cannot distinguish a flake from a regression, so a genuinely broken change gets a second chance to pass and can ship. It also doubles the worst-case stage time and removes any signal about how flaky the suite really is. Retry the specific unreliable operation, record when retries fire, and treat repeat offenders as bugs.

saying these in an interview costs you the question

  • Thinking retry(3) means one attempt plus three retries
  • Expecting retry to resume from the failed step
  • Assuming the workspace is cleaned between attempts
  • Believing retry backs off automatically between attempts
  • Using retry to make a genuinely failing test suite pass

context