skip to content

In GitHub Actions, why is a final cleanup step skipped when an earlier step fails?

level: seniorimportance: should knowfreq 62%

answer

  1. no condition is still a condition
  2. success is assumed, not neutrality
  3. four status check functions exist
  4. one of them ignores cancellation too
  5. cancellation is not failure

basics

~20 s

Every GitHub Actions step carries an implicit condition that all previous steps succeeded, so one failure skips the rest of the job. A cleanup step must opt out with an explicit if:, such as if: always() or if: !cancelled().

solid answer

~50 s

A step with no `if:` in GitHub Actions is not unconditional — it behaves as `if: success()`, meaning it runs only while no previous step in the job has failed. The first failing step therefore skips every step after it, including the one that tears down a test container, revokes a temporary credential, or uploads the logs you now need. The fix is an explicit condition using the status check functions: `if: always()` runs the step in every case including cancellation, `if: !cancelled()` runs after success or failure but lets a cancelled run stop promptly, and `if: failure()` runs only on the failure path — the right choice for a diagnostics dump. GitHub recommends `!cancelled()` over `always()` for anything slow, because `always()` keeps work going after someone hits cancel. The same rules apply at job level with `needs:`.

code

yaml · 20 lines
yaml
steps:
  - uses: actions/checkout@v4
  - name: Start dependencies
    run: docker compose up -d
  - name: Integration tests
    id: it
    run: ./gradlew integrationTest
  - name: Dump logs on failure
    if: failure()
    run: docker compose logs --no-color > compose.log
  - name: Upload diagnostics
    if: always()
    uses: actions/upload-artifact@v4
    with:
      name: diagnostics
      path: compose.log
      if-no-files-found: ignore
  - name: Tear down
    if: '!cancelled()'
    run: docker compose down -v

go deeper

for a junior

Recall that a GitHub Actions step with no if: runs only when the previous steps succeeded, and that if: always() is what makes a step run anyway. Recognise the leaked test container as the symptom.

for a middle

Name the four status check functions — success(), failure(), cancelled(), always() — and explain what each covers. Be able to say why success() || failure() is not the same as always().

for a senior

Reason about cancellation as a distinct outcome from failure, justify !cancelled() over always() for slow teardown, and separate this from continue-on-error and from steps.<id>.outcome versus conclusion. Extend the same reasoning to job-level needs:.

for a principal

Own the convention across the organisation: which resources must have guaranteed release, whether teardown belongs in a composite action with its own post step rather than a copy-pasted final step, and how self-hosted runner hygiene depends on it.

## The implicit condition The single sentence that explains the whole bug: **a GitHub Actions step without an `if:` key runs only if every previous step in the job succeeded.** The absence of a condition is itself a condition, functionally `if: success()`. So a job written in the obvious order — ``` - run: docker compose up -d - run: ./gradlew integrationTest - run: docker compose down ``` — tears down the environment exactly when it did not need tearing down, and leaks it every time the tests fail. On a hosted runner that leak is mostly harmless because the machine is discarded, but on a self-hosted runner it poisons the next job, and the same shape applies to revoking a lease, releasing a lock, closing a maintenance window, or uploading the log file that would have told you why the tests failed. ## The status check functions GitHub provides four functions usable in an `if:`: - `success()` — true when no previous step (or, at job level, no dependency) has failed. The implicit default. - `failure()` — true when a previous step failed. Use for notify-on-broken-build and diagnostics dumps. - `cancelled()` — true when the run has been cancelled. - `always()` — true unconditionally, including during cancellation. They compose with normal boolean operators, and `if:` accepts an expression with or without the `${{ }}` wrapper. The three idioms worth memorising: | Intent | Condition | |---|---| | Always tear down | `if: always()` | | Tear down, but let cancel win | `if: !cancelled()` | | Only on the failure path | `if: failure()` | ## `always()` versus `!cancelled()` `always()` is the blunt instrument, and GitHub's own documentation advises against it for anything long-running: a step marked `always()` keeps executing after a cancellation is requested, so a cancelled run does not actually stop when the operator expects it to. `!cancelled()` gives you the same success-and-failure coverage while letting cancellation take effect. Reach for `always()` only when the cleanup genuinely must happen even on cancel — releasing a shared lock is the classic legitimate case. ## Combining with a condition of your own A subtlety that trips people: as soon as you write *any* `if:`, the implicit success gate is gone. So ``` if: github.ref == 'refs/heads/main' ``` on a deploy step does **not** silently keep "...and everything before succeeded" — except that at step level the status functions are still implied unless you use one. The safe habit is to be explicit about both halves when you mean both: ``` if: always() && github.event_name == 'push' ``` and to read any condition you write as the complete truth about when the step runs. ## The same problem one level up Jobs behave identically. A job with `needs: [test]` is skipped when `test` fails, so a `cleanup` or `notify` job needs `if: always()` (or `!cancelled()`) to run at all. Once it does run, it can inspect `needs.test.result` to decide what to say. The pattern: ``` notify: needs: [build, test] if: failure() runs-on: ubuntu-latest steps: - run: ./notify.sh "build=${{ needs.build.result }} test=${{ needs.test.result }}" ``` ## Neighbouring mechanics people confuse with this **`continue-on-error: true`** on a step means the step's failure does not fail the job — so later steps still run *and the job can still be green*. That is a different tool: it hides a failure rather than guaranteeing cleanup. When you use it, `steps.<id>.outcome` holds what really happened (`failure`) while `steps.<id>.conclusion` holds what the job was told (`success`), and inspecting `outcome` is how you branch on the real result. **Post steps.** Some actions register their own cleanup that runs at the end of the job regardless — `actions/checkout` removing its credentials is the familiar example. That is the action's own guarantee, not something your `run:` step inherits, so it does not save your teardown script. **Timeouts and cancellation.** A job killed by `timeout-minutes` or by a cancellation is not a step failure, which is precisely why `always()` and `!cancelled()` differ, and why a teardown that matters under timeout needs `always()`. ## The interview answer in one breath Steps default to "only if everything before me passed"; cleanup is by definition the thing that must run when something before it did not pass; therefore cleanup always carries an explicit `if:`, and which of `always()`, `!cancelled()` or `failure()` you choose is a statement about cancellation and about whether the step is diagnostics or teardown.

  • What is the difference between if: always() and if: !cancelled() in GitHub Actions?
    Both run the step after success or failure of earlier steps. `always()` additionally runs it when the run is being cancelled, so the work continues after an operator hits cancel — GitHub advises against it for long tasks for exactly that reason. `!cancelled()` stops with the cancellation. Use `always()` only when the step must happen even on cancel, such as releasing a shared lock.
  • How does continue-on-error: true differ from putting if: always() on later steps?
    `continue-on-error` changes the *verdict*: the step may fail without failing the job, so subsequent steps run and the job can still be reported green. `if: always()` changes only *whether a specific step runs*, leaving the job's failure intact. Use continue-on-error for genuinely optional work, and inspect `steps.<id>.outcome` to see what really happened versus the reported `conclusion`.
  • How do you make a notification job run only when something in the GitHub Actions workflow failed?
    Give it `needs:` on the jobs it watches and `if: failure()`. Because a job with needs: is otherwise skipped when a dependency fails, the explicit condition is what lets it run at all. Inside, read `needs.<job_id>.result` for each dependency to report which stage broke, remembering that a `skipped` result is also not a success.
  • Does a step killed by timeout-minutes count as a failure for later steps in GitHub Actions?
    A step that exceeds its `timeout-minutes` is terminated and the job fails, so subsequent steps guarded only by the implicit success gate are skipped — the same as any failure. A job-level timeout ends the job outright. Teardown that must survive either case needs `if: always()`, since `!cancelled()` will not protect you from an operator cancellation arriving at the same moment.

saying these in an interview costs you the question

  • Assumes a step with no if: runs unconditionally
  • Uses if: success() || failure() and expects cancellation coverage
  • Confuses continue-on-error with guaranteeing cleanup runs
  • Puts teardown last and calls it a finally block
  • Thinks a cancelled run counts as failure()

context