skip to content

In a GitHub Actions matrix build, what does setting fail-fast: false change?

level: middleimportance: should knowfreq 54%

answer

  1. the default cancels the siblings
  2. cancelled is not the same as failed
  3. one setting buys information, the other protects a resource
  4. the throttle applies within this matrix only
  5. serialising across runs is a different key

basics

~20 s

By default a failing matrix job cancels every other in-progress and queued job in that matrix. fail-fast: false lets all legs run to completion, so one pull request shows every platform or version that is broken instead of only the first one to fail.

solid answer

~40 s

`strategy.fail-fast` defaults to `true`, meaning the first matrix job that fails cancels all in-progress and pending siblings in the same matrix. That saves runner minutes but destroys information: a change that breaks Windows and Java 17 reports a single cancelled mess, and you fix one thing, push, and discover the next. Setting `fail-fast: false` lets every leg finish, so one run tells you the full blast radius — which is what you want on a matrix whose legs represent distinct platforms rather than redundant shards. The companion knob is `max-parallel`, which caps how many matrix jobs run concurrently; it is for protecting a shared resource such as a licence pool, a rate-limited API, or a small self-hosted runner fleet, not for correctness. Neither setting affects jobs outside the matrix.

code

yaml · 21 lines
yaml
jobs:
  test:
    runs-on: ${{ matrix.os }}
    continue-on-error: ${{ matrix.experimental == true }}
    strategy:
      fail-fast: false      # report every broken platform in one run
      max-parallel: 3       # shared staging DB tolerates 3 writers
      matrix:
        os: [ubuntu-latest, windows-latest]
        java: [17, 21]
        include:
          - os: ubuntu-latest
            java: 25
            experimental: true
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with:
          distribution: temurin
          java-version: ${{ matrix.java }}
      - run: ./gradlew test

go deeper

for a junior

Know that a matrix runs many jobs and that by default one failing job cancels the rest, and that fail-fast: false is how you let them all finish.

for a middle

Explain the trade explicitly — runner minutes versus complete feedback — and describe max-parallel as a concurrency cap within the matrix, distinct from the concurrency key.

for a senior

Choose per matrix: default on redundant shards, off on compatibility matrices, max-parallel where a shared resource or a small runner fleet is the constraint, and continue-on-error for a genuinely experimental leg.

for a principal

Own the policy across the organization: how much concurrency a shared runner pool can absorb, what feedback latency the team is buying, and when broad coverage moves off the pull-request path to a scheduled run.

## The default behaviour strategy: fail-fast: true # this is the default matrix: os: [ubuntu-latest, windows-latest, macos-latest] With `fail-fast` at its default, the moment any matrix job fails, GitHub cancels the remaining jobs of that matrix — both the ones already running and the ones still queued. Cancelled jobs report as cancelled, not failed, which is itself a common source of confusion when reading a run summary. The trade being made is **minutes versus information**. Fail-fast is right when the legs are largely redundant — shards of the same test suite, for example — because the first failure tells you everything the others would. It is wrong when the legs are genuinely different environments, because the one thing you want from a multi-platform matrix is to learn about every platform in one round trip. ## Turning it off strategy: fail-fast: false matrix: os: [ubuntu-latest, windows-latest, macos-latest] java: [17, 21] Now all six legs run regardless of failures. The run's conclusion is still a failure if any leg failed; you simply get the complete picture. The cost is real — a broken change burns the full matrix's minutes — so the sensible policy is usually `fail-fast: false` on a small, meaningful compatibility matrix and the default on a large shard matrix. ## max-parallel strategy: max-parallel: 2 matrix: shard: [1, 2, 3, 4, 5, 6] `max-parallel` limits how many jobs from this matrix run at the same time; the rest queue. Reasons to set it: - A shared external dependency that cannot take the concurrency — a staging database, a rate-limited third-party API, a fixed pool of test licences. - A small self-hosted runner fleet you do not want one workflow to monopolise while other repositories wait. - Cost smoothing, where you would rather trade wall-clock time for a flatter concurrency profile. Without it, GitHub runs as many matrix jobs concurrently as your available runners and account concurrency limits allow. Note the scope: `max-parallel` constrains only jobs within this matrix, in this run. Serialising *across* runs — for instance one deploy at a time per branch — is what the `concurrency` key is for, and confusing the two is a common mistake. ## Tolerating a known-flaky leg Sometimes one leg is experimental and should not fail the run. The idiom is a job-level `continue-on-error` driven by a matrix value: strategy: fail-fast: false matrix: java: [17, 21] include: - java: 25 experimental: true continue-on-error: ${{ matrix.experimental == true }} The experimental leg still runs and still reports, but its failure does not fail the workflow. Keep this rare: a permanently red-but-tolerated leg is quickly ignored by everyone. ## Related context values The `strategy` context exposes `strategy.fail-fast`, `strategy.max-parallel`, `strategy.job-index`, and `strategy.job-total`, which are occasionally useful for sharding logic or for labelling output. `job-index` in particular lets a script decide which slice of a test suite to run without duplicating the list in YAML. ## What to say "`fail-fast: false` stops one failing leg from cancelling the others, so a single run shows every environment that is broken; `max-parallel` throttles how many legs run at once to protect a shared resource. The first is about feedback quality, the second about resource contention."

  • When is the default fail-fast: true the better choice?
    When the matrix legs are redundant rather than distinct — shards of one test suite, or repeated runs of the same configuration. The first failure already tells you what the others would, so cancelling saves minutes without losing information. On a compatibility matrix across operating systems or language versions, the opposite is true.
  • How does max-parallel differ from the concurrency key in GitHub Actions?
    `max-parallel` caps how many jobs of one matrix run simultaneously within a single run. `concurrency` groups runs or jobs by an expression and ensures only one of a group is in flight, optionally cancelling the older one — that is what serialises deploys across runs. They solve different problems and are frequently confused.
  • How do you let one experimental matrix leg fail without failing the workflow?
    Add the leg through `include` with a marker value such as `experimental: true`, then set the job's `continue-on-error` from that value. The leg still runs and reports, but its failure does not fail the run. Use it sparingly — a leg that is permanently allowed to be red stops being read.

saying these in an interview costs you the question

  • Thinks fail-fast controls retries on failure
  • Believes cancelled matrix jobs are reported as failures
  • Uses max-parallel to serialise deploys across runs
  • Sets fail-fast: false everywhere without cost thought
  • Expects fail-fast to cancel jobs outside the matrix

context