In a GitHub Actions matrix build, what does setting fail-fast: false change?
answer
- the default cancels the siblings
- cancelled is not the same as failed
- one setting buys information, the other protects a resource
- the throttle applies within this matrix only
- serialising across runs is a different key
basics
~20 sBy default a failing matrix job cancels every other in-progress and queued job in that matrix. fail-fast: false lets all legs run to completion, so one pull request shows every platform or version that is broken instead of only the first one to fail.
solid answer
~40 s`strategy.fail-fast` defaults to `true`, meaning the first matrix job that fails cancels all in-progress and pending siblings in the same matrix. That saves runner minutes but destroys information: a change that breaks Windows and Java 17 reports a single cancelled mess, and you fix one thing, push, and discover the next. Setting `fail-fast: false` lets every leg finish, so one run tells you the full blast radius — which is what you want on a matrix whose legs represent distinct platforms rather than redundant shards. The companion knob is `max-parallel`, which caps how many matrix jobs run concurrently; it is for protecting a shared resource such as a licence pool, a rate-limited API, or a small self-hosted runner fleet, not for correctness. Neither setting affects jobs outside the matrix.
code
yaml · 21 linesjobs:
test:
runs-on: ${{ matrix.os }}
continue-on-error: ${{ matrix.experimental == true }}
strategy:
fail-fast: false # report every broken platform in one run
max-parallel: 3 # shared staging DB tolerates 3 writers
matrix:
os: [ubuntu-latest, windows-latest]
java: [17, 21]
include:
- os: ubuntu-latest
java: 25
experimental: true
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: ${{ matrix.java }}
- run: ./gradlew testgo deeper
Know that a matrix runs many jobs and that by default one failing job cancels the rest, and that fail-fast: false is how you let them all finish.
Explain the trade explicitly — runner minutes versus complete feedback — and describe max-parallel as a concurrency cap within the matrix, distinct from the concurrency key.
Choose per matrix: default on redundant shards, off on compatibility matrices, max-parallel where a shared resource or a small runner fleet is the constraint, and continue-on-error for a genuinely experimental leg.
Own the policy across the organization: how much concurrency a shared runner pool can absorb, what feedback latency the team is buying, and when broad coverage moves off the pull-request path to a scheduled run.
## The default behaviour strategy: fail-fast: true # this is the default matrix: os: [ubuntu-latest, windows-latest, macos-latest] With `fail-fast` at its default, the moment any matrix job fails, GitHub cancels the remaining jobs of that matrix — both the ones already running and the ones still queued. Cancelled jobs report as cancelled, not failed, which is itself a common source of confusion when reading a run summary. The trade being made is **minutes versus information**. Fail-fast is right when the legs are largely redundant — shards of the same test suite, for example — because the first failure tells you everything the others would. It is wrong when the legs are genuinely different environments, because the one thing you want from a multi-platform matrix is to learn about every platform in one round trip. ## Turning it off strategy: fail-fast: false matrix: os: [ubuntu-latest, windows-latest, macos-latest] java: [17, 21] Now all six legs run regardless of failures. The run's conclusion is still a failure if any leg failed; you simply get the complete picture. The cost is real — a broken change burns the full matrix's minutes — so the sensible policy is usually `fail-fast: false` on a small, meaningful compatibility matrix and the default on a large shard matrix. ## max-parallel strategy: max-parallel: 2 matrix: shard: [1, 2, 3, 4, 5, 6] `max-parallel` limits how many jobs from this matrix run at the same time; the rest queue. Reasons to set it: - A shared external dependency that cannot take the concurrency — a staging database, a rate-limited third-party API, a fixed pool of test licences. - A small self-hosted runner fleet you do not want one workflow to monopolise while other repositories wait. - Cost smoothing, where you would rather trade wall-clock time for a flatter concurrency profile. Without it, GitHub runs as many matrix jobs concurrently as your available runners and account concurrency limits allow. Note the scope: `max-parallel` constrains only jobs within this matrix, in this run. Serialising *across* runs — for instance one deploy at a time per branch — is what the `concurrency` key is for, and confusing the two is a common mistake. ## Tolerating a known-flaky leg Sometimes one leg is experimental and should not fail the run. The idiom is a job-level `continue-on-error` driven by a matrix value: strategy: fail-fast: false matrix: java: [17, 21] include: - java: 25 experimental: true continue-on-error: ${{ matrix.experimental == true }} The experimental leg still runs and still reports, but its failure does not fail the workflow. Keep this rare: a permanently red-but-tolerated leg is quickly ignored by everyone. ## Related context values The `strategy` context exposes `strategy.fail-fast`, `strategy.max-parallel`, `strategy.job-index`, and `strategy.job-total`, which are occasionally useful for sharding logic or for labelling output. `job-index` in particular lets a script decide which slice of a test suite to run without duplicating the list in YAML. ## What to say "`fail-fast: false` stops one failing leg from cancelling the others, so a single run shows every environment that is broken; `max-parallel` throttles how many legs run at once to protect a shared resource. The first is about feedback quality, the second about resource contention."
- When is the default fail-fast: true the better choice?When the matrix legs are redundant rather than distinct — shards of one test suite, or repeated runs of the same configuration. The first failure already tells you what the others would, so cancelling saves minutes without losing information. On a compatibility matrix across operating systems or language versions, the opposite is true.
- How does max-parallel differ from the concurrency key in GitHub Actions?`max-parallel` caps how many jobs of one matrix run simultaneously within a single run. `concurrency` groups runs or jobs by an expression and ensures only one of a group is in flight, optionally cancelling the older one — that is what serialises deploys across runs. They solve different problems and are frequently confused.
- How do you let one experimental matrix leg fail without failing the workflow?Add the leg through `include` with a marker value such as `experimental: true`, then set the job's `continue-on-error` from that value. The leg still runs and reports, but its failure does not fail the run. Use it sparingly — a leg that is permanently allowed to be red stops being read.
saying these in an interview costs you the question
- Thinks fail-fast controls retries on failure
- Believes cancelled matrix jobs are reported as failures
- Uses max-parallel to serialise deploys across runs
- Sets fail-fast: false everywhere without cost thought
- Expects fail-fast to cancel jobs outside the matrix