skip to content

Your GitHub Actions matrix has grown to 60 jobs per pull request. How do you decide what to keep?

level: principalimportance: nice to knowfreq 34%

answer

  1. treat it as a portfolio, not a checklist
  2. which legs failed alone?
  3. move coverage rather than deleting it
  4. the biggest saving is not running the leg at all
  5. renaming a leg can break a required check

basics

~20 s

Rank each dimension by the failures it has actually caught, keep on every pull request only the legs that catch defects a merge would ship, and move the rest to a scheduled or pre-release run. Then attack the per-leg cost with caching and change detection.

solid answer

~50 s

I treat the matrix as a portfolio and ask what each dimension has actually caught. Pull the failure history and classify: legs that catch real, unique defects stay on the pull-request path; legs that have never failed independently of another leg are redundant and move to a nightly or pre-release run; legs that fail only flakily are a reliability bug to fix or delete, because a leg nobody trusts is worse than no leg. Then I reduce the cost of what remains — dependency caching with a properly fingerprinted key, a dynamic matrix that builds only the packages a pull request touched, and shorter smoke coverage on the pull request with full coverage on merge. Finally I set the policy: `fail-fast: false` on the small compatibility matrix so one round trip shows everything, `max-parallel` only where a shared resource forces it, and a stated target for pull-request feedback latency that the matrix has to fit inside.

go deeper

for a junior

Understand that every matrix leg is a full job with its own setup and cost, so adding a dimension multiplies the work rather than adding a little to it.

for a middle

Explain the levers you would reach for: caching with a good key, dropping redundant combinations, and running the broad matrix on a schedule instead of on every pull request.

for a senior

Bring evidence: failure history per leg, unique yield, feedback latency, re-run rate — then tier the coverage and cut per-leg cost with change detection and verified cache hits.

for a principal

Own the policy and its consequences: a stated latency budget, rules for adding a dimension, flaky-leg deadlines, shared runner capacity across teams, and coordination with required-check policy so a matrix change never becomes a merge outage.

## Frame it as value per leg, not job count Sixty jobs is not automatically wrong; sixty jobs where fifty-five have never independently caught a defect is. The question a lead has to answer is what marginal risk each dimension removes and what it costs in feedback latency, runner spend, and queue contention with every other repository sharing the capacity. ## Step 1: measure what the matrix actually catches Use the run history — the Actions API exposes runs, jobs, and conclusions — to build a simple table over a meaningful window: for each matrix leg, how often it failed, and how often it failed **when no other leg failed**. That second number is the leg's unique yield, and it is usually brutally revealing. Legs with high unique yield are earning their place. Legs that only ever fail alongside the Linux job are duplicates of it. Legs that fail intermittently and pass on re-run are flaky and are costing trust as well as minutes. Also measure the human cost: the median wall-clock time from push to a complete verdict, and the re-run rate. If the matrix takes 40 minutes and a third of runs are re-run, the real cost is engineers context-switching, which dwarfs the compute bill. ## Step 2: tier the coverage Most matrices collapse cleanly into three tiers: - **Pull request** — the smallest set that catches the defects a merge would ship: usually the primary platform, the oldest and newest supported toolchain, and anything with a genuine history of unique failures. - **Merge or pre-release** — broader coverage, run once per merge to the default branch, where a slower result is acceptable because nobody is waiting on it to press a button. - **Scheduled** — the exhaustive sweep, run nightly or weekly, catching environment drift and long-tail combinations. Failures here need an owner and a triage path, or the schedule becomes noise. The tiering conversation is easier than the deletion conversation: nobody is losing coverage, it is moving to where it costs less. ## Step 3: reduce the cost of what remains - **Change detection.** A dynamic matrix generated from changed paths turns a monorepo's fixed 30 legs into the two that the pull request touched. This is often the single largest win. - **Caching that actually hits.** A key fingerprinted on the lockfile with a stable restore-key prefix, and per-leg keys that include the OS and toolchain so legs do not collide. Verify the hit rate rather than assuming it. - **Split slow legs from fast ones.** Put quick checks in one job that fails early, and keep the expensive matrix behind it, so an obvious mistake does not spend the whole product. - **Right-size the runner.** A larger hosted runner or a self-hosted machine can be cheaper in wall-clock terms than parallelising further, especially where the workload is single-process. ## Step 4: set the policy explicitly Write down the rules so the matrix does not regrow: - A stated feedback-latency target for pull-request CI; anything that breaks it must justify itself. - `fail-fast: false` on the small compatibility matrix — with few legs, complete information is worth more than the saved minutes — and the default on large shard matrices. - `max-parallel` only where a shared resource or a shared runner fleet genuinely requires it, not as a general throttle. - A rule for adding a dimension: it comes with the failure it is expected to catch, and it is reviewed after a quarter against its unique yield. - Flaky legs get fixed or removed on a deadline; "re-run it" is not a policy. ## Step 5: mind the check surface Matrix legs surface as individually named checks, and repository policy may require specific check names. Renaming or removing legs can therefore silently leave a required check that never reports — blocking every merge — or quietly drop one that was protecting the branch. Coordinate matrix changes with whoever owns the branch policy, and prefer a single aggregating job that the policy requires, with the matrix behind it, so the check surface stays stable while the matrix evolves. ## What a strong answer sounds like Evidence first (unique yield per leg, latency, re-run rate), then tiering rather than deletion, then cost reduction through change detection and caching, then a written policy with a latency budget — and an explicit note about the required-check surface, which is the part that turns a tidy matrix cleanup into a merge outage.

  • What single metric best identifies a matrix leg that is not earning its place?
    How often it failed when no other leg failed in the same run — its unique yield. A leg that only ever fails alongside the primary platform is duplicating that leg's signal and can move to a scheduled sweep. A leg that has never failed alone in six months is paying nothing for its slot on every pull request.
  • How do you shrink a monorepo matrix without losing coverage?
    Generate the matrix dynamically from the paths a pull request touched, so only affected packages build, and run the full matrix on merge to the default branch and on a schedule. Coverage is unchanged in aggregate; what changes is when it runs, which is what the latency budget is actually about.
  • What breaks when you remove or rename matrix legs that are required checks?
    A required check that no longer reports blocks every merge until the policy is updated, and a removed one silently stops protecting the branch. The durable fix is to require a single aggregating job that depends on the matrix, so the check surface stays constant while legs are added or removed underneath it.
  • How do you handle a matrix leg that fails intermittently?
    Treat it as a defect with an owner and a deadline, not as a fact of life. Quantify the re-run rate, isolate whether it is the test, the environment, or resource contention, and fix or delete it. A leg engineers habitually re-run has negative value: it costs minutes and it trains people to ignore red.

saying these in an interview costs you the question

  • Optimises runner cost without measuring what legs catch
  • Deletes coverage instead of moving it to a schedule
  • Adds max-parallel as a blanket throttle
  • Ignores that legs are named required checks
  • Accepts routine re-runs as normal for flaky legs

context