skip to content

How do you build a GitHub Actions matrix at runtime from a previous job's output?

level: seniorimportance: should knowfreq 46%

answer

  1. job outputs are always strings
  2. one expression function does the parsing
  3. the producer writes to a file path in the environment
  4. an empty list means zero jobs
  5. either an array for one key or an object for the whole matrix

basics

~20 s

A setup job computes a JSON array, writes it to GITHUB_OUTPUT, and declares it as a job output. The downstream job depends on it with needs and sets strategy.matrix to fromJSON of that output, which parses the string into the matrix structure.

solid answer

~50 s

Matrix values must be known when the job is scheduled, but they can come from an earlier job. A setup job runs a script — often a directory scan, a changed-files query, or a `gh` API call — that emits a JSON array, writes it with `echo "targets=$JSON" >> "$GITHUB_OUTPUT"`, and exposes it under the job's `outputs:`. The consuming job declares `needs: setup` and writes `matrix: ${{ fromJSON(needs.setup.outputs.targets) }}`, because job outputs are always strings and `fromJSON` is what turns the string back into a list or object the matrix expander can use. Emit compact single-line JSON, and guard the empty case: an empty array produces zero jobs, and downstream jobs using `needs` on a skipped job need an explicit `if:` such as one built on `always()` or a non-empty check. The same function is also how you read a JSON string anywhere else in an expression.

code

yaml · 26 lines
yaml
jobs:
  setup:
    runs-on: ubuntu-latest
    outputs:
      services: ${{ steps.list.outputs.services }}
    steps:
      - uses: actions/checkout@v4
      - id: list
        run: |
          json=$(ls -d services/*/ | xargs -n1 basename | jq -R -s -c 'split("\n")[:-1]')
          echo "services=$json"
          echo "services=$json" >> "$GITHUB_OUTPUT"

  build:
    needs: setup
    if: needs.setup.outputs.services != '[]'
    runs-on: ubuntu-latest
    strategy:
      fail-fast: false
      matrix:
        service: ${{ fromJSON(needs.setup.outputs.services) }}
    steps:
      - uses: actions/checkout@v4
      - run: ./build.sh "$SERVICE"
        env:
          SERVICE: ${{ matrix.service }}

go deeper

for a junior

Recognise the pattern: one job produces a value, a later job with needs: reads it. Know that a matrix can come from that value rather than being written out in the YAML.

for a middle

Explain why fromJSON is necessary — outputs are strings — and write the three pieces: GITHUB_OUTPUT, the job outputs: map, and matrix: ${{ fromJSON(needs.x.outputs.y) }}.

for a senior

Cover the operational edges: compact JSON, the empty-list skip and its effect on required checks, the 256-job ceiling, the added critical-path cost of the generator, and not splicing generated values into shell commands.

for a principal

Own when dynamic beats static: monorepo change detection, the feedback-latency versus coverage trade, how skipped legs interact with required checks, and how the generator itself is tested so a bad list cannot silently disable CI.

## Why a dynamic matrix A static matrix is fine while the list of things to build is stable. It stops being fine when the list lives in the repository — one job per service directory in a monorepo, one job per changed package, one job per supported version read from a config file. Duplicating that list in YAML means it silently drifts. Generating it means the workflow follows the repository. ## The mechanic Three pieces have to fit together. **1. Produce JSON in a step and expose it as a job output.** jobs: setup: runs-on: ubuntu-latest outputs: services: ${{ steps.list.outputs.services }} steps: - uses: actions/checkout@v4 - id: list run: | json=$(ls -d services/*/ | xargs -n1 basename | jq -R -s -c 'split("\n")[:-1]') echo "services=$json" >> "$GITHUB_OUTPUT" Step outputs are written as `name=value` lines to the file named by `GITHUB_OUTPUT`, then surfaced at job level through the job's `outputs:` map. **2. Parse it in the consumer.** Job outputs are strings, so the matrix cannot use the value directly: build: needs: setup strategy: matrix: service: ${{ fromJSON(needs.setup.outputs.services) }} runs-on: ubuntu-latest steps: - run: ./build.sh "${{ matrix.service }}" `fromJSON` parses a JSON string into a value the expression engine understands — here an array bound to the `service` matrix key. **3. Choose the shape.** Two forms are common. Assigning to a single key, as above, needs a JSON **array** of scalars. Assigning to the whole matrix — `matrix: ${{ fromJSON(needs.setup.outputs.config) }}` — needs a JSON **object** with the same structure the YAML would have had, for example `{"include":[{"service":"api","os":"ubuntu-latest"}]}`. The object form is what you use when the legs are heterogeneous, since it can carry an `include` list directly. ## Practical rules - **Emit compact, single-line JSON.** `jq -c` exists for this. A multi-line value needs the heredoc delimiter form of `GITHUB_OUTPUT` and is easy to get wrong. - **Handle the empty list.** An empty array produces zero matrix jobs and the job is skipped, which then makes anything depending on it skip too. Either emit a placeholder, or gate the consumer with an explicit condition, or make downstream jobs tolerate the skip with an `if:` that uses `always()` plus a result check. - **Print the JSON before publishing it.** A malformed string fails at expression evaluation with a message that points at the matrix rather than at the generator, so `echo` it in the setup job to make debugging one glance instead of five runs. - **Keep the generator cheap.** The setup job's own startup — runner allocation, checkout — is pure overhead added in front of every downstream job, so it should be a small ubuntu job with a shallow checkout, not a full build. - **Respect the ceiling.** One matrix in a run is capped at 256 jobs. A generator over a large monorepo can hit that; cap or bucket the list yourself rather than discovering the limit on a busy day. - **Treat generated values as data.** If any part of the JSON derives from a branch name, a pull-request title, or another attacker-influenceable field, do not splice it into a `run:` command via `${{ }}`; pass it through `env:` and quote it. ## Where the value usually comes from Common generators are a directory listing (one leg per service), a parse of a config or manifest file with `jq`, a changed-paths computation so only affected packages build, and a `gh api` call to list something like open release branches. The changed-paths variant is the highest-value one in a monorepo: it turns a fixed 30-job matrix into a two-job matrix on the average pull request. ## Summary answer "Setup job writes a JSON array to `GITHUB_OUTPUT` and exposes it as a job output; the build job takes `needs:` on it and sets `matrix:` to `fromJSON(...)` of that output. `fromJSON` is required because outputs are strings, and the empty-list case has to be handled deliberately."

  • Why is fromJSON required rather than referencing the output directly?
    Because job and step outputs are strings. A matrix key needs a list, and the whole-matrix form needs an object, so the string has to be parsed. `fromJSON` converts the JSON text into a structured value the expression engine can expand, and it is equally useful for reading any JSON string in an `if:` or in an input.
  • What happens when the generator produces an empty JSON array?
    The matrix expands to zero jobs, so the consuming job does not run and is reported as skipped. Anything that declares `needs:` on it is skipped in turn, which can quietly make a required check disappear. Handle it explicitly: gate the consumer on a non-empty condition, or emit a no-op placeholder leg.
  • How do you generate a heterogeneous matrix where legs have different keys?
    Emit a JSON object containing an `include` array and assign it to the whole matrix: `matrix: ${{ fromJSON(needs.setup.outputs.config) }}` where the string is `{"include":[{"service":"api","os":"ubuntu-latest"},{"service":"web","os":"windows-latest"}]}`. That mirrors the include-only static form and avoids generating an unwanted cross product.
  • What is the cost of adding a setup job in front of the matrix?
    A whole extra job's startup — queueing, runner allocation, checkout — sits in front of every downstream leg and lengthens the critical path on every run. Keep the generator to a small Linux job with a shallow checkout and a fast script, and only pay for it where the list genuinely varies rather than for a stable set of targets.

saying these in an interview costs you the question

  • Assigns the raw output string to matrix without parsing
  • Emits pretty-printed multi-line JSON to GITHUB_OUTPUT
  • Ignores the empty-array case and loses a required check
  • Forgets needs: so the output is unavailable
  • Runs a full build in the generator job

context