skip to content

In GitLab CI, what does adding `needs:` to a job change about when it starts and which artifacts it receives?

level: middleimportance: should knowfreq 62%

answer

  1. stages are a barrier, this is an edge
  2. start when your dependency finishes, not your stage
  3. the download set narrows to what you named
  4. empty array means start immediately
  5. failure stops only the branch below it

basics

~20 s

needs: turns stage-ordered execution into a DAG: the job starts as soon as the jobs it lists have finished, even if its own stage has not been reached, and it downloads artifacts only from those jobs instead of from all earlier stages.

solid answer

~50 s

By default a GitLab pipeline is a sequence of stages: no job in `test` starts until every job in `build` has finished. `needs:` replaces that coarse barrier with per-job dependencies, so a job runs as soon as its named jobs complete — potentially long before its own stage would have started. That turns the pipeline into a directed acyclic graph and is the main way to cut wall-clock time on a wide pipeline where one slow job was holding an entire stage. It also changes artifact behaviour: a job with `needs:` downloads artifacts **only** from the jobs it needs, not from every earlier stage, and `artifacts: false` on an entry keeps the ordering without the transfer. `needs: []` means no dependencies at all, so the job starts immediately when the pipeline begins, regardless of its stage. `optional: true` tolerates a needed job that `rules:` excluded from this pipeline.

code

yaml · 19 lines
yaml
stages: [build, test, deploy]

pull-base-image:
  stage: deploy
  needs: []
  script: docker pull registry.example.com/base:latest

build:linux:
  stage: build
  script: make linux
  artifacts:
    paths: [out/linux/]

test:unit:
  stage: test
  needs:
    - job: build:linux
      artifacts: true
  script: ./run-tests.sh out/linux/

go deeper

for a junior

Know that jobs run stage by stage by default and that needs: lets a job start earlier by naming the jobs it depends on. Be able to read a small DAG pipeline.

for a middle

Explain both effects together: scheduling by dependency rather than by stage, and artifact download narrowing to the needed jobs. Mention needs: [] for an immediate start and optional: true for jobs a rule may exclude.

for a senior

Demonstrate that you add needs: against a measured critical path, verify artifacts still arrive, and account for the change in failure propagation when stage barriers stop protecting downstream work.

for a principal

Weigh pipeline speed against comprehensibility across many repositories: how much DAG complexity a team can maintain, whether stage discipline should remain the default convention, and what pipeline duration target justifies the change.

## The default model, and what it costs A `.gitlab-ci.yml` declares `stages:` and assigns each job to one. Execution is strictly staged: every job of stage *n* must reach a final state before any job of stage *n+1* starts. This is easy to reason about and it wastes time. If `build` contains a two-minute Linux build and a twenty-minute Windows build, the fast lint and unit-test jobs that depend only on the Linux build still wait twenty minutes. `needs:` breaks that barrier per job: ```yaml stages: [build, test, deploy] build:linux: stage: build script: make linux artifacts: paths: [out/linux/] build:windows: stage: build script: make windows test:linux: stage: test needs: [build:linux] script: ./run-tests.sh out/linux/ ``` `test:linux` now starts the moment `build:linux` succeeds, while `build:windows` is still running. GitLab shows such a pipeline in the "Job dependencies" view as a graph rather than columns. ## What `needs:` changes, precisely **Scheduling.** The job becomes eligible as soon as all needed jobs have completed successfully. Stages still exist and still determine the visual grouping and the fallback ordering for jobs *without* `needs:`, but for a job with `needs:` the stage is no longer a gate. **Artifact selection.** Without `needs:`, a job downloads artifacts from all jobs in all earlier stages. With `needs:`, it downloads from exactly the needed jobs. This is usually what you want, and it is also a trap when you add `needs:` for speed and silently lose an artifact the job was quietly relying on. Per entry: ```yaml deploy: needs: - job: build:linux artifacts: true - job: sign artifacts: false script: ./deploy.sh ``` **Immediate start.** `needs: []` (an empty array) declares no dependencies, so the job is eligible the instant the pipeline is created even if it sits in a late stage. This is how you start a slow, independent job — pulling a large base image, warming a environment — in parallel with everything else. **Optional needs.** If a needed job may or may not exist in a given pipeline because its own `rules:` excluded it, a plain `needs:` reference makes the pipeline invalid. `needs: [{job: build:windows, optional: true}]` makes the dependency apply only when the job is present. ## Constraints to know - A needed job must be in the same stage or an earlier one; you cannot need a job that comes later. (Same-stage `needs:` is allowed, which is how you sequence two jobs inside one stage.) - There is a documented cap on how many jobs one job may need — 50 by default on a modern GitLab, configurable on self-managed instances. A pipeline that bumps into it is usually a design smell. - The graph must be acyclic; GitLab rejects a cycle at pipeline creation. - `needs:` can reach outside the pipeline in specific forms: `needs:project` downloads artifacts from another project's pipeline, and `needs:pipeline:job` fetches from a parent or upstream pipeline. These are artifact-fetch mechanisms, not scheduling links across pipelines. ## Where it changes behaviour you did not intend The subtle one is failure propagation. In a staged pipeline, a failure in `build` stops the entire `test` stage. In a DAG, only the branches downstream of the failed job are affected; unrelated branches keep running. That is normally desirable, but a team that used the stage barrier as an implicit "stop everything if the build is broken" safeguard will notice deploy-adjacent jobs continuing where they previously did not. The second is readability. A modest number of `needs:` edges makes a pipeline both faster and clearer. A pipeline where every job names five needed jobs is harder to reason about than the stage list it replaced, and the graph view becomes the only way to understand it. Add `needs:` where a real critical path exists, measure the change in pipeline duration, and leave the rest on stages. ## Sanity checks - After adding `needs:`, confirm the job still receives every artifact it reads — the download set just narrowed. - Check that no job now runs before something that was implicitly protecting it, such as a lint or policy gate that was relying on stage order rather than an explicit dependency. - Compare pipeline duration before and after; if the critical path did not move, the added complexity bought nothing.

  • What does `needs: []` do that omitting `needs:` entirely does not?
    Omitting it leaves the job under stage ordering, so it waits for every earlier stage. An empty array declares that the job depends on nothing, so it becomes eligible the moment the pipeline is created even though it sits in a later stage — useful for a slow, independent job you want running in parallel from the start. It also means the job downloads no artifacts.
  • You add `needs:` to speed a job up and it starts failing with a missing file. What happened?
    Artifact selection changed with it. Without `needs:` the job downloaded artifacts from every job in every earlier stage, and it was silently consuming one you never declared. With `needs:` it fetches only from the jobs it needs. Add the missing producer to the `needs:` list, or set `artifacts: true` on the entry that produces the file.
  • How does a failure behave differently in a DAG pipeline compared with a purely staged one?
    In a staged pipeline a failure blocks the whole next stage. With `needs:`, only jobs downstream of the failed job are blocked; independent branches continue to completion. Teams that relied on the stage barrier as a global stop should replace that assumption with an explicit dependency or a gating job rather than expecting stages to do it.

saying these in an interview costs you the question

  • Thinks needs: only reorders the pipeline graph visually
  • Assumes a job with needs: still downloads all earlier artifacts
  • Believes stages are removed once any job uses needs:
  • Tries to need a job in a later stage
  • Adds needs: everywhere without checking the critical path

context