In GitLab CI, how do you pass a file produced by one job to a job in a later stage, and what controls which artifacts a job downloads?
answer
- each job starts on a clean working directory
- the producer declares, the consumer usually does not
- later stages download earlier stages by default
- an empty list turns the download off
- cache is for inputs, this is for outputs
basics
~10 sDeclare the file under artifacts:paths in the producing job; GitLab uploads it and later-stage jobs download artifacts from all earlier-stage jobs automatically. Narrow that with dependencies: (an empty list downloads nothing) or with needs:.
solid answer
~50 sThe producing job lists the file under `artifacts:paths:`. When the job finishes, the runner uploads those paths to GitLab as a job artifact. By default **every job in a later stage downloads the artifacts of every job in every earlier stage** and unpacks them into its working directory, so the consumer usually needs no extra keyword at all. That default is often too broad: `dependencies:` restricts the download to a named list of jobs, and `dependencies: []` disables it entirely, which is the standard fix for a deploy job that was pulling gigabytes of build output it never used. If a job uses `needs:`, artifact download follows the `needs` list instead of the stage rule, and `needs: [{job: build, artifacts: false}]` keeps the ordering without the download. Artifacts are also how `artifacts:reports` feeds test, coverage and security results into the merge request widget.
code
yaml · 20 linesbuild:
stage: build
script: make dist
artifacts:
paths:
- dist/
expire_in: 1 week
test:
stage: test
script: make test
artifacts:
when: always
reports:
junit: reports/junit.xml
notify:
stage: deploy
dependencies: []
script: ./post-to-chat.shgo deeper
Know that each job starts clean, that artifacts:paths is how a file travels forward, and that a later-stage job usually receives it without extra configuration.
Explain the default download rule and how dependencies: and needs: narrow it, including dependencies: [] to switch it off. Distinguish artifacts from cache in one clear sentence.
Show that you think about cost and correctness: expire_in and storage growth, when: always for failure evidence, narrow paths so downstream jobs start fast, and dotenv reports rather than ad-hoc file parsing.
Own the policy across teams: artifact retention against storage spend, which reports are mandatory on merge requests, and the rule that a released binary is the promoted artifact rather than something a later job rebuilds.
## Artifacts are the supported way to move files between jobs Each GitLab CI job runs in its own environment — a fresh container, a fresh working directory, often a different machine from the job before it. Nothing written to disk in one job is visible to the next unless it is explicitly carried across. Artifacts are that carrier: the runner archives the declared paths, uploads them to GitLab, and later jobs download and unpack them. ```yaml build: stage: build script: ./gradlew assemble artifacts: paths: - build/libs/*.jar expire_in: 1 week deploy: stage: deploy script: ./upload.sh build/libs/app.jar ``` `deploy` needs no keyword to receive the jar: it is in a later stage, so it downloads `build`'s artifacts automatically. ## Paths and the rules around them `artifacts:paths` takes paths relative to the project directory; you cannot reach outside it with `../`. Globs are supported. Related keywords: - `artifacts:expire_in` — retention (`30 days`, `1 week`, `never`). Without it the instance default applies. Retention matters: artifacts are the single biggest consumer of object storage on a busy GitLab instance. - `artifacts:when` — `on_success` (default), `on_failure`, or `always`. Use `always` for logs and screenshots you want precisely when the job failed. - `artifacts:exclude` — drop paths that a glob would otherwise sweep in. - `artifacts:name` — names the downloadable archive, commonly built from `CI_COMMIT_REF_SLUG`. - `artifacts:reports` — typed reports (JUnit, coverage, dotenv and others). These are parsed by GitLab and surfaced in the merge request rather than just stored. `artifacts:reports:dotenv` is special: the variables in that file become environment variables in the jobs that consume the artifact, which is the sanctioned way to pass a computed *value*, not a file, to a later job. ## Controlling the download side The default — later stages download everything from earlier stages — is convenient and quickly becomes wasteful. Two keywords narrow it: `dependencies:` takes an explicit list of job names whose artifacts should be fetched. The listed jobs must be in an earlier stage (or be in the job's `needs`). An empty list is the important case: ```yaml deploy: stage: deploy dependencies: [] script: ./trigger-release.sh ``` That job now downloads nothing, which can turn a multi-minute job start into seconds on a pipeline with large build outputs. `needs:` also implies artifact selection: a job with `needs:` downloads artifacts from exactly the jobs it needs, ignoring the stage rule. Per entry you can set `artifacts: false` to keep the ordering dependency without transferring files. ## Artifacts versus cache — the distinction interviewers probe They look similar and solve opposite problems: - **Artifacts** are *output*: build results and reports, produced by one job and consumed by another or by a human downloading them. They are versioned per pipeline, guaranteed to be there if the producing job succeeded, and eventually expire. - **Cache** is *input reuse*: dependency directories such as `node_modules/` or `~/.m2`, restored to make a job faster. A cache miss is not an error — the job must still work when the cache is empty. Using cache to move build output between jobs is a classic mistake: it works until the day the runner that has the cache is not the runner that gets the job, and then you ship whatever was in the directory. A related failure: a job that consumes an artifact but does not declare a dependency on the producer works by accident when the producer happens to be in an earlier stage, then breaks the moment someone reorders stages or adds `needs:`. ## Practical notes - If the producing job fails and `artifacts:when` is the default, nothing is uploaded — so a failing test job leaves no report unless you set `when: always`. - Artifacts are also downloadable from the job page and through the API, which is how release pipelines fetch the exact binary a build produced. - Size limits are enforced per job by the instance; oversized artifacts fail the upload, not the script, so the job can go red at the very end with a confusing message. - Declare the narrowest paths that are actually needed downstream. `artifacts: paths: [.]` is the anti-pattern that makes every later job slow.
- What is the difference between `cache:` and `artifacts:` in GitLab CI?Artifacts carry a job's output forward — build results and reports, uploaded on completion, guaranteed present for consumers, expiring on a schedule. Cache speeds up inputs such as dependency directories and is best-effort: a miss is normal and the job must still succeed. Using cache to hand build output to a later job is unsafe because the cache is not guaranteed to be restored.
- How does one GitLab CI job pass a computed value, not a file, to a later job?Write it as `KEY=value` lines to a file and declare that file under `artifacts:reports:dotenv`. GitLab parses it and injects those keys as environment variables into jobs that consume the artifact, so the later job reads `$KEY` directly. Plain artifacts would only give it a file to parse itself.
- A test job fails and you find no JUnit report attached to the merge request. Why?Artifacts default to `when: on_success`, so a failing job uploads nothing. Set `artifacts:when: always` on the test job so the report is uploaded on failure too — which is exactly the run you wanted the report for. The same applies to logs and screenshots collected for debugging.
saying these in an interview costs you the question
- Assumes files persist between jobs without artifacts
- Uses cache to hand build output to the next job
- Thinks the consumer must always name dependencies explicitly
- Believes artifacts upload even when the job fails by default
- Declares the whole workspace as an artifact path