skip to content

In an azure-pipelines.yml file, how do stages, jobs and steps nest, and which of the three is the unit that gets an agent?

level: juniorimportance: must knowfreq 78%

answer

  1. three nesting levels, one allocation unit
  2. the agent is handed to something
  3. steps share a workspace; jobs may not
  4. stages sequential, jobs parallel by default
  5. a bare steps: block implies both

basics

~10 s

Stages contain jobs and jobs contain steps. The job is the unit of agent allocation: one agent runs all of a job's steps in one workspace, so two separate jobs share no filesystem.

solid answer

~50 s

An `azure-pipelines.yml` file nests three constructs. A **stage** is the coarsest grouping and normally maps to a phase of delivery such as build or test. A **stage** contains **jobs**, and a **job** contains **steps**, where each step is a single `task:`, `script:`, `bash:` or `pwsh:` invocation. The important part is that the *job* — not the stage and not the step — is what Azure DevOps allocates an agent to. Every step in a job runs on the same agent, in the same workspace, in written order; two jobs may land on two different agents and get two different working directories, even inside one stage. That is why crossing a job boundary requires publishing and downloading a pipeline artifact or declaring an output variable, while crossing a step boundary requires nothing. Stages run sequentially by default; jobs inside a stage run in parallel by default, bounded by how many parallel job slots the organisation has.

code

yaml · 32 lines
yaml
trigger:
  branches:
    include:
      - main

stages:
  - stage: Build
    jobs:
      - job: Compile
        pool:
          vmImage: ubuntu-latest
        steps:
          - script: npm ci && npm run build
            displayName: Build
          - task: PublishPipelineArtifact@1
            inputs:
              targetPath: dist
              artifact: web

  - stage: Test
    dependsOn: Build
    jobs:
      - job: Unit
        pool:
          vmImage: ubuntu-latest
        steps:
          - task: DownloadPipelineArtifact@1
            inputs:
              artifact: web
              path: dist
          - script: npm ci && npm test
            displayName: Unit tests

go deeper

for a junior

Be able to say the nesting out loud — stages hold jobs, jobs hold steps — and name the job as the thing that gets an agent. Knowing that steps share a workspace and jobs do not is the point of the question.

for a middle

Explain the defaults and how to change them: sequential stages, parallel jobs, dependsOn to build an explicit graph, and the implicit stage/job Azure inserts around a bare steps: block. Show how a value or a file crosses a job boundary.

for a senior

Demonstrate the operational consequence: job splits cost artifact upload and download time and a fresh checkout, so more jobs is not automatically faster. Be ready to explain why a pipeline that fans out twenty jobs still executes a few at a time.

for a principal

Own the tradeoff between pipeline shape and delivery speed: how many parallel slots the organisation buys, whether stage boundaries should mirror approval boundaries, and when a wide job graph stops paying for its own coordination overhead.

## The three levels Azure Pipelines describes a run with three nested constructs, written top-down in a YAML file that is conventionally named `azure-pipelines.yml`: - **stages** — the coarsest grouping, usually one per delivery phase (`Build`, `Test`, `Deploy`). - **jobs** — a unit of work inside a stage. - **steps** — the individual commands: a `task:` (a packaged unit from the task library, referenced as `TaskName@version`), or one of the script shorthands `script:`, `bash:`, `pwsh:`, `powershell:`. ```yaml stages: - stage: Build jobs: - job: Compile pool: vmImage: ubuntu-latest steps: - script: ./gradlew assemble displayName: Compile - task: PublishPipelineArtifact@1 inputs: targetPath: build/libs artifact: app ``` ## The job is the allocation unit The single fact that makes this hierarchy behave the way it does: **an agent is assigned per job**. Azure DevOps takes a job off the queue, finds an agent in the job's pool that satisfies its demands, and that one agent executes every step of the job, sequentially, in one workspace directory. Steps therefore share everything for free — the filesystem, the installed toolchain, the process environment, the checked-out source. Jobs share none of that. Two jobs in the same stage can be dispatched to two different agents on two different machines. Even on a single self-hosted agent, there is no guarantee both jobs land there, and Azure will not reuse one job's workspace as another's. So the moment you split work into two jobs, you have to *transport* anything they share: - **Files** — publish from the producer (`PublishPipelineArtifact@1`) and download in the consumer (`DownloadPipelineArtifact@1`); a `checkout` also happens per job by default. - **Values** — mark a variable as an output variable when you set it, then read it through the `dependencies` context in the consuming job. This is the number-one surprise for people new to the platform: a `cd` or a `npm install` in job A means nothing to job B. ## Ordering: sequential stages, parallel jobs By default, **stages run one after another** in the order written, and a stage only starts if the previous one succeeded. **Jobs inside a stage run in parallel** by default. Both defaults are overridden with `dependsOn`, which turns the flat list into an explicit dependency graph: ```yaml jobs: - job: A steps: [ { script: echo a } ] - job: B steps: [ { script: echo b } ] - job: C dependsOn: [ A, B ] # fan-in: waits for both steps: [ { script: echo c } ] ``` An empty `dependsOn: []` on a stage removes its implicit dependency on the previous stage, letting stages run concurrently. How much of that parallelism you actually get is capped by the parallel-job slots the organisation has bought or been granted, so a fan-out of twenty jobs may still execute a few at a time. ## Implicit shorthands Azure lets you omit the outer levels when you do not need them, and quietly supplies them: - A file with only `steps:` at the root becomes **one stage containing one job containing those steps**. - A file with only `jobs:` becomes one implicit stage. That shorthand is why small pipelines look nothing like the three-level model, and why the model still applies to them: the single implicit job is still what receives the agent. ```yaml trigger: [ main ] pool: vmImage: ubuntu-latest steps: - script: npm ci && npm test ``` One caveat worth remembering: as soon as you introduce `stages:`, every job must live under a stage — you cannot mix a top-level `steps:` or `jobs:` with `stages:` in the same file. ## Where `pool` sits Because the job is what gets an agent, `pool` is fundamentally a job-level property. You may write it once at the root or on a stage as a default, but it is inherited *down to each job*, and a job-level `pool` wins over the inherited one. A stage with two jobs can perfectly well run one on a Linux image and the other on Windows. ## Why interviewers ask The question separates people who have written a pipeline from people who have only read one. The three-level vocabulary is easy to recite; the consequence — that the job boundary is a machine boundary, so state must be published across it — is what actually shows up as a broken pipeline on someone's first week.

  • If a file contains only a top-level steps: block, what does Azure DevOps actually run?
    It supplies the missing levels: one implicit stage containing one implicit job containing those steps. The run behaves identically to the fully-written-out three-level form, and that single implicit job is still what receives an agent. The shorthand stops working the moment you add an explicit `stages:` key — then every job must live under a stage.
  • Two jobs in one stage both need the compiled binary. How does the second one get it?
    Not from the filesystem — the jobs may be on different agents. The producing job publishes the output with `PublishPipelineArtifact@1`, the consuming job declares `dependsOn` on it so ordering is guaranteed, then retrieves it with `DownloadPipelineArtifact@1`. Without `dependsOn` the jobs would run in parallel and the download would race the publish.
  • Do stages ever run in parallel in Azure Pipelines?
    Yes. A stage implicitly depends on the one before it, but `dependsOn: []` clears that, and an explicit `dependsOn` list can express any DAG — two stages both depending only on `Build` run concurrently. Actual concurrency is still bounded by the organisation's parallel-job slots, so declaring parallelism does not guarantee getting it.

saying these in an interview costs you the question

  • Thinks every step gets its own agent
  • Assumes files written in one job are visible in the next
  • Believes stages are only labels for the run summary
  • Says jobs in a stage always run one after another
  • Thinks pool is a stage-level concept, not a job-level one

context