skip to content

CI/CD Concepts

The tool-agnostic vocabulary of delivery: what a stage, runner, artifact, environment and deployment strategy actually are, independent of which YAML dialect you type them in. Interviewers ask these first because a candidate who only knows one vendor's syntax cannot design a pipeline for a different shop.

on this pageshow

questions

page 1 of 2

A CI/CD pipeline builds the application from source separately for the dev, staging and production deployments. What does the "build once, deploy many" principle say to do instead, and what risk does rebuilding per environment introduce?

level: juniorimportance: must knowfreq 74%

answer

  1. the evidence chain from test to production
  2. one build, many deployments
  3. same commit, two different artifacts
  4. configuration injected, never compiled in
  5. rollback redeploys, it never rebuilds

basics

~20 s

Build once, deploy many means the pipeline produces a single versioned artifact and moves that exact artifact through every environment. Rebuilding per environment means production runs bits nobody tested, so the staging result no longer proves anything.

solid answer

~50 s

The pipeline should compile and package exactly once, publish an immutable, versioned artifact — a container image, a jar, a wheel, a signed binary — and have every later stage deploy *that* artifact by its identity. Deploy jobs take an artifact identity as input; they never invoke the compiler. Rebuilding per environment breaks the evidence chain: two builds of the same commit can differ because dependency ranges, base images, system packages and the runner's toolchain all resolve at build time, so what production runs is not what staging approved. It also wrecks rollback — you can no longer redeploy the thing that was working, only rebuild it, which takes a full build cycle during an outage and may not reproduce the same bits. The precondition is that per-environment configuration lives outside the artifact and is injected at deploy time.

go deeper

for a junior

Be ready to state the principle plainly — one build produces one artifact that every environment deploys — and to name the risk in a sentence: what you tested is not what ships.

for a middle

Explain the mechanics of drift: dependency ranges, base images and runner toolchains all resolve at build time, so two builds of one commit can differ. Show where configuration enters instead.

for a senior

Show the operational consequences you have lived: rollback that becomes a rebuild during an incident, an unanswerable "what is running right now", and a diagnosis with two variables instead of one.

for a principal

Own the tradeoff between artifact storage cost and recovery time, and set the organisational rule — deploy stages carry no build toolchain — plus how you would migrate teams whose configuration is currently compiled in.

## What the principle actually says "Build once, deploy many" puts exactly one place in the delivery path where source becomes a runnable thing: the build stage. That stage compiles, packages and publishes a **versioned, immutable artifact** — a container image, a jar or war, a Python wheel, a tarball, a signed binary — and records its identity. Every later stage, in every environment, consumes that artifact by that identity. Deploy jobs are dumb: they take an artifact identity plus environment configuration and put the two together. They never check out source, never call the compiler, never run a packaging step. ## Why two builds of the same commit are not the same artifact The instinct behind rebuilding is "same commit, same output". That is a property you have to engineer and verify, not one you get for free. Things that routinely drift between two builds of one commit: - **Dependency resolution.** A manifest that permits ranges (`^1.4.0`, `1.4.+`, an unpinned plugin) resolves against a package registry *at build time*. A lockfile narrows this for the application's direct graph, but build plugins, toolchains and container base images are frequently outside it. - **Base images and system packages.** `FROM node:22` or an `apt-get install`/`apk add` line resolves to whatever the upstream index publishes today. - **The build machine.** The hosted runner image is upgraded on its own schedule; the second build may compile with a different compiler patch level or link a different system library. - **Build metadata and ordering.** Timestamps, hostname, locale, file ordering in archives, minifier or code-generator nondeterminism. So the artifact production runs is a *different* artifact from the one your staging tests approved — and nobody can say how it differs, because nothing compared them. ## What that costs **Test evidence.** Every pre-production test you ran was evidence about an artifact you then discarded. A green staging run tells you the staging build works. **Diagnosis.** "Works in staging, fails in production" now has two candidate causes — the environment *and* the artifact — instead of one, and you cannot cheaply eliminate either. **Rollback.** Rollback means running the thing that was working ten minutes ago. If that thing exists only as source, rollback becomes a rebuild: minutes to tens of minutes of build time exactly when you have none, and it can fail outright because an upstream dependency was yanked, a base image tag moved, or a build secret expired. **Auditability.** "What is running in production?" should be answerable with an artifact identity, not with a branch name and a hope. ## The shape of the pipeline instead ``` build: checkout -> test -> package -> publish emits: app@<immutable-id> deploy dev: app@<immutable-id> + dev config deploy stg: app@<immutable-id> + staging config deploy prod: app@<immutable-id> + production config ``` The value that flows down the pipeline is the artifact identity, not the source ref. Everything downstream is parameterised by it. ## What must be true for it to work - **Configuration and secrets live outside the artifact** and are supplied at deploy or start-up. An artifact with a hostname compiled into it cannot be promoted; it can only be rebuilt. - **The artifact store outlives the pipeline run**, with retention that covers your rollback horizon. - **Environments reference an identity that cannot be reassigned** — a content digest or a version that is never republished — so the reference means one specific set of bytes forever. - **The same artifact must be able to run at every size.** Replica counts, memory limits and pool sizes are environment configuration, not build inputs. ## The usual pushback *"Our build is deterministic, so rebuilding is harmless."* Sometimes true, and worth having — reproducibility is a real engineering goal, verified by rebuilding and comparing hashes rather than assumed. But even a perfectly reproducible rebuild costs the wall-clock time you do not have during an incident, and it still gives you no artifact to point at when someone asks what production ran last Tuesday. *"Each environment needs different settings, so each needs its own build."* This confuses code with configuration. The fix is to move the settings out, which is the same work you need for build-once anyway.

  • If the build were fully reproducible, would rebuilding per environment become acceptable?
    It would remove the correctness argument but not the operational one. You would still spend a full build cycle to roll back, still have no stored artifact to identify what production ran, and still depend on every upstream source being reachable at rollback time. Reproducibility is worth having as a verification tool — rebuild and compare hashes — rather than as a licence to rebuild in the deploy path.
  • What does a deploy job actually take as its input in a build-once pipeline?
    An artifact identity plus the target environment. The identity must be immutable — a content digest or a version that is never republished — so the job resolves to exactly one set of bytes. Everything that differs per environment (endpoints, credentials, sizing) arrives as configuration alongside it. The job needs no source checkout and no build toolchain at all.
  • Where does the artifact's identity come from if the pipeline only builds once?
    The build stamps it: a version derived from the commit or tag being built, plus the content-addressed identity the artifact store returns on publish. The pipeline then passes that identity to every downstream stage as an output, and the running application reports the same value, so a health endpoint and the deployment record agree on what is live.

It is the difference between serving the dish you cooked and inspected, and cooking a second one from the same recipe for the customer. The recipe matches; the plate that goes out was never checked.

saying these in an interview costs you the question

  • The same commit always produces the same artifact
  • Rebuilding is fine because the tests run again
  • Rollback just means rebuilding the previous tag
  • Each environment needs its own build to get its settings
  • Containers make builds reproducible automatically

context

open as a page

In a monorepo CI pipeline, what does a changed-file (path) filter do, and what does it fail to notice?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A path filter compares the files a change touched against glob patterns and runs a job only when one matches. It sees file paths only, so it misses packages affected indirectly through dependency edges or shared root files.

open as a page

CI pipelines are usually described as stages that contain jobs that contain steps. What is each of those three units actually responsible for?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A stage is a unit of ordering — a named phase that acts as a barrier before the next one starts. A job is a unit of allocation — one machine and one workspace, and the thing that runs in parallel. A step is a unit of execution — one command inside a job.

open as a page

Under Semantic Versioning 2.0.0, what do the MAJOR, MINOR and PATCH fields mean, and how do you decide which one a given change bumps?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Semantic Versioning encodes compatibility: bump MAJOR for a backward-incompatible change to the public API, MINOR for backward-compatible new functionality, PATCH for a backward-compatible bug fix. The public contract decides the bump, not how much code changed.

open as a page

In a CI/CD system, what is the difference between a provider-hosted runner and a self-hosted runner, and what does choosing self-hosted actually change?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A provider-hosted runner is a fresh machine the CI vendor creates per job, bills per minute, then destroys. A self-hosted runner is a machine you own and register yourself: usually cheaper at volume and able to reach private networks, but you inherit patching, isolation and cleanup.

open as a page

What does putting a credential in a CI platform's secret store actually protect against, and what does it not?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A CI secret store keeps the value out of the repository and out of history, encrypts it at rest, hides it from the UI once saved, and masks it in logs. It does not limit what the credential can do, or stop a job that receives it from sending it anywhere.

open as a page

Why is a CI/CD pipeline treated as production infrastructure rather than as a developer convenience, and what does that change about how you secure it?

level: juniorimportance: must knowfreq 48%

basics

~10 s

A CI/CD pipeline holds the credentials that reach production and its output is deployed without anyone re-reading it, so compromising the build is equivalent to compromising every system that build can deploy to.

open as a page

A web application's CI pipeline bakes the API base URL into the JavaScript bundle at build time, so staging and production each get a separately built bundle. How would you restructure this so one artifact serves both environments?

level: middleimportance: must knowfreq 58%

basics

~20 s

Take the URL out of the build so one bundle serves every environment, and supply it at deploy time — substituted into the served files at start-up, fetched from a per-environment config file, or derived from the serving origin.

open as a page

How would you design the cache key for a CI job's dependency cache, and what are the two ways a cache key can be wrong?

level: middleimportance: must knowfreq 68%

basics

~20 s

Derive the key from a digest of every input that determines the cached content: the lockfile, the toolchain version, the operating system and architecture. Too-specific keys never hit and waste effort; too-loose keys hit while stale and silently corrupt builds.

open as a page

In a rolling update you can configure how many instances may run above the desired count (surge) and how many of the desired count may be missing at once (unavailable). What do these two budgets trade off, and how does each extreme setting fail?

level: middleimportance: must knowfreq 60%

basics

~20 s

Surge buys capacity with spare resources: extra instances start before old ones stop. Unavailable buys speed with reduced capacity: old instances stop first. Surge zero degrades a busy service mid-rollout; unavailable zero stalls the rollout when there is no headroom to schedule the extras.

open as a page

In a CI/CD system, what does it mean to model a deployment target such as staging or production as a first-class environment object rather than just a set of variables?

level: middleimportance: must knowfreq 60%

basics

~20 s

A first-class environment is a named deployment target the delivery platform owns: it carries its own scoped credentials, protection rules that decide who and what may deploy to it, and a recorded deployment history that serves as the audit trail.

open as a page

In a monorepo build tool such as Nx, Turborepo or Bazel, how is the "affected" set of targets computed, and why is that stronger than a path filter?

level: middleimportance: must knowfreq 62%

basics

~20 s

The tool builds a dependency graph of the repository's projects, maps the changed files to the projects that own them, then walks the graph backwards to include every project that transitively depends on those. Path filters see only the directly edited directories.

open as a page

What are the common ways a CI/CD pipeline run gets started, and how do they differ in what they cost and what feedback they give?

level: middleimportance: must knowfreq 68%

basics

~20 s

Runs start from a branch push, a pull/merge request, a tag, a schedule, a manual button, an upstream pipeline, or an external API call. Each differs in how often it fires, who is waiting for the result, and how much compute it burns per merge.

open as a page

In a CI/CD pipeline, what is OIDC workload-identity federation, and what does the pipeline actually exchange for a cloud credential?

level: middleimportance: must knowfreq 72%

basics

~20 s

OIDC federation replaces a stored cloud key with a trust relationship. The CI platform mints a short-lived signed identity token describing the specific run; the cloud verifies its signature and claims against a trust policy and hands back temporary credentials.

open as a page

Nearly everything a CI job pulls in at run time — shared build steps, container images, tool installers, dependency version ranges — is referenced by a mutable name. Why is that a code-execution risk, and what does pinning to a digest or commit change?

level: middleimportance: must knowfreq 52%

basics

~20 s

A mutable reference resolves at build time to whatever the publisher has there now, so a third party can change what your build executes with no commit and no review in your repository. A digest names the exact bytes, making substitution impossible.

open as a page

In a CI pipeline, what is the difference between a cache and a build artifact, and why must a job still succeed when the cache is empty?

level: juniorimportance: should knowfreq 55%

basics

~20 s

A cache is a best-effort copy of data a job could regenerate itself, kept only to make later runs faster; an artifact is a required output a later job or a release consumes. Every job must still succeed on an empty cache.

open as a page

What is a recreate deployment strategy, what does it cost, and when is it still the right choice?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Recreate stops every old instance before starting any new one, so there is a deliberate downtime window and never two versions running at once. Choose it when mixed versions would corrupt shared state or fight over an exclusive resource.

open as a page

In a deployment pipeline, what is a manual approval gate, and what is the pipeline actually doing while it waits for a human to approve?

level: juniorimportance: should knowfreq 58%

basics

~20 s

A manual approval gate pauses a run at a boundary before a protected stage until a permitted person approves it. A platform-level gate holds the run as pending state with no agent allocated, and expires after a configured timeout.

open as a page

A pipeline blocks merges whenever total test coverage falls below 80 percent. What makes a threshold gate like that useful or useless, and what would you gate on instead?

level: middleimportance: should knowfreq 48%

basics

~20 s

A whole-repository coverage threshold is dominated by legacy code, so one change barely moves it and the gate rarely fires on the change that deserves it. Gate on coverage of the lines the change touched, or ratchet the number so it can never fall.

open as a page

In a monorepo that publishes many packages, what is the difference between giving every package one locked version and versioning each package independently?

level: middleimportance: should knowfreq 46%

basics

~20 s

Locked (fixed) mode gives every package the same version number and republishes them together on each release. Independent mode bumps and publishes only the packages that changed, each on its own version line, at the cost of far more bookkeeping.

open as a page

In CI, why can a pipeline whose phases run as strict sequential barriers finish later than the same work expressed as a dependency graph between jobs?

level: middleimportance: should knowfreq 58%

basics

~20 s

A barrier makes every job in a phase wait for the slowest job in the previous phase, so total time is the sum of those slowest jobs. A dependency graph lets each job start as soon as its own inputs are ready, so total time is the longest actual chain.

open as a page

How can a CI job derive the next version number and write the changelog automatically from commit messages, without a human editing a version file?

level: middleimportance: should knowfreq 54%

basics

~20 s

Conventional Commits give each commit a type — fix, feat, or a breaking marker. A release job reads every commit since the last release tag, derives the SemVer bump from the highest-ranked type it finds, generates the changelog from the same messages, then tags and publishes.

open as a page

A CI job can execute in a container on a shared host, in a virtual machine created for it, or directly on a shared bare-metal machine. What does each boundary actually contain, and what leaks across it?

level: middleimportance: should knowfreq 46%

basics

~20 s

A container isolates the filesystem view, processes and network but shares the host kernel, so a kernel bug or a privileged mount escapes it. A virtual machine gives the job its own kernel and is the strongest boundary. A shared bare-metal machine contains nothing: jobs share a user, disk and credentials.

open as a page

A build passes on your team's long-lived shared CI machines but fails on a freshly provisioned one. What kinds of leftover state on a reused runner cause that, and what removes the whole class of problem?

level: middleimportance: should knowfreq 56%

basics

~20 s

A reused runner carries state forward: leftover workspace files, globally installed tools, package-manager config and credentials in the agent user's home directory, stray processes, and cached images. The build silently depends on it. Single-use runners — one job per instance, then destroyed — remove the class.

open as a page

A CI job's shell step embeds the pull request title into a command using the CI engine's template syntax. Why does that hand a contributor code execution on the runner, and how do you write the step safely?

level: middleimportance: should knowfreq 45%

basics

~20 s

The CI engine substitutes the title into the script text before any shell runs, so an attacker-chosen title becomes part of the program rather than data. Bind such values to environment variables and reference them quoted instead.

open as a page

An emergency rollback to a release from six weeks ago fails because the delivery system's retention policy already deleted that build's artifact. How would you design artifact retention so this cannot happen?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Retain by class, not by age alone: expire per-commit and pull-request builds quickly, exempt any artifact an environment has actually run, and keep released artifacts for at least as long as your stated rollback horizon.

open as a page

In a build-once pipeline, promoting a release from staging to production never changes the artifact's bytes. What does promotion actually change, and what must the production deployment reference for that promotion to mean anything?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Promotion changes a reference, not the artifact — the record of which build production is declared to run. For it to mean anything, production must deploy an immutable identity such as a content digest, never a tag that promotion moves.

open as a page

A pull request from an outside fork runs CI and writes an entry into a shared dependency cache that later builds on the main branch restore. Why is that a supply-chain risk, and what contains it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Restoring a cache places files chosen by whoever wrote it onto disk, and a build then executes them with the trusted job's credentials. Containment is scoping: untrusted refs get their own cache namespace, read from trusted entries but never write them.

open as a page

A CI test suite that takes 30 minutes on one job is split across 10 parallel jobs, but wall-clock time only drops to 12 minutes. What explains the gap, and how would you choose the split count?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Each parallel job repays a fixed cost — queue wait, checkout, dependency restore, image pull, service startup — and that cost does not divide. Add unbalanced shards and shared bottlenecks and the speedup flattens. Choose the split where the next shard's saving stops exceeding that fixed cost.

open as a page

Your green environment passed every smoke test, but seconds after the blue-green cutover sent it 100% of production traffic, latency spiked and errors appeared for several minutes before settling. What causes this, and what does keeping a warm standby actually involve?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A standby that passed smoke tests at near-zero traffic is still cold: empty caches, unestablished connection pools, uncompiled hot paths and an autoscaler sized for no load. Warming means giving it real load — shadow traffic, pre-scaling and pre-opened pools — before the switch.

open as a page

showing 1–30 of 52