skip to content

In GitLab, what actually happens when you use the rollback button on an older deployment in an environment's deployment history, and why might it not restore the previous version?

level: seniorimportance: should knowfreq 38%

answer

  1. it re-runs a job, not a snapshot
  2. new deployment appended, history append-only
  3. artifact expiry outlives no incident
  4. floating tag makes rollback a no-op
  5. data changes are not rolled back

basics

~20 s

GitLab re-runs that older pipeline's deployment job against the old commit and records it as a new deployment. It restores the previous version only if that job is reproducible — if it rebuilds from source, depends on expired artifacts, or pulls a mutable tag, you can get something else entirely.

solid answer

~60 s

The button is not a state restore; it is a **job re-run**. GitLab takes the deployment job from the pipeline of the commit you selected, runs it again with that commit checked out, and appends the result as a new deployment at the top of the environment's history — so the history is append-only and the old entry stays where it is. That works beautifully when the deploy job is a thin promotion step that pulls an immutable, digest-identified artifact and points the platform at it. It fails when the job does real work: artifacts from the original pipeline may have passed their expiry and no longer be downloadable; a job that rebuilds from source will rebuild with today's dependencies rather than the ones from six weeks ago; a `docker pull myapp:latest` will fetch the current image regardless of which commit is checked out. Also worth knowing: a project can be configured to skip outdated deployment jobs, which is precisely a job older than the current deployment, so check that setting before you need it in an incident.

go deeper

for a junior

Know that GitLab keeps a deployment history per environment and that rolling back re-runs the deploy job for the older commit rather than undoing anything by magic.

for a middle

Explain that a new deployment is appended rather than the history rewound, and that the replay uses the old commit's script with today's variables, runners and registry contents.

for a senior

Show the failure modes you would check under pressure: expired job artifacts, a rebuild that is not reproducible, a floating image tag that makes the rollback a no-op, and outdated-deployment-job protection blocking the retry.

for a principal

Own the property that makes rollback real — build once, promote an immutable digest, set retention that outlives an incident — and be explicit that stateful changes need a forward-fix strategy instead.

## What GitLab stores Every job carrying `environment:` appends a **deployment** to that environment: commit SHA, pipeline, job, who triggered it, when, and status. The environment page renders that list newest-first and marks the current one. Older entries offer a re-deploy control, usually described as rolling back. ## What the button does It re-runs the deployment job belonging to the pipeline of the commit you chose. Concretely: the old commit is checked out, the job's script runs again, and a **new** deployment is appended to the history — GitLab does not rewind, and the entry you clicked is not resurrected in place. So the history remains a truthful, append-only log: "we deployed A, then B, then A again", not "B never happened". The crucial consequence is that **rollback is only as reliable as the deploy job is reproducible**. GitLab replays a job; it does not snapshot and restore system state. ## Why the replay can produce something else **Artifacts expire.** If the deployment job consumes artifacts produced earlier in the *same* pipeline — a built binary, a rendered manifest bundle — those artifacts have an expiry. A rollback attempted after that window finds nothing to download and either fails outright or, worse, proceeds with a missing input. Long-lived releases need their deployable kept somewhere with retention that outlives the emergency: an artifact whose expiry is set to never for release pipelines, or a package registry entry. **Rebuilds are not reproducible.** A deploy job that compiles or builds an image as part of deploying will resolve dependencies at *replay* time. Floating version ranges, a base image tag that moved, or a package that was yanked mean the artifact you get today is not the artifact that was running last month — you have deployed a new build of old source, which is not the same thing as the thing that worked. **Mutable references defeat the whole exercise.** If the job runs `kubectl set image ... myapp:latest` or `helm upgrade` with a floating tag, the old commit's script still resolves to the newest image. The rollback "succeeds", the history shows the old commit, and the bad code is still live. This is the failure that costs the most incident time, because every signal says it worked. **Configuration drift.** The replay uses the old commit's script but *today's* project and group CI/CD variables, today's runners, and today's cluster. A credential that was rotated, a variable that was renamed, or a runner image that no longer exists breaks a job that ran cleanly at the time. **Outdated-job protection.** GitLab has a project setting that prevents older deployment jobs from running, so that a slow pipeline cannot overwrite a newer deployment. Since a rollback is by definition an older job, that protection interacts with it directly; GitLab exposes a companion option to permit retries for rollback deployments. Whichever way your project is configured, find out *before* an incident rather than while the site is down. ## The shape that makes rollback trustworthy Build once, promote the same bytes, and make the deployment job a thin step that names an immutable reference: ```yaml deploy_prod: stage: deploy script: - helm upgrade --install myapp ./chart \ --set image.digest="$IMAGE_DIGEST" \ --namespace prod environment: name: production url: https://app.example.com ``` where `IMAGE_DIGEST` was resolved and pinned when the commit was first built, and is carried with the commit rather than recomputed. Replaying that job is deterministic: it re-applies a specific, content-addressed artifact. ## Where the button is the wrong tool Redeploying old code does not undo a database migration, restore deleted rows, or un-expire a cache. If the release included a destructive schema change, the honest answer in an interview is that the rollback path is a forward fix or an expand/contract migration pattern, not this button. Similarly, an environment with a lot of state — queues drained, feature flags flipped, external webhooks re-registered — is not restored by re-running one job, and pretending otherwise is how a bad five-minute outage becomes a bad two-hour one. ## What to say in an interview Name the mechanism first (it re-runs the old pipeline's deploy job and appends a new deployment), then the precondition (the job must be reproducible), then the two named failures you have actually seen: expired artifacts and a floating tag that made a successful-looking rollback change nothing.

  • How do you make sure a rollback deploys the exact bytes that were running before?
    Build the artifact once, publish it to a registry, and record its immutable digest with the commit so the deployment job only ever references that digest. Then replaying the job is deterministic. Also set retention on release artifacts long enough that an emergency months later still finds them.
  • Why can a GitLab rollback report success while the bad version is still serving traffic?
    Because the deploy job resolved a mutable reference. If the script sets an image tag such as `latest` or `main`, checking out the old commit changes nothing about what the tag points to, so the platform pulls the current image. The deployment record shows the old commit; the running code is new.
  • What does redeploying an old commit fail to undo?
    Anything outside the application binary: applied database migrations, deleted or transformed rows, consumed messages, flipped feature flags, invalidated caches, and state written into third-party systems. If the release included a destructive schema change, the recovery path is a forward fix or an expand/contract migration, not a redeploy.

saying these in an interview costs you the question

  • Believes rollback restores a saved system snapshot
  • Assumes old pipeline artifacts are still available
  • Rebuilds from source and calls it the same version
  • Thinks a redeploy reverses database migrations
  • Expects the old deployment entry to move back to current

context