Review apps created by your GitLab CI pipeline are never cleaned up after branches merge. Which environment keywords are supposed to stop them, and why does the stop job so often fail to run?
answer
- on_stop names a job, action: stop
- expiry timer resets on each deploy
- the stop job is created early, run later
- deleted branch means no checkout
- stopped record does not mean destroyed
basics
~20 sTeardown comes from environment:on_stop, which names a job with action: stop, plus environment:auto_stop_in for an expiry. Stop jobs fail most often because the job was never created in the original pipeline, or because it tries to check out a branch that has since been deleted.
solid answer
~50 sTwo keywords do the work. `on_stop:` names another job in the same pipeline that carries `environment:action: stop`; GitLab runs that job when the environment is stopped — manually, when the merge request is merged or closed, or when the source branch is deleted. `auto_stop_in:` gives the environment an expiry such as `3 days`, and the countdown restarts on every new deployment, so an active branch stays up and a forgotten one dies. The failures are boringly consistent. The stop job must actually be **created** in the deploy pipeline — if its rules excluded it, `on_stop` points at nothing and the environment can only be stopped as a record. It must declare the same `environment:name` with `action: stop`. And it usually needs `GIT_STRATEGY: none`, because by the time it runs the branch may be gone and the checkout fails before your teardown script ever runs. Also remember `auto_stop_in` with no `on_stop` marks the environment stopped in GitLab while leaving the real infrastructure running.
code
yaml · 11 linesstop_review:
stage: deploy
variables:
GIT_STRATEGY: none
script:
- kubectl delete namespace "$CI_ENVIRONMENT_SLUG" --ignore-not-found
when: manual
allow_failure: true
environment:
name: review/$CI_COMMIT_REF_SLUG
action: stopgo deeper
Know that a review app does not disappear on its own: the deploy job needs on_stop pointing at a teardown job, and auto_stop_in gives it an expiry.
Explain that the stop job is created in the deploy pipeline and executed later, must repeat the same environment name with action: stop, and typically needs GIT_STRATEGY: none.
Diagnose the silent versions: rules that never create the stop job, a checkout failing on a deleted branch, and auto_stop_in with no on_stop leaving live infrastructure behind a clean-looking Environments page.
Treat pipeline-driven cleanup as best-effort and require an independent reconciliation loop plus resource labelling, so a leak is bounded by a scheduled sweep rather than by whether a YAML edit was correct.
## The two keywords ```yaml deploy_review: script: ./deploy.sh "$CI_ENVIRONMENT_SLUG" environment: name: review/$CI_COMMIT_REF_SLUG url: https://$CI_COMMIT_REF_SLUG.review.example.com on_stop: stop_review auto_stop_in: 3 days stop_review: variables: GIT_STRATEGY: none script: ./teardown.sh "$CI_ENVIRONMENT_SLUG" when: manual environment: name: review/$CI_COMMIT_REF_SLUG action: stop ``` `on_stop` names the job that destroys the thing. `auto_stop_in` sets an expiry on the environment, expressed in human units (`30 minutes`, `1 day`, `3 days`, `1 week`). The timer is reset by each new deployment, so a branch under active development keeps its review app alive and an abandoned one expires. When the expiry passes, GitLab stops the environment, which means running its `on_stop` job. GitLab also triggers the stop path when a merge request is merged or closed, and when the source branch is deleted — which is why review apps mostly clean themselves up on well-run projects and why the ones that do not are usually broken in one of a small number of ways. ## Failure one: the stop job was never created The stop job is not conjured on demand. It is created **in the same pipeline as the deploy job**, in a state where it does not run, and GitLab executes it later when the environment is stopped. If the job's own selection rules excluded it from that pipeline — a `rules:` block that only matches merge request events while the stop is triggered in a different context, or a stage that was skipped — then `on_stop` refers to a job that does not exist for that environment. GitLab will then let you "stop" the environment in the UI, but all that does is flip the record to stopped: no script runs, and the namespace, DNS record and database survive. The practical rule is that the stop job must be created under exactly the same conditions as the deploy job, and typically in the same stage, so the two always appear together. ## Failure two: the checkout dies before your script By the time a stop job runs, the branch that produced it may have been deleted — deleting the source branch on merge is the default in many projects, and it is one of the events that triggers the stop. The runner's default behaviour is to fetch the repository at the job's ref first. No ref, no checkout, and the job fails before line one of your teardown script. Setting `GIT_STRATEGY: none` on the stop job skips the fetch entirely. That is fine because a teardown script normally needs no source code — it needs a name, and `CI_ENVIRONMENT_SLUG` is still populated. If your teardown genuinely needs files from the repo, put them in an image or pass them as artifacts rather than relying on a ref that will not exist. ## Failure three: mismatched declaration The stop job must declare the *same* `environment:name` as the deploy job and set `environment:action: stop`. Getting the name expression subtly different — hardcoding `review/$CI_COMMIT_REF_NAME` in one and `review/$CI_COMMIT_REF_SLUG` in the other — produces two environments, one of which is never stopped. A manual stop job should also be allowed to fail rather than block the pipeline, since it deliberately does not run at deploy time. ## Failure four: expiry without teardown `auto_stop_in` on an environment whose `on_stop` is missing does exactly what it says and no more: GitLab marks the environment stopped. Your Kubernetes namespace, your load balancer, your cloud database and your DNS record are all still there and still billing. This is the version of the bug that hides longest, because the Environments page looks clean — the leak is only visible in the infrastructure. ## Belt and braces Because every one of these failures is silent, teams that run review apps at scale add a defence that does not depend on the pipeline: a scheduled job that lists the actual namespaces or DNS records, compares them against the environments GitLab believes are running, and deletes or reports the orphans. Label each resource with the environment slug and the merge request IID at creation time so that reconciliation is a simple set difference rather than a guess. The habit to take away is that in GitLab CI the create half and the destroy half of a review app are written **together, in the same commit** — an `environment:` block with `on_stop` and `auto_stop_in` filled in, and a stop job created under the same conditions as the deploy. Anything else leaks by default.
- Why does `GIT_STRATEGY: none` matter specifically on a stop job?The stop job frequently runs after its branch has been deleted — merging with "delete source branch" is one of the events that triggers it. The runner would otherwise try to fetch that ref and fail before the teardown script starts. `GIT_STRATEGY: none` skips the fetch; the job still has `CI_ENVIRONMENT_SLUG`, which is all a teardown usually needs.
- What does `auto_stop_in` do when a branch keeps receiving commits?Nothing visible — each new deployment to that environment restarts the countdown, so an actively developed branch keeps its review app for as long as work continues. The expiry only bites once deployments stop, which is exactly the abandoned-branch case you want cleaned up.
- How would you catch review-app resources that leaked despite all of this?Reconcile outside the pipeline. Label every created resource with the environment slug and merge request IID, then run a scheduled job that lists real namespaces, DNS records and databases and compares them with the environments GitLab reports as available. Delete or alert on anything with no matching live environment.
saying these in an interview costs you the question
- Assumes GitLab tears down review apps automatically
- Thinks a stopped environment means destroyed infrastructure
- Writes on_stop but never creates the stop job
- Lets the stop job attempt a checkout of a deleted branch
- Uses different environment name expressions in start and stop