In a GitOps setup where an in-cluster agent syncs Kubernetes manifests from Git, how do you promote a version that is already running in staging to production?
answer
- the deploy button is a merge
- only one line should change
- the artifact already exists
- immutable reference beats moving tag
- undo is git revert
basics
~20 sPromotion is a Git commit, not a deploy job. You change the production manifests to reference the exact image the staging environment already runs, merge that pull request, and the cluster agent converges production onto it.
solid answer
~40 sIn GitOps there is no deploy button, so the only way to change production is to change the file the agent reads. The manifest repository normally holds shared manifests in a `base` plus one directory per environment carrying the differences — replicas, hostnames, and the image reference. Promotion is a commit that copies the **exact** image reference from the staging directory into the production directory, opened as a pull request so it is reviewed and recorded. Nothing is rebuilt: the artifact that passed staging is the one that ships, ideally pinned by digest rather than a mutable tag. Once merged, the agent picks up the new revision and applies the diff. Rollback is the same mechanism in reverse — `git revert` of the promotion commit.
code
yaml · 10 linesapiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ../../base
replicas:
- name: checkout
count: 6
images:
- name: registry.example.com/checkout
newTag: v1.4.2go deeper
Be able to say plainly that promotion is a commit or pull request against the production path of the manifest repository, and that the cluster agent applies it. Know that rollback is reverting that commit.
Explain the base plus per-environment directory layout and why only the image reference should differ in a promotion diff. Be ready to justify pinning by digest over a mutable tag.
Show the operational judgment: the promotion is not finished at merge, it is finished when the agent reports synced and healthy, and reverting a manifest does not undo migrations or lost state. Talk about how you detect a promotion that silently never synced.
Own the tradeoff between promotion friction and safety — which environments get an automatic bump, where a human approval genuinely adds signal, and what the repository layout has to guarantee so promotion stays a one-line, reviewable change across many services.
## What "no deploy button" means In a GitOps delivery model an agent running inside the cluster continuously compares a path in a Git repository against what is actually running, and converges the cluster onto whatever Git declares. Git is the source of desired state, and the agent is the only thing that writes to the cluster. That has one direct consequence for promotion: the only supported way to make production run a new version is to change the files the agent reads for production. There is no job you run, no button, and no `kubectl` you type — there is a commit. ## What the manifest repository actually holds A typical layout separates what is common from what differs per environment: - a **base** holding the manifests that are identical everywhere — the Deployment, Service, ConfigMap skeleton; - one **directory per environment** holding only the deltas: replica count, hostnames, resource limits, and above all the image reference. Kustomize expresses this as `base/` plus `overlays/staging/` and `overlays/production/`, each with its own `kustomization.yaml`. Helm expresses the same idea as one chart plus per-environment values files. Either way, the environment directory is the unit the agent is pointed at: an Argo CD `Application` names a repository, a `path` and a `targetRevision`; a Flux `Kustomization` names a `path` on a `GitRepository`. Pointing two agents at two directories in the same branch is what makes "staging" and "production" two views of one repository rather than two forks of it. ## The promotion commit The change itself is small and boring, and that is the point: ```yaml # overlays/production/kustomization.yaml images: - name: registry.example.com/checkout newTag: v1.4.2 # was v1.4.1 ``` The reviewer sees exactly one line: the version production is being asked to run. Everything else about production — its replica count, its hostname, its limits — stays where it was, because those values live in the production directory and were never touched. Crucially, **nothing is rebuilt**. Promotion moves a reference to an artifact that already exists and has already been tested. If the pipeline rebuilt the image at promotion time, you would get a different binary — different base-image layers, different transitive dependencies — and the thing you tested would not be the thing you shipped. ## Tags versus digests A tag such as `v1.4.2` is a mutable pointer: someone can push a different image to the same tag. A digest, `sha256:…`, names the image content itself and cannot be moved. Promoting a digest makes the Git commit an exact, verifiable statement of what production runs, and makes the history a real audit trail. Many teams write both — a tag so humans can read the diff, a digest so the runtime is unambiguous. ## How the change reaches the cluster The pull request merges, the agent notices the new revision (on its polling interval, or immediately if a webhook nudges it), and applies the difference. The rollout itself — surging new Pods, waiting for readiness — is the Deployment controller's job, not the agent's. Practically this means a promotion is *done* when the agent reports the target revision applied and the workload healthy, not when the merge button is clicked; treating the merge as the finish line is how teams end up announcing a release that never synced. ## Rollback Because the previous state is in Git history and the previous image is immutable in the registry, rolling back is `git revert` of the promotion commit — the same mechanism as rolling forward, with the same review trail. The important caveat is that reverting restores only what Git owns. A database migration that ran, a persistent volume that was deleted, or state the application wrote to an external system does not come back because the manifest went back. ## Failure modes worth naming - **Editing the cluster directly.** A hotfix applied with `kubectl` is not in Git, so it is either reverted or reported as drift, and it silently disappears at the next promotion. - **Promoting a floating tag.** If every environment references `latest` or `main`, promotion does nothing meaningful — each environment resolves to whatever the registry currently holds. - **Rebuilding at promotion time**, which destroys artifact identity. - **Copying the whole staging directory into production**, which drags staging's hostnames and replica counts along with the version bump.
- Why promote by image digest rather than by tag?A tag is a mutable pointer — someone can push a new image over `v1.4.2`, so the same manifest can resolve to different bits over time. A digest names the content itself, so the commit is an exact statement of what runs and the Git history becomes a genuine audit trail. Tags stay useful for human readability; the digest is what makes the promotion verifiable.
- What happens if someone hotfixes production with kubectl instead of committing?The change is invisible to Git, so the agent either reverts it on the next reconcile or flags production as out of sync. Even where it survives, the next promotion commit overwrites it, and nobody reviewing the repository can tell production is not what the manifests say. The correct escape hatch is a fast-tracked commit, not a manual edit.
- Does merging the promotion PR mean the release is finished?No. The merge only changes desired state; the agent still has to notice the revision and apply it, and the workload still has to roll out and become healthy. A release is complete when the agent reports the target revision synced and the workload healthy, so release automation should watch that signal rather than the merge event.
saying these in an interview costs you the question
- Says promotion means re-running the build for production
- Thinks you run kubectl apply to promote
- Promotes by moving a floating latest tag
- Copies the entire staging directory over production
- Calls the release done the moment the PR merges