skip to content

A team keeps long-lived staging and production branches in their Kubernetes manifest repository and promotes by merging staging into production. What goes wrong with that layout, and what is usually preferred?

level: seniorimportance: should knowfreq 52%

answer

  1. environments differ on purpose
  2. merges fight the intended differences
  3. merge takes everything, not one change
  4. hotfix never flows back
  5. directories make the delta visible

basics

~20 s

Environment branches diverge because each carries its own environment-specific values, so every promotion merge conflicts on the same files and drags unrelated changes along. The common alternative is one branch with a directory per environment, where promotion is an explicit edit to the production directory.

solid answer

~50 s

The trouble is that a branch is a poor container for configuration that is *meant* to differ. Production's replica counts, hostnames and limits live in the same files as staging's, so every merge conflicts on exactly those lines, and engineers learn to resolve conflicts by hand — which is how environments silently drift. Merging also promotes everything unmerged, not the one change you intended, so teams fall back on cherry-picking and lose any honest answer to "what is in production". Hotfixes committed straight to the production branch never flow back, so staging stops representing production. The usual layout instead is a single branch with a shared base plus `overlays/staging` and `overlays/production`; each agent tracks its own path, promotion is a one-line edit to the production overlay, and the diff between environments is visible at any moment without comparing branches.

go deeper

for a junior

Recognize that keeping each environment in its own directory on one branch is the common layout, and that promotion there means editing the production directory rather than merging a branch.

for a middle

Explain concretely why merges conflict — the environment differences live in the same files — and why a merge promotes the whole backlog rather than the single change you intended.

for a senior

Bring the operational evidence: hotfixes that never flow back, hand-resolved conflicts changing production settings, and no reliable answer to what production runs. Describe the migration to overlays and how you keep the delta enforced.

for a principal

Own the estate-level tradeoff: a shared base makes the blast radius global, tenant isolation needs separate repositories rather than branches, and the replacement promotion flow has to be as obvious as "merge to deploy" or the team will route around it.

## Why the branch model is attractive first Branch-per-environment feels natural because engineers already think in branches, and "merge staging into production" sounds exactly like promotion. Delivery agents support it directly — Argo CD's `Application` has a `targetRevision`, Flux's `GitRepository` has `ref.branch` — so pointing production at a `production` branch is a two-minute setup. The failure is not immediate; it appears after a few months of divergence. ## Failure one: the merge conflicts are structural, not accidental Environments differ on purpose. Production runs more replicas, different hostnames, different resource limits, different feature flags. In the branch model those differences are *diffs between branches on the same files*. Every promotion merge therefore touches the very lines that are supposed to differ, and Git cannot tell an intentional environment difference from a change awaiting promotion. Engineers resolve these conflicts by hand, under time pressure, and a single mistaken resolution silently changes production's replica count or hostname. Over time the conflict resolution *is* the deployment process, which is not a process anyone reviews. ## Failure two: merge promotes everything, not one thing A merge takes the whole delta. If staging carries three merged changes and you only trust one, merging ships all three. The workaround is cherry-picking, and once a team cherry-picks routinely the branches no longer share history: `git log production..staging` stops meaning "not yet promoted" because commits exist in both with different SHAs. At that point nobody can answer what is in production without diffing rendered manifests. ## Failure three: hotfixes only flow one way An urgent fix committed directly to the production branch is, by construction, absent from staging. Unless someone remembers to back-merge, staging is now behind production on that file, and the next promotion merge will happily revert the hotfix. This is the classic release-branch back-merge problem, imported into infrastructure where the consequence is an outage rather than a bad build. ## Failure four: you cannot see the estate With one branch per environment, comparing environments means checking out two refs and rendering both. With one branch and one directory per environment, `diff -r overlays/staging overlays/production` answers it instantly, code review shows the environment delta inline, and a policy check can *enforce* that the delta stays within an allowlist. ## The preferred layout One branch — `main` — containing: ``` base/ # identical everywhere overlays/dev/ overlays/staging/ overlays/production/ ``` Each environment's agent is pointed at its own path on the same revision: an Argo CD `Application` per environment with a different `path`, or a Flux `Kustomization` per environment with a different `path`. Shared truth lives in `base` and changes everywhere at once when it should; per-environment values live in the overlay that owns them and never move; promotion is an explicit edit of the production overlay's image reference. The reviewer of a promotion sees one line, and the reviewer of a base change knows it hits every environment. ## The honest caveats This layout is not free. A change to `base` reaches every environment on the next reconcile, so a broken base is an estate-wide event — teams mitigate with a staged rollout of the base change itself, or by having environments pin a released version of the base. It also means everyone can read every environment's configuration, which matters where tenants must be isolated; that is a case for separate repositories, not for separate branches. And branches are not universally wrong. A short-lived branch per preview environment is fine precisely because it never lives long enough to diverge. The anti-pattern is specifically the *long-lived* environment branch that accumulates its own configuration. ## Migrating away Do it once, deliberately: render each branch's manifests, extract everything common into `base`, keep each branch's remaining differences as its overlay, repoint each agent at its new path, and delete the branches so nobody commits to them again. The hardest part is usually social — the team's muscle memory says "merge to deploy", and the replacement must be equally obvious, which is what an automated promotion pull request provides.

  • Aren't there cases where a branch per environment is legitimate?
    Short-lived branches are fine — a preview environment per pull request never lives long enough to diverge. The anti-pattern is the long-lived branch that accumulates its own configuration. Genuine isolation requirements — a tenant who must not read another tenant's manifests — are solved with separate repositories and separate credentials, not with branches in a shared repository.
  • With one shared base, a bad base change hits every environment at once. How do you contain that?
    Treat the base as a released component rather than a live one: environments reference a tagged or versioned base and adopt it in the same staged order as an application version, so the base change is itself promoted. Alternatively keep the base deliberately thin so most changes land in a single overlay, and add a policy check that flags base edits for wider review.
  • How do you stop the environment overlays from quietly drifting apart even in the directory layout?
    Make the delta explicit and testable: render each overlay in CI and diff them, then fail the build if they differ on anything outside an allowlist such as image, replicas, hostnames and limits. That turns "production has a setting staging never had" from an incident discovery into a pull-request comment.

saying these in an interview costs you the question

  • Says merge conflicts just mean sloppy engineers
  • Claims cherry-picking is a normal promotion mechanism
  • Believes branch-per-environment prevents drift
  • Cannot explain how a production hotfix returns to staging
  • Thinks rendering and diffing environments requires the agent

context