You are designing GitOps for a platform with many tenant teams across a fleet of Kubernetes clusters. How do you decide between one shared GitOps control plane for everyone and a separate instance per tenant or per cluster?
answer
- scoping is soft, not hard, isolation
- source, destination, identity
- one version for everyone
- split multiplies upgrades and dashboards
- shard before you fragment
basics
~20 sDecide on isolation strength, upgrade coupling, and operational cost. A shared control plane gives one inventory and one thing to run but couples every tenant to one version and one failure domain; per-tenant instances give hard isolation at the price of N upgrades and no single view.
solid answer
~50 sStart from what a shared instance can and cannot enforce. Scoping inside one control plane — restricting which source repositories a tenant may deploy from, which clusters and namespaces they may deploy to, and giving each tenant its own service account and permissions — handles most organisations and gives you one inventory, one upgrade, one set of dashboards. Split when something that scoping cannot deliver is required: a tenant that must not share a failure domain or a controller version, a regulated workload where soft multi-tenancy is not acceptable, or a reconcile loop that no longer fits in one control plane at fleet scale. The middle path is usually the right answer — a shared or sharded control plane per environment tier, with per-tenant scoping inside it, and separate instances only for the handful of tenants whose requirements genuinely differ. Whatever you choose, keep one fleet inventory, because losing the single view is the real cost of splitting.
go deeper
Know that one GitOps control plane can serve many teams by restricting which repositories they deploy from and which clusters and namespaces they deploy to, and that some organisations run separate instances instead.
Explain the three scoping levers — source, destination, identity — and why they are soft multi-tenancy: enforced by configuration inside a process that still has broad reach across the fleet.
Weigh the operational reality: separate instances multiply upgrades, alerts and dashboards and destroy the single fleet view, so reach for sharding before fragmentation and split only for isolation or cadence reasons you can name.
Own the decision as trigger conditions rather than a preference — hard isolation requirements, divergent upgrade cadence, or a loop that no longer fits — and state what stays single regardless: the fleet inventory, the promotion path and the aggregated view.
## The question behind the question "One GitOps instance or many" is the same shape as every multi-tenancy question: how much isolation do you actually need, and what does soft isolation fail to give you? Answer it with axes, not with a preference. ## Axis 1 — what soft scoping can enforce A shared control plane is not undefended. Every serious GitOps tool can constrain a tenant's reconciled units along three lines: - **source**: which repositories or paths a tenant is allowed to deploy from, so team A cannot point a unit at team B's manifests; - **destination**: which clusters and namespaces a tenant may deploy into, so a misconfigured unit cannot land in production or in someone else's namespace; - **identity**: which service account and permissions the reconcile runs as, so the controller's power over a tenant's objects is bounded by ordinary cluster authorization rather than by one omnipotent identity. Add to that per-tenant limits on what kinds a tenant may create — cluster-scoped objects, CRDs, RBAC bindings — and most organisations are adequately served. The honest framing is that this is *soft* multi-tenancy: it is enforced by configuration inside a process that, taken as a whole, still has broad reach. ## Axis 2 — failure domain and credential concentration One control plane managing a fleet is one thing that can be broken. Its outage stops convergence everywhere, its misconfiguration can be fleet-wide, and it is the single most attractive target on the platform because it holds the means to change every cluster. That risk is inherent to centralisation and is only partly mitigated by scoping. Per-cluster agents shrink the blast radius to one cluster each; per-tenant instances shrink it to one tenant. Whether that matters depends on your threat model and on how bad a fleet-wide bad convergence would be — which is largely a question of whether pruning and self-heal are enabled. ## Axis 3 — upgrade coupling A shared instance means one controller version for everyone. That is an advantage (one upgrade, one CVE response, one behaviour to document) until a tenant needs to stay behind, or needs a feature only available in a newer version. Separate instances decouple the cadence and immediately multiply the work: N upgrades, N sets of alerts, N configurations to keep consistent, and drift between instances that nobody owns. Teams routinely underestimate this and end up with three instances on three different versions, which is worse than either extreme. ## Axis 4 — scale of the loop There is a practical ceiling. A single control plane reconciling thousands of units across many clusters spends real CPU and memory, holds large caches, and produces a lot of API traffic against every target. Before that ceiling, splitting for performance is premature. At it, the usual first move is not per-tenant instances but **sharding**: several controller replicas that split the set of managed clusters or units between them, keeping one logical platform while distributing the work. ## Axis 5 — the single view The cost people notice last is visibility. One control plane answers "what is deployed where, and is it healthy" in one place. Ten instances mean ten dashboards, ten sets of credentials for engineers, and no fleet-wide answer unless you build aggregation. If you split, budget for the inventory and the aggregated view up front — the fleet inventory list should remain single even when the control planes are not. ## A defensible default For most platforms: **one control plane per environment tier** (a production control plane, a non-production one), tenants scoped inside it by source, destination and identity, sharded when the loop outgrows one process, and dedicated instances only for tenants whose requirements are structurally different — hard isolation for regulatory reasons, a divergent version cadence, or an air-gapped or customer-managed cluster the platform team does not operate. Separating production from non-production is worth doing even when nothing else is separated: it is the split where a mistake is most expensive and where the two populations genuinely have different change rates and different reviewers. ## How to justify it in an interview Do not answer "shared" or "per-tenant". Answer with the trigger conditions: split when isolation must be enforced by something stronger than configuration, when a tenant needs its own upgrade cadence, or when the loop no longer fits — and stay shared otherwise, because every extra control plane is another thing to upgrade, monitor and reconcile with the others. Then say what you keep single regardless: the fleet inventory, the promotion path, and the view.
- What can a shared GitOps control plane realistically enforce between tenants?Three scopes: which source repositories or paths a tenant may deploy from, which clusters and namespaces it may deploy into, and which service account the reconcile runs as, bounded further by limits on the object kinds it may create. That covers accidents and most misconfiguration, but it is configuration-enforced soft multi-tenancy inside one broadly privileged process — not a hard boundary.
- A team asks for its own GitOps instance because reconciliation feels slow. Is that a good reason to split?Usually not. Slowness is a capacity problem first: check the interval, the number of managed units, cache and API pressure, and whether the controller can be sharded across replicas by cluster or unit set. Sharding keeps one logical platform and one inventory. Splitting into a separate instance to fix latency also buys you a separate upgrade, alerting and audit surface you will maintain forever.
- If you do split into several control planes, what should remain single?The fleet inventory and the view over it. The list of clusters and tenants should stay one artifact even when several control planes consume it, and health and version status should be aggregated somewhere. Losing the single answer to "what runs where" is the real cost of splitting, and it is the part teams discover months later during an incident.
saying these in an interview costs you the question
- Treats per-tenant scoping as a hard security boundary
- Gives every team its own instance without counting the upgrade cost
- Ignores that one control plane is one fleet-wide failure domain
- Splits for performance before trying to shard
- Keeps production and non-production in the same instance for convenience