skip to content

You are standardizing a platform where 40 teams want managed databases, caches and queues via Kubernetes Operators across a fleet of shared clusters. How do you decide the scoping, permissions and upgrade strategy for those operators?

level: principalimportance: nice to knowfreq 30%

answer

  1. CRDs cluster-scoped = fleet-wide API contract
  2. Curate a small catalog; each one is a permanent commitment
  3. Namespaced vs cluster-wide vs separate cluster by blast radius
  4. Bound tenants with RBAC + policy + quota on the CR spec
  5. Upgrades staged like an API migration; rehearse rollback

basics

~20 s

Curate a small vetted set of operators; decide per operator whether it runs namespaced or cluster-wide, grant the narrowest RBAC that works, and treat its CRDs as a cluster-wide singleton API needing versioned, staged upgrades. Monitor operators as tier-1 services and define who is on call for them.

solid answer

~50 s

Three decisions dominate. **Scope.** CRDs are cluster-scoped, so one operator version serves the whole cluster. Choose deliberately between a single cluster-wide instance (simple, one upgrade, one blast radius) and namespaced instances per tenant (isolation and independent versions, more overhead, only if the project supports it). Anything genuinely per-tenant and high-risk may deserve its own cluster. **Permissions.** Each operator is effectively privileged infrastructure. Grant per-resource-type, per-namespace rights; avoid blanket Secret access; run with a restricted security context; and gate the custom resources themselves with RBAC plus policy so teams cannot request 10 TB volumes or disable backups. **Upgrades.** Because the CRD is a shared API, upgrades are fleet events: version CRDs properly, stage through non-production clusters, upgrade the controller before adopting new spec fields, and rehearse rollback. Add an operator catalog with an owner, a support tier and an exit plan per entry. Above all, curate: five well-run operators beat twenty adopted ad hoc.

code

yaml · 18 lines
yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: postgrescluster-guardrails
spec:
  matchConstraints:
    resourceRules:
      - apiGroups: ["db.example.com"]
        apiVersions: ["v1"]
        operations: ["CREATE", "UPDATE"]
        resources: ["postgresclusters"]
  validations:
    - expression: "object.spec.storage.size <= '2Ti'"
      message: "storage above 2Ti requires platform approval"
    - expression: "has(object.spec.backup) && object.spec.backup.schedule != ''"
      message: "backups may not be disabled"
    - expression: "object.spec.replicas <= 7"
      message: "replica count above 7 requires platform approval"

go deeper

for a junior

Recognize that operators run with real permissions and are shared infrastructure, so they are not something a single team adopts unilaterally.

for a middle

Explain namespaced versus cluster-wide operation, that CRDs are cluster-scoped, and why the controller needs its own monitoring.

for a senior

Give an operational plan: scoped RBAC, guardrails on custom resource specs, staged upgrades with rollback rehearsal, and alerting on reconcile staleness and backup age.

for a principal

Frame it as platform strategy — a curated catalog with owners, support tiers and exit paths; isolation decisions driven by blast radius; and an explicit acceptance that operators trade toil for permanent maintenance and fleet-wide API coupling.

## Frame the problem At fleet scale the interesting questions are not "how does an operator work" but "what does it mean to run dozens of privileged control loops as shared platform infrastructure". Three properties make this different from adopting a library: - **CRDs are cluster-scoped singletons.** Two teams cannot run incompatible versions of the same operator's API in one cluster. The operator's API is therefore a *platform-level* contract, not a team-level choice. - **Operators are privileged.** They usually need write access to workloads, Secrets, PVCs, sometimes nodes. A compromised or buggy operator can affect every tenant it serves. - **They are in the availability path for day-2.** When the operator is down, failovers and backups quietly stop, even though applications keep serving. ## Curation before configuration Start by deciding *which* operators exist at all. Maintain a small catalog with, per entry: the business need it serves, maturity against the capability levels, project health and release cadence, demanded RBAC, whether it ships admission webhooks, the support tier, the named owning team, and the exit path (can we get the data out and run the software without the operator?). Rejecting operators is the highest-leverage decision; every one you accept becomes a permanent operational commitment. ## Scope: cluster-wide, namespaced, or separate cluster - **Single cluster-wide instance** — most common. One deployment, one version, one upgrade. Simple, but the blast radius is every tenant, and noisy-neighbour effects appear in the shared work queue: one team's 500 custom resources can starve reconciliation for everyone. - **Namespaced instances** — one controller per tenant or per group, watching only its namespaces. Better isolation and independent rollouts, but the CRDs are still shared cluster-wide, so the *API version* remains a common dependency, and you multiply the number of controllers to monitor. - **Separate clusters** — for genuinely different risk profiles (regulated data, a tenant that needs a different operator version, or an operator with cluster-admin-like requirements). Expensive but the only complete isolation. Decide per operator using blast radius and version-coupling, not uniformly. ## Permissions Treat the operator's ServiceAccount as a security boundary. Practical rules: enumerate resource types and verbs instead of wildcards; restrict to the namespaces it manages when supported; avoid `get`/`list` on all Secrets cluster-wide if a narrower scheme exists; run the controller with a hardened Pod security context and no host access; and separate the operator's namespace from tenant namespaces. Then govern the **custom resources** themselves, which is where tenants exert power: RBAC on who may create which kinds in which namespaces, plus policy (admission policy or a policy engine) and quota to bound what a spec may ask for — storage size, replica counts, instance classes, whether backups may be disabled, which storage classes and regions are allowed. A platform default that is safe and a validated envelope for deviations beats free-form YAML. ## Upgrades as fleet events Because the CRD is a shared API, plan upgrades like an API migration: 1. **Version CRDs properly** with a served/storage version strategy and conversion where the operator supports it; never rely on breaking schema changes. 2. **Stage rollouts**: development, then a canary production cluster, then the fleet, with soak time in between. 3. **Upgrade the controller before tenants adopt new spec fields**, so the API is available before it is used, and keep the previous version's objects reconciling. 4. **Rehearse rollback**, including what happens to in-flight operations and to objects written under the newer schema. 5. **Decouple the tenant's application version from the operator version** where possible, so upgrading the platform does not force 40 teams to move at once. 6. **Watch for webhook coupling**: an operator whose admission webhook is unavailable can block writes for its resources; enforce failure policy choices consciously and keep the webhook highly available. ## Run them like tier-1 services Monitor per operator: reconcile error rate, work-queue depth and latency, age of the last successful reconcile per resource, leader-election flapping, controller restarts, plus domain SLIs like backup age and replication lag. Alert on the operator being unable to act — the silent failure mode. Define on-call ownership: the platform team owns the operator, tenants own their custom resources, and there must be a documented, safe way to pause reconciliation during an incident without deleting anything. ## The judgment to voice The strategic point is that adopting operators converts operational toil into a **platform API surface with a permanent maintenance cost**. That trade is worth it when the automation is used by many teams and replaces frequent, risky manual work; it is a poor trade for one team's convenience. A principal-level answer names the criteria, admits the fleet-wide coupling that CRDs create, and sets a standard: few operators, explicit ownership, staged upgrades, bounded tenant power, and an exit path for each.

  • What breaks when two teams need different versions of the same operator in one cluster?
    The CRDs are cluster-scoped, so the API schema is shared even if you run two controllers. Different controller versions watching the same kinds will fight over the same objects unless the project supports namespaced scoping and non-overlapping watch namespaces. In practice you either align both teams on one version or separate them into different clusters.
  • How do you detect that an operator has silently stopped doing its job?
    Alert on the absence of work, not just on errors: age of the last successful reconcile per resource, backup age exceeding the schedule, growing work-queue depth or latency, leader-election churn and controller restarts. Applications keep serving while the operator is wedged, so without those signals the first symptom is a failed failover or a missing backup during an incident.
  • What is an exit path and why require one per operator?
    It is the documented ability to keep running the software if you stop using the operator — data in a standard format, credentials you control, manifests you could manage directly. Requiring it prevents a single-vendor controller from becoming an unremovable dependency, and it gives you an incident option when the operator itself is the problem.

saying these in an interview costs you the question

  • Treating operator adoption as a per-team decision when CRDs make the API cluster-wide
  • Granting cluster-admin or wildcard RBAC because 'the operator needs it' without reading what it actually uses
  • Assuming namespaced operator instances give full isolation, ignoring the shared CRD schema
  • Upgrading operators in place across the fleet with no staging, canary or rollback rehearsal
  • Monitoring only the managed application and not the controller's own reconcile health
  • Letting tenants specify unbounded storage, replicas or disabled backups in their custom resources

context