skip to content

A platform team runs 612 External Secrets Operator ExternalSecrets for many teams on one 48-node Kubernetes cluster: how would you design store scoping, refresh cadence and outage behaviour?

level: principalimportance: nice to knowfreq 26%

answer

  1. store scope is the auth boundary
  2. per-team identity, provider enforces
  3. objects times frequency equals bill
  4. tier intervals, force planned syncs
  5. outage keeps value, stops rotation

basics

~20 s

Give each team a namespaced SecretStore with its own provider identity, and keep ClusterSecretStores few and condition-limited. Tier refreshInterval by credential class within the provider call budget. Alert on failed syncs, because an outage keeps values but stops rotation.

solid answer

~50 s

I would treat this as a blast-radius and budget design. **Scoping:** each team gets a namespaced `SecretStore` authenticating with its own identity, so a mistaken `remoteRef.key` fails at the provider instead of leaking another team's data. A few shared values go through a `ClusterSecretStore` whose `spec.conditions` restrict the namespaces, and admission policy keeps teams from pointing at stores they do not own. **Cadence:** refresh is a provider call, so 612 objects at `1m` is about 10 reads per second, while at the `1h` default it is about 10 per minute. I would tier it: short intervals for credentials on a rotation schedule, the default for the rest, and forced syncs for planned rotations. **Outage:** a failed fetch keeps the last Secret, so Pods keep running, but rotation stops. I would alert on `Ready=False`, watch the controller's reconcile backlog (`--concurrent` defaults to 1), and rehearse a provider outage during something like a node-by-node upgrade.

code

yaml · 18 lines
yaml
apiVersion: external-secrets.io/v1
kind: ClusterSecretStore
metadata:
  name: shared-registry
spec:
  conditions:
    - namespaceSelector:
        matchLabels:
          platform.example.com/registry-access: "true"
  provider:
    vault:
      server: https://vault.internal.example.com:8200
      path: shared
      version: v2
      auth:
        kubernetes:
          mountPath: kubernetes
          role: eso-shared-registry

go deeper

for a junior

Remember that a SecretStore belongs to one namespace, while a ClusterSecretStore can be used from any namespace unless restricted.

for a middle

Be able to compute provider load from object count and refresh interval, and explain what happens to Secrets when a sync fails.

for a senior

Show operational controls: alerts on failed syncs, controller concurrency and queue depth, and a rehearsed provider outage during node maintenance.

for a principal

Own the trade between tenant isolation, freshness, provider cost and resilience. State where you land and which constraint, such as etcd exposure or per-call billing, would change it.

## Framing the decision Once a cluster hosts hundreds of `ExternalSecret` objects owned by many teams, the External Secrets Operator stops being a convenience and becomes **shared infrastructure** sitting between every workload and the secret manager. The design has no single right answer. It trades **isolation**, **freshness**, **provider cost and rate limits** and **resilience** against each other. The scenario here is 612 ExternalSecrets across about 40 team namespaces on a 48-node cluster that is also mid-upgrade, with nodes draining one by one. ## 1. Store scoping: who can read what The store holds the credential the operator uses against the provider, so **store scope is the authorisation boundary**. | Model | Isolation | Operational cost | Main risk | |---|---|---|---| | One `ClusterSecretStore` with a platform-wide identity | Weak. Any namespace can name any key | Lowest | One typo or malicious `remoteRef.key` reads another team's secret | | `ClusterSecretStore` limited with `spec.conditions` | Medium. Only the listed or selected namespaces may use it | Low | Label drift on namespaces widens access | | A namespaced `SecretStore` per team, with a per-team provider identity | Strong. The provider enforces the path boundary | Higher, with one identity per team | Store sprawl and identity lifecycle | The usual landing point: - A **per-team `SecretStore`**, authenticated with a cloud IAM role bound to that namespace's ServiceAccount, or an equivalent identity, so the **provider** refuses cross-team reads. - A small number of **`ClusterSecretStore`s** for genuinely shared data (for example a shared registry credential), each restricted with `conditions` using `namespaceSelector` or an explicit `namespaces` list. - **Admission policy** that rejects an ExternalSecret whose `secretStoreRef.kind` is `ClusterSecretStore` unless the namespace is allowed. The policy engine is a neighbouring system, and the rule is what matters here. - RBAC that lets teams create ExternalSecrets but not SecretStores, if the platform issues the identities. ## 2. Refresh cadence: freshness against budget Every refresh is at least one provider request, and more with `dataFrom` or multiple `data` entries. For 612 objects with one read each: | `refreshInterval` | Reads per minute | Reads per second | |---|---|---| | `1h` (default) | 10.2 | 0.17 | | `15m` | 40.8 | 0.68 | | `1m` | 612 | 10.2 | Per-request pricing and provider rate limits make the `1m` row expensive, and a burst after a controller restart makes it worse. The recommendation is to **tier by credential class**: 1. **Rotated on a schedule** (database passwords, API tokens): an interval sized so the staleness bound fits inside the overlap window the credential owner grants, often `5m` to `15m`. 2. **Static configuration-like secrets**: the `1h` default, or `refreshPolicy: OnChange` if the value is only changed together with the manifest. 3. **Planned rotations**: a **forced sync** by changing the ExternalSecret, rather than a permanently short interval. ## 3. Controller capacity - The operator's `--concurrent` flag defaults to **1** reconcile at a time. With 612 objects and a slow provider, the work queue can grow faster than it drains after a restart or a mass change, so raise concurrency deliberately and measure the queue depth. - The controller also talks to the Kubernetes API about every Secret it manages. Keep its client limits in view on a busy control plane, especially during an upgrade. - Run it with more than one replica and leader election for availability. Only the leader reconciles, so replicas add resilience, not throughput. ## 4. Outage behaviour When the provider is unreachable, the controller marks each affected ExternalSecret `Ready=False` with reason `SecretSyncedError` and **leaves the existing Secret untouched**. - **Good:** Pods rescheduled during the node-by-node upgrade still start, because the Secrets already exist in the cluster. That is a structural advantage over a mount-time CSI fetch. - **Bad:** rotation silently stops. If a credential is revoked upstream during the outage, Pods keep an invalid value. - **Controls:** - alert on the count of `Ready=False` ExternalSecrets per namespace, not per object; - forbid `deletionPolicy: Delete` for critical secrets, so a provider-side mistake does not delete live Secrets; - rehearse: block the provider in a staging cluster and confirm that Pods still roll during an upgrade. ## 5. Where I would land, and what would change it - Per-team namespaced stores with per-team identities, few condition-limited cluster stores, tiered intervals, raised concurrency with alerting, and the default `Retain` deletion policy. - **Change the answer** if teams must keep secrets out of etcd (move those workloads to the CSI driver's file-only path), if the provider bills per call at a level that dominates cost (lengthen intervals and rely on forced syncs), or if a single team's volume dominates (give it a dedicated operator instance scoped to its namespaces).

  • A team asks for refreshInterval 30s on all 70 of its ExternalSecrets. How do you respond?
    Ask what staleness bound they actually need. 70 objects at 30s is 140 reads a minute from one team, which is about 2.3 per second. If the goal is fast rotation, a forced sync at rotation time gives the same freshness for a fraction of the calls. I would set a platform floor on the interval and document the forced-sync path.
  • When would you run more than one External Secrets Operator instance?
    When one tenant's volume or provider latency starves the shared reconcile queue, when a tenant needs a different provider trust boundary, or when compliance demands separate controller credentials. Scope each instance to its namespaces so they never reconcile the same objects.

saying these in an interview costs you the question

  • One cluster-wide store with an admin identity is simplest and fine
  • Refresh frequency has no cost, so set everything to seconds
  • Adding operator replicas multiplies reconcile throughput
  • A provider outage deletes synced Secrets, so Pods fail fast
  • Namespace RBAC alone stops cross-team reads through a shared store