skip to content

For a production platform, would you keep application credentials in native Kubernetes Secret objects or source them from an external secret manager? Make the design argument.

level: principalimportance: nice to knowfreq 28%

answer

  1. native = delivery layer, not a secret manager
  2. gaps: versioning, rotation, revocation, cross-cluster, git
  3. SOPS/Sealed = git-safe only
  4. ESO syncs into etcd; CSI mounts past it
  5. short-lived workload identity beats managing static secrets

basics

~20 s

Native Secrets are a distribution mechanism with no versioning, rotation or cross-cluster source of truth. Keep them as the delivery layer, but make an external manager the source of truth for anything long-lived — and prefer short-lived workload identity so there is no static credential to manage at all.

solid answer

~1 min

I would not frame it as either/or. Native Secrets are how material reaches a Pod; the question is where the material comes from and how long it lives. **Native alone** is defensible for a small estate: no extra dependency, RBAC and audit already exist, encryption at rest covers the storage tier. Its gaps are real — no versioning or rotation, no lease/revocation, namespace-scoped blast radius, and nothing safe to commit to git. **Options above it**, in rising order of ambition: - *SOPS / Sealed Secrets* — encrypted in git, materialised as native Secrets. Cheap, fixes the git problem, does not fix rotation. - *External Secrets Operator* — an external manager is the source of truth and syncs into native Secrets. Rotation and central audit; material still lands in etcd. - *Secrets Store CSI driver* — mounts straight from the manager into the Pod's tmpfs, optionally syncing. Keeps values out of etcd; adds a start-time dependency on the manager. - *Workload identity / short-lived tokens* (IRSA, projected ServiceAccount tokens, SPIFFE, database dynamic credentials) — the strongest answer where the downstream supports it, because there is no long-lived secret to steal or rotate. Decide on rotation SLA, compliance, blast radius, and the failure mode when the external store is down.

code

yaml · 18 lines
yaml
apiVersion: external-secrets.io/v1beta1
kind: ExternalSecret
metadata:
  name: db-credentials
  namespace: team-a
spec:
  refreshInterval: 1h
  secretStoreRef:
    name: vault-backend
    kind: SecretStore
  target:
    name: db-credentials     # a normal Opaque Secret is created here
    creationPolicy: Owner
  data:
    - secretKey: password
      remoteRef:
        key: prod/checkout/db
        property: password

go deeper

for a junior

Know that native Secrets have no rotation or versioning and are unsafe to commit to git, and that external managers exist to fill those gaps.

for a middle

Name the concrete options — SOPS/Sealed Secrets, External Secrets Operator, Secrets Store CSI — and what each specifically fixes.

for a senior

Argue from failure modes and blast radius: what still lands in etcd, what happens when the manager is down at Pod start, and how rotation actually reaches running workloads.

for a principal

Lead with eliminating static credentials via workload identity, treat native Secrets as the delivery layer, and justify each added component against a named gap, a compliance requirement and the operational cost of running it.

## Reframe the question "Native or external" is a false binary, and saying so is most of the answer. Native Secrets are a **delivery** mechanism: namespaced storage plus a kubelet path that projects material into a Pod as files or environment variables. Almost every external option still ends with a native Secret or a tmpfs file. The real decisions are: where the **source of truth** lives, how **long-lived** the credential is, and what happens when the dependency chain breaks. ## What native Secrets genuinely give you - Namespaced objects with RBAC, admission control and audit already in place. - Optional encryption at rest via the API server's EncryptionConfiguration (KMS-backed on managed clusters). - Node scoping — the Node authorizer plus NodeRestriction limit a kubelet to Secrets used by its own Pods. - tmpfs delivery, so values are not written to node disk, with in-place refresh for volume mounts. - Zero additional components to run, upgrade and page on. For a single cluster with a few dozen credentials and a team that rotates by hand, that is a coherent, defensible position. Do not over-engineer past it without a reason you can name. ## What they do not give you - **No versioning.** A Secret has one current value; there is no history and no "pin to version 7". - **No rotation.** Nothing expires anything. Rotation is a human or a controller you write. - **No leases or revocation.** You cannot revoke a credential that has already been read. - **No cross-cluster source of truth.** Ten clusters means ten copies drifting apart. - **Nothing git-safe.** A Secret manifest in git is a plaintext credential in git. - **Coarse blast radius.** `list secrets` in a namespace exposes everything in it, and anyone who can create a Pod there can mount and print any of them. Each gap maps to a specific external capability, which is how you should justify adding one. ## The options, and what each actually buys **SOPS / Sealed Secrets.** Encrypt values so the manifest can live in git; a controller or a decryption step turns them back into native Secrets at apply time. Cheap, no runtime dependency at Pod start, and it solves exactly one problem — git safety. Rotation, versioning and revocation remain manual, and key management becomes your problem. **External Secrets Operator (or similar sync controllers).** Declare an `ExternalSecret` naming a remote key in Vault, AWS Secrets Manager, GCP Secret Manager or Azure Key Vault; the operator writes a native Secret and refreshes it on an interval. You get one central source of truth across clusters, rotation driven by the manager, and central audit of the value. The material still lands in etcd, so encryption at rest and RBAC still matter — this raises the floor, it does not remove the native Secret from the picture. Note the interaction with immutability: this controller writes in place, so its target Secrets cannot be immutable. **Secrets Store CSI driver.** Mounts secrets from the external manager directly into the Pod's tmpfs volume; nothing need touch etcd unless you enable the optional sync-to-Secret. Strongest confidentiality of the sync-based options, but it introduces a **Pod-start dependency** on the external manager: if it is unreachable during a large-scale reschedule, Pods do not start. That failure mode has to be an explicit, accepted risk with a mitigation (caching, regional endpoints, staged evacuation), not a surprise. **Workload identity and dynamic credentials.** IRSA/EKS Pod Identity, GKE Workload Identity, projected ServiceAccount tokens with an audience, SPIFFE/SPIRE, Vault dynamic database credentials. The Pod proves *who it is* with a short-lived, automatically rotated token and the downstream system issues an equally short-lived credential. This eliminates the static secret rather than managing it, so rotation, revocation and leakage-blast-radius largely stop being problems. Where the downstream supports it, this is the answer; the constraint is legacy systems that only accept a static password. ## How I would decide Ask five questions: 1. **Rotation SLA.** If credentials must rotate every 24 hours or on demand, native alone is out. 2. **Compliance.** HSM-backed keys, evidence of who read what, retention of rotation history — those force an external manager, and often a specific one. 3. **Estate shape.** Multiple clusters or regions sharing credentials makes a central source of truth worth its cost; one cluster rarely does. 4. **Blast radius tolerance.** Who can read Secrets in a namespace, and who can create Pods there? If those sets are wide, moving material out of etcd via CSI or to short-lived identity is worth real effort. 5. **Failure mode.** What happens on a mass reschedule when the manager is unavailable? Sync-based approaches degrade gracefully (the last synced Secret is still there); direct-mount approaches block Pod start. ## Where I land Default: **workload identity wherever the downstream supports it**, so most services carry no static credential at all. For the residue that needs static material — legacy databases, third-party API keys — an external manager as source of truth synced into native Secrets, with encryption at rest and tight RBAC underneath, and SOPS-style encryption for anything that must live in git. Reserve direct CSI mounting for the highest-sensitivity workloads where keeping the value out of etcd justifies the start-time dependency. And say the operational truth: every layer you add is a component to run, upgrade and be paged for. A well-run native-Secrets setup with encryption at rest and disciplined RBAC beats a poorly-run Vault integration whose sidecar is the top cause of failed deploys.

  • External Secrets Operator still writes the value into a native Secret in etcd. What did you actually gain?
    A single source of truth across clusters, rotation driven by the manager rather than by humans, central audit and revocation of the upstream value, and no credential in git. You did not gain confidentiality against someone with read access to Secrets in that namespace — encryption at rest and RBAC still carry that. If keeping the value out of etcd is the requirement, you need the CSI driver or short-lived identity instead.
  • What is the operational risk of mounting secrets directly from an external manager with the CSI driver?
    You add the manager to the Pod-start critical path. During a mass reschedule — node pool replacement, zone failure — every starting Pod calls the manager, and if it is unreachable or rate-limits you, Pods stay in ContainerCreating. Sync-based approaches degrade better because the last synced Secret is still in the cluster. Accept the risk deliberately with caching, regional endpoints and staged evacuation, or keep sync for tier-1 workloads.
  • When is plain native Secrets genuinely the right answer?
    A single cluster, a modest number of credentials, no regulatory requirement for HSM-backed keys or read auditing of values, and a rotation cadence the team can meet manually. With encryption at rest enabled and disciplined RBAC on `get`/`list`, that is a coherent posture — and it beats a half-maintained external integration that becomes the top cause of deploy failures.

Native Secrets are the delivery van; an external manager is the warehouse with an inventory system. Arguing van-versus-warehouse misses that you usually want both — and that the best shipment is the one you never had to store.

saying these in an interview costs you the question

  • Treating it as binary and ignoring that most external options still end in a native Secret.
  • Claiming an external manager removes the need for encryption at rest and tight RBAC.
  • Not mentioning short-lived workload identity, which removes the static credential entirely.
  • Ignoring the Pod-start dependency the CSI driver introduces during a mass reschedule.
  • Recommending Vault reflexively for a single small cluster without naming which gap it closes.

context