skip to content

How can a Pod obtain short-lived cloud IAM credentials, for example to read an object-storage bucket, without any static cloud access key being stored in the cluster? Describe the trust chain that makes it work.

level: seniorimportance: must knowfreq 50%

answer

  1. cluster = OIDC issuer, publishes JWKS
  2. cloud trust policy conditions on iss + aud + sub
  3. sub = system:serviceaccount:<ns>:<name>
  4. projected token audience = sts.amazonaws.com (or equivalent)
  5. STS exchange -> short-lived creds; no stored key

basics

~20 s

Workload identity federation: the cluster publishes an OIDC discovery document and JWKS, the cloud IAM provider is configured to trust that issuer, the Pod gets a projected ServiceAccount token minted for the cloud's audience, and the cloud STS exchanges that token for short-lived credentials scoped to a role mapped from the ServiceAccount.

solid answer

~60 s

Use the projected ServiceAccount token as a federated identity assertion instead of storing a cloud key. The chain has four links. **One:** the API server is configured with a public `--service-account-issuer` URL and serves an OIDC discovery document plus a JWKS of its token-signing public keys. **Two:** in cloud IAM you register that issuer as an identity provider and write a trust policy on a role saying "tokens from this issuer whose `sub` equals `system:serviceaccount:<ns>:<name>` may assume me". **Three:** the Pod mounts a projected token with the cloud's audience (`sts.amazonaws.com`, or the provider's equivalent), typically injected automatically by a webhook when the ServiceAccount carries the right annotation. **Four:** the SDK calls STS with that token, STS validates the signature against the JWKS and checks issuer, audience, and subject, and returns short-lived credentials. Nothing long-lived exists anywhere: the projected token expires in about an hour and is bound to the Pod, and the STS credentials expire sooner. The mapping unit is the ServiceAccount, which is one more reason to have one per workload.

code

yaml · 31 lines
yaml
apiVersion: v1
kind: ServiceAccount
metadata:
  name: reports
  namespace: analytics
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::111122223333:role/reports-reader
---
# what the mutating webhook effectively adds to the Pod
spec:
  serviceAccountName: reports
  volumes:
    - name: aws-token
      projected:
        sources:
          - serviceAccountToken:
              path: token
              audience: sts.amazonaws.com
              expirationSeconds: 3600
  containers:
    - name: app
      image: reports:3.2
      env:
        - name: AWS_ROLE_ARN
          value: arn:aws:iam::111122223333:role/reports-reader
        - name: AWS_WEB_IDENTITY_TOKEN_FILE
          value: /var/run/secrets/aws/token
      volumeMounts:
        - name: aws-token
          mountPath: /var/run/secrets/aws
          readOnly: true

go deeper

for a junior

Know that the Pod's ServiceAccount token can be exchanged with the cloud for temporary credentials, so no cloud key needs to be stored in the cluster.

for a middle

Describe the four links of the chain and where the audience and role annotation appear in the manifests.

for a senior

Cover the trust-policy conditions in detail, the failure modes (missing sub condition, silent node-role fallback, token caching), and how revocation and audit attribution work.

for a principal

Set it as the platform standard: one ServiceAccount per workload mapped to one cloud role, trust policies pinned to issuer, audience and subject, node instance roles stripped to nothing, and a policy that no static cloud key is ever admitted into a cluster.

## The problem being removed The naive way to let a Pod reach a cloud API is to put a long-lived access key in a Kubernetes Secret. That key does not expire, is readable by anyone with Secret access or Pod-create in the namespace, is invisible to cloud-side audit as "which workload used it", and has to be rotated manually forever. The alternative on many platforms was to give the *node* an instance role, which is worse: every Pod on the node inherits it, so the permission set is the union of every workload's needs and attribution is impossible. Workload identity federation removes the stored credential entirely by reusing something the cluster already mints: a signed, short-lived, audience-scoped assertion of the Pod's ServiceAccount identity. ## The trust chain, link by link **1. The cluster becomes an OIDC issuer.** The API server is started with `--service-account-issuer=<https URL>` and signs ServiceAccount tokens with a key whose public half is published at `<issuer>/.well-known/openid-configuration` and the JWKS URL it names. The endpoint must be reachable by the cloud provider, which on managed clusters is handled for you (EKS gives each cluster an OIDC issuer URL; GKE and AKS expose equivalents), and on self-managed clusters is commonly done by publishing the two static documents to an object-storage bucket. **2. The cloud trusts the issuer.** You register the issuer as an OIDC identity provider in the cloud account, then attach a trust (assume-role) policy to a role conditioned on the token's claims: the issuer, the audience (`aud`), and the subject (`sub`), which for a Kubernetes ServiceAccount token is `system:serviceaccount:<namespace>:<name>`. Getting the condition right matters enormously: a policy that checks the issuer but not the subject lets *any* ServiceAccount in the cluster assume the role, and one that omits the audience check weakens replay protection. **3. The Pod receives a scoped token.** The Pod mounts a `projected` volume with a `serviceAccountToken` source whose `audience` is the cloud STS (`sts.amazonaws.com` for AWS; the provider-specific audience elsewhere) and a chosen `expirationSeconds`. In practice a mutating webhook does this automatically when the ServiceAccount carries the provider's annotation (`eks.amazonaws.com/role-arn`, `iam.gke.io/gcp-service-account`, or the Azure Workload Identity label plus client-id annotation), and the same webhook injects the environment variables the SDKs look for. The audience scoping is what makes it safe to hand this token to an external party: the Kubernetes API server rejects it, and the cloud rejects tokens minted for the API server. **4. The exchange.** The SDK's credential provider reads the token file and calls the cloud STS (`AssumeRoleWithWebIdentity` on AWS, the STS token exchange on GCP, and the equivalent on Azure). STS fetches the cluster's JWKS, verifies the signature, checks `iss`, `aud`, `exp`, and `sub` against the role's trust policy, and returns temporary credentials, typically valid for an hour and refreshed by the SDK automatically. The SDK also re-reads the projected token file, which is necessary because the kubelet rotates it. ## Properties worth naming - **No stored secret.** Every credential in the chain is short-lived and derived. There is nothing to rotate and nothing worth exfiltrating for long. - **Per-workload attribution.** Cloud audit logs record the assumed role and the federated subject, so a bucket read traces back to a specific ServiceAccount rather than to a node. - **Least privilege at the right granularity.** The unit of mapping is the ServiceAccount, so it composes with the one-ServiceAccount-per-workload rule; sharing an account across workloads collapses their cloud permissions together too. - **Revocation.** Deleting the Pod invalidates its projected token (it is bound to the Pod UID); removing the trust condition or the role stops the exchange for everyone immediately. ## Failure modes to recognise The common ones are: a trust policy that matches the issuer but not the subject, which is a cluster-wide privilege escalation waiting to happen; a missing or wrong `aud`, so STS rejects the token; a Pod that does not name the annotated ServiceAccount, so no token is injected and the SDK silently falls back to the node instance role, which usually *works* and hides the misconfiguration; an old SDK that reads the token file once and fails after rotation; and clock skew or an unreachable JWKS endpoint causing intermittent validation failures. AWS's newer EKS Pod Identity is a variant that replaces the per-cluster OIDC provider with an agent on the node and an association object, trading the standards-based federation for simpler setup. ## What to say Walk the four links (cluster as OIDC issuer, cloud trusts the issuer with issuer/audience/subject conditions, Pod gets an audience-scoped projected token, SDK exchanges it at STS for short-lived credentials), then state the payoff: no stored key, per-ServiceAccount attribution, and automatic expiry. Mention the subject-condition mistake, because it is the one that turns this from a hardening win into a cluster-wide escalation.

  • A role's trust policy checks the OIDC issuer and the audience but not the sub claim. What is the impact?
    Any ServiceAccount in that cluster can assume the role, because every token the cluster issues carries the same issuer and can request the same audience. A developer who can create a Pod in any namespace can therefore obtain the role's cloud permissions. The subject condition is what binds a cloud role to one specific ServiceAccount, and omitting it turns per-workload identity back into cluster-wide access.
  • How is this better than giving the node an instance role that Pods inherit?
    A node role is shared by every Pod scheduled there, so its permission set is the union of all their needs and the cloud audit log attributes actions to the node, not the workload. Workload identity scopes credentials to a single ServiceAccount, so permissions are per-workload, actions are attributable, and a compromised Pod cannot use permissions belonging to its neighbours. It also removes the incentive to over-grant the node role.
  • What breaks if the application's SDK caches the projected token file at startup?
    The kubelet rotates that file roughly every hour, so a cached copy eventually expires and the STS exchange starts failing, taking cloud API access down with it, often after the deployment has looked healthy for an hour. Current cloud SDKs re-read the web-identity token file on each refresh; old ones may not, which is why SDK version is part of the migration checklist.

The cluster acts like a passport office and the cloud like a border that has agreed to accept its passports: the Pod carries a passport stamped with the destination and an expiry, and the border issues a short-stay visa, so no permanent visa ever has to be stored anywhere.

saying these in an interview costs you the question

  • Storing a long-lived cloud access key in a Kubernetes Secret and calling it workload identity
  • Writing a trust policy that omits the sub condition, allowing any ServiceAccount in the cluster to assume the role
  • Believing the Kubernetes API server would accept the cloud-audience token, or that one token works for both
  • Falling back to a node instance role without noticing, because the SDK silently succeeds when no projected token is injected
  • Sharing one annotated ServiceAccount across unrelated workloads, which merges their cloud permissions

context