skip to content

In a Terraform pipeline, what cloud credentials should the plan job hold compared with the apply job, and why is a single long-lived access key stored in CI a poor choice for both?

level: middleimportance: should knowfreq 48%

answer

  1. read for the preview, write only after merge
  2. plan still writes the lock
  3. no secret at rest, no rotation
  4. the token names repo and branch
  5. scope the trust, not just the policy

basics

~20 s

The plan job should get read-only access to the managed resources; only the apply job needs permission to create, change and destroy. Both should receive short-lived credentials issued per run through federation, not one static key sitting in CI forever.

solid answer

~50 s

Plan and apply have genuinely different needs. Plan reads the configuration and queries the provider APIs to refresh what exists, so it can run under a read-only role — with one wrinkle: it still writes the backend's lock, so the plan role needs write access to the lock even though it changes nothing in the estate. Apply is the only job that needs create, update and delete permissions, and it runs only from the protected branch, so the powerful role is reachable from far fewer code paths. A single static access key collapses that separation: every run, including one triggered from a branch, holds full write power. It also never expires, so a leak in a log or a compromised dependency is exploitable until someone notices and rotates it. Short-lived credentials issued per run through OIDC federation expire in minutes, leave no secret at rest, and can be scoped so only runs from the right repository and branch can assume the apply role.

go deeper

for a junior

Know that the plan job only needs to read while the apply job needs permission to change things, and that credentials should not be long-lived keys pasted into CI.

for a middle

Explain the split precisely, including the exception that plan still writes the state lock, and describe how per-run federated credentials remove the stored secret and the rotation problem.

for a senior

Bring the operational judgment: scoping the apply role per environment, session duration versus long applies, and the honest limits of a read-only plan role when a plan can read sensitive data sources.

for a principal

Own the trust model across the estate — which claims each role's trust condition pins, how environments are isolated from one another, and how you evidence that no long-lived cloud key exists in CI anywhere.

## Two jobs, two permission sets The pipeline's two halves are asymmetric and their credentials should be too. **The plan job** reads. It parses the configuration, downloads providers, reads the current state from the backend, and calls provider APIs to check what the recorded resources look like now. It creates nothing. A read-only role over the relevant services is enough — plus two things people forget: - **Write access to the backend's lock.** Terraform locks state before planning, so a role with read-only access to the backend fails immediately with a lock error. Whatever your backend uses for locking must be writable by the plan role. - **Read access to the state object itself**, which is a different permission from reading the resources it describes. **The apply job** writes. It needs create, update and delete on everything the configuration manages, plus write access to the state. This is the powerful role, and its value is that exactly one code path can reach it: the job that runs after a merge to the protected branch. Splitting them buys containment. A pull request from any branch — including, in some setups, an untrusted contributor's fork — can trigger a plan. If the plan job's credentials could write, that trigger is a path to production, because a plan job runs code from the branch: the configuration itself, and anything the pipeline evaluates. With a read-only plan role, the worst outcome is disclosure, which is serious but bounded. Be honest about that bound in an interview: a read-only role can still read a lot. Plan output describes your infrastructure, and a plan can be made to read a data source that surfaces something sensitive. Read-only is a reduction in blast radius, not a guarantee of harmlessness. ## Why one static key is the wrong shape A single long-lived key stored as a CI secret has four properties you do not want: 1. **No separation.** One key means the plan job and the apply job have identical power, so the containment above does not exist. Every plan run is a potential write. 2. **No expiry.** A key that never expires is exploitable from the moment it leaks until someone notices — which, for a credential nobody rotates, can be years. Leaks are mundane: a `terraform apply` run with debug logging on, a crash dump, an error message echoed into a public build log. 3. **A secret at rest.** It exists in the CI secret store, in whatever spreadsheet or vault it was copied from, and on the laptop of whoever created it. Every copy is an exposure. 4. **Painful rotation.** Because rotating means editing every pipeline that references it, rotation slips, which is why these keys are old. ## What replaces it OIDC federation. The CI system issues each job a short-lived, signed identity token describing the run — the repository, the branch or environment, the workflow. The cloud provider is configured to trust that issuer and to exchange the token for temporary credentials against a specific role, with conditions on the token's claims. Nothing is stored: the token is minted for the job, the credentials it yields expire in minutes, and there is no secret to leak, rotate or exfiltrate. Terraform picks the resulting credentials up through the provider's ordinary environment-based authentication, so the configuration itself needs no change. The conditions on the trust relationship are what make the split enforceable rather than merely conventional. The apply role's trust policy names the specific repository and restricts it to runs from the protected branch or a protected deployment environment; the plan role trusts a broader set of runs but grants only reads. A job on a feature branch that asks for the apply role is refused by the cloud provider, not by a convention in the pipeline file. ## Practical points - **One role per environment, not one for the estate.** The pipeline that applies staging should not be able to touch production. - **The apply role's permissions should still be scoped.** "Administrator" is easy and wrong; scope to the services the configuration actually manages, accepting that Terraform's need to create IAM constructs makes this genuinely hard and worth deliberate design. - **Match run length to credential lifetime.** A long apply against a large estate can outlive a short session, so the session duration must cover the slowest realistic run. - **Do not let the plan role read secret stores** unless the configuration genuinely needs it; that is the most common way a read-only plan job becomes a disclosure incident.

  • Why can't the plan job's role be strictly read-only on everything?
    Because Terraform locks the state before planning, so the plan role needs write access to whatever the backend uses for locking, and read access to the state object itself — a separate permission from reading the resources state describes. Everything else it touches genuinely can be read-only. Making the lock write the sole exception is what keeps the separation meaningful.
  • How does the trust configuration stop a feature branch from assuming the apply role?
    The identity token the CI system mints carries claims about the run — repository, and branch or deployment environment. The role's trust relationship is written to accept only tokens whose claims match the protected branch or environment. A job elsewhere presents a token with different claims and the exchange is refused by the cloud provider, so the restriction is enforced outside any file a contributor can edit.
  • A large production apply sometimes fails partway through with an authentication error. What would you check?
    Whether the temporary credentials expired mid-run. Federated sessions have a maximum duration, and an apply against a big estate — or one waiting on slow resources like a database or a certificate — can outlive a short one. Raise the session duration to cover the slowest realistic run, and treat a very long apply as a signal that the state should be split.

saying these in an interview costs you the question

  • One CI key for plan and apply is simpler and fine
  • Plan needs no cloud permissions at all
  • A read-only plan role makes fork pull requests completely safe
  • Static keys are safe because the secret store encrypts them
  • Give the apply role administrator access and move on

context