For a team adopting Pulumi, how would you decide between Pulumi Cloud and a self-managed state backend such as an S3 or GCS bucket, and what do you take on with the self-managed option?
answer
- it is a login target, not a code block
- checkpoint durability becomes yours
- no service key means an explicit secrets provider
- bucket read equals state read
- two pipelines, one stack, no referee
basics
~20 sPulumi Cloud is the managed backend: it holds each stack's checkpoint, manages a per-stack encryption key, and provides accounts, permissions and history. Self-managed means pulumi login s3://bucket, and you then own bucket durability, the secrets provider and key distribution, access control, and making sure two pipelines never update one stack at once.
solid answer
~50 sThe decision is about who operates the state, not about features of the language. `pulumi login` targets the managed service; `pulumi login s3://my-bucket`, `gs://`, `azblob://` or `file://~` targets a self-managed ("DIY") backend where checkpoints are just objects you own. Choose self-managed when policy says the state must stay in your own account, or the team is small enough that the surrounding features are noise. Choose the service when you want identity, per-stack permissions, deployment history and organization-wide guardrails without building them. With DIY you inherit: bucket versioning and backups, because the checkpoint is the record of what exists; an explicit secrets-provider decision, since there is no service key — a shared passphrase you must distribute, or better a cloud KMS key governed by IAM; access control collapsing to "who can read the bucket", which is also who can read the state; and serializing updates yourself so two CI runs never apply to one stack concurrently.
go deeper
Know that the backend is chosen by pulumi login — the managed Pulumi Cloud, or an S3/GCS/Azure bucket you own — and that the stack's checkpoint is stored there rather than in your repository.
Explain what changes on a self-managed backend: you supply the secrets provider yourself, bucket permissions become state permissions, and there is no built-in history or per-stack role. Mention versioning the bucket.
Show that you have operated one: versioning and backup on the state bucket, a KMS-backed secrets provider instead of a shared passphrase, CI concurrency so two runs never hit one stack, and knowing pulumi cancel plus a refresh after a killed update.
Own the decision and its cost. Be explicit that self-managed moves durability, key governance, access control and concurrency onto your team, name the requirement that would justify it, and state what you would fund on day one rather than after the first incident.
## What the backend actually is Pulumi stores a **checkpoint** per stack — the record of every resource it manages, with inputs, outputs and provider information. The backend is where that lives, and it is selected by login rather than by a block in code: ```bash pulumi login # Pulumi Cloud (the managed service) pulumi login s3://my-pulumi-state pulumi login gs://my-pulumi-state pulumi login azblob://my-container pulumi login file://~ # local filesystem, single-machine only ``` Everything else about the program is identical. This is a deployment/ops decision, not a code decision, and you can move a stack between backends by exporting and importing its checkpoint. ## What you are buying with the managed service - **Identity and permissions.** Accounts, organizations and per-stack access, so "who may update prod" is expressible without inventing it out of bucket policies. - **A managed encryption key per stack.** Secrets work with nothing extra to distribute; access to decrypt follows the same permissions as access to the stack. - **History and diffs.** A record of updates, who ran them and what changed — the audit answer to "who changed prod", which is usually why the organization adopted IaC in the first place. - **Concurrency control.** The service prevents two updates racing on one stack, which on a bucket you would have to arrange yourself. - **Organization-level guardrails** such as centrally applied policy — available as a service feature rather than something each pipeline opts into. ## What you take on going self-managed **Durability of the checkpoint.** The bucket is now the system of record for your infrastructure. Enable object versioning and back it up. Losing the checkpoint does not delete anything in the cloud — it deletes your *knowledge* of it, and recovery means importing resources back into a fresh stack one by one. **The secrets provider becomes an explicit decision.** With no service key, the stack needs `passphrase` or a cloud KMS URL. A passphrase is a shared secret every engineer and every CI job needs (`PULUMI_CONFIG_PASSPHRASE`), it cannot be revoked from one person, and losing it makes the stack's secrets unrecoverable. A KMS key turns decryption into an IAM grant you can audit and revoke — for a team, that is worth the extra setup. **Access control collapses.** Bucket read is state read is, effectively, seeing every resource property Pulumi recorded. There are no per-stack roles unless you build them out of prefixes and IAM policies, and "can deploy dev" versus "can deploy prod" has to be modelled with separate buckets or prefixes and separate CI identities. **Concurrency is yours.** Two pipelines running `pulumi up` on the same stack at once is the failure you must design out — pipeline-level concurrency groups, an environment lock in your CI system, or simply never having two automated paths to one stack. Also learn `pulumi cancel`, which clears a stack's pending-operation record after a run is killed mid-update; expect to reconcile with a `pulumi refresh` afterwards, since a killed update may have created resources the checkpoint does not know about. **Everything else you would have got.** History, review UI, notifications and organization policy either go unbuilt or get reimplemented as pipeline logs. ## How to actually decide Ask three questions in order: 1. **Is there a hard requirement that state never leaves our accounts?** If yes, self-managed with a KMS secrets provider, and budget for the operational work above. This is the honest, common reason. 2. **How many people and pipelines touch prod?** One or two engineers with one pipeline barely notice the missing features. Twenty engineers across five teams will rebuild permissions and audit badly. 3. **What is the cost of losing the checkpoint or leaking it?** That number sets how much you invest in versioning, backup, and key governance either way. ## The answer that lands Say plainly that self-managed is not free — it moves durability, key management, access control and concurrency from a vendor's problem to your backlog — and that the deciding factor is usually a data-residency or vendor policy rather than a technical preference. Then say which one you would pick for the team in front of you, and name the first thing you would set up: bucket versioning and a KMS-backed secrets provider on day one, not after the first incident.
- The S3 object holding a stack's checkpoint is deleted. What now?Restore it from bucket versioning or your backup — this is exactly why versioning is mandatory on a self-managed backend. Without a copy, nothing in the cloud is lost but Pulumi no longer knows about any of it, so recovery means creating a fresh stack and importing each existing resource back under management, which is slow and error-prone at estate scale.
- How does secret handling differ once you move off Pulumi Cloud?There is no service-managed per-stack key, so every stack needs an explicit secrets provider: `passphrase`, with `PULUMI_CONFIG_PASSPHRASE` distributed to every engineer and CI job, or a cloud KMS URL. Prefer KMS — the right to decrypt becomes an IAM grant you can audit and revoke per identity, instead of a shared string nobody can un-share.
- How would you stop two CI runs updating the same stack at once on a self-managed backend?Design it out at the pipeline level: a concurrency group or environment lock in the CI system so only one job per stack runs at a time, and only one automated path to each stack. Also know `pulumi cancel` for clearing a stack's pending-operation record after a run is killed, followed by a refresh to reconcile whatever that run actually created.
saying these in an interview costs you the question
- Says self-managed is free because a bucket is cheap
- Assumes the service and a bucket differ only in cost
- Forgets bucket versioning until the checkpoint is lost
- Ships a shared passphrase to the whole team and calls it access control
- Believes losing the checkpoint deletes the real infrastructure