skip to content

Environment Separation

Putting production in an account of its own rather than behind a name prefix, a label or a namespace, and paying to build the setup twice. Probed because a mis-scoped script respects no convention.

on this pageshow

questions

4

Your staging ledger shares production's account, kept apart only by a name prefix and an environment label — why is that not an isolation boundary?

level: juniorimportance: must knowfreq 72%

answer

  1. convention against enforcement
  2. who reads the prefix at call time?
  3. isolation, quota and bill attach to the account
  4. one account is one permission surface
  5. a mis-scoped call reads no prefix

basics

~20 s

A name prefix and an environment label are data every caller in the account can read or ignore; nothing enforces them. Isolation, quota and billing attach to the account, so only a separate account makes the split real.

solid answer

~50 s

A prefix and a label are **conventions**: data attached to resources that anything holding credentials in the account can read, and that the platform never consults before authorizing a call. One account is one isolation boundary, one quota pool and one bill, so the staging ledger and the live ledger share the same permission surface, the same service limits and the same blast radius. A job run with account-wide credentials that matches one character too few, or meets a resource created before the labelling rule existed, reaches both. Moving production into an account of its own makes the split something the platform holds: a credential issued for the non-production account cannot name a production resource unless someone deliberately grants that. You pay for it with the setup and the access built twice, which is the honest trade.

go deeper

for a junior

Be able to say that a prefix and a label are data, not a control: nothing in the platform checks them before it runs a call. The account is the thing that is enforced.

for a middle

Explain what an account is the unit of — isolation, quota, billing and the controls applied above it — and why sharing one means staging and production share all four.

for a senior

Show the failure you have lived: automation holding account-wide credentials whose filter was the only thing narrowing its reach, and what you changed so the credential could not name the other environment.

for a principal

Argue the price. Two accounts mean the setup and the access built twice and every later change made twice; say which environments earn an account of their own and which can share one.

## Convention against boundary Separating environments with a **name prefix** (`stg-ledger-api`) or an **environment label** (`environment=staging`) is a convention: a rule that lives in a document, in code review, and in the habits of whoever writes the scripts. The platform stores the prefix as part of a resource's name and the label as a key-value pair hanging off it, and it will return both to anyone who asks. What it will not do is consult either one before it authorizes a call. A convention is therefore exactly as strong as the least careful caller that ever holds credentials in that account — and that set includes automation nobody remembers writing. An **account** is a different kind of thing, because it is the unit the platform itself uses: - **Isolation** — a credential is issued for one account, and a call made with it can name only resources in that account. - **Quota** — service limits are counted per account, so one environment's burst draws down the pool the other needs. - **Billing** — charges roll up per account, so the non-production number exists without anyone maintaining tags. - **Policy** — the controls an organisation applies attach to an account, or to the grouping above accounts; nothing attaches to a prefix. Put staging and the live ledger in one account and they share all four. The prefix *describes* the split; the account is where the split would have to be *enforced*, and it is not. ## What each one actually buys | | name prefix or label | production in its own account | |---|---|---| | Enforced by | people, habit and review | the platform, on every call | | A mis-scoped call | matches whatever it matches | cannot name the other environment | | Quota | one shared pool | one pool per environment | | Bill | attributable only while tagging holds | separate by construction | | A leaked credential | reaches both environments | reaches one | | Cost to adopt | almost nothing | the setup and the access, built twice | ## The failure the split is bought for The scenario is never dramatic. It goes like this: 1. Something runs with credentials valid for the whole account — a nightly cleanup job, an interactive session, a one-off migration. 2. Its filter is the only thing narrowing what it touches: a prefix match, a label match, a wildcard. 3. The filter is wrong for a reason nobody predicted — resources created before the convention existed carry no label, a pattern matches one character short, a condition is inverted, a command is pasted without its last argument. 4. The call succeeds, because every resource it named sat inside the account the credential was issued for. Nothing in that chain is malicious, and none of it reads your prefix. That is the whole argument: a convention cannot fail closed, because it was never in the path of the decision. Moving production into its own account changes step 4 — the credential the job holds cannot name a production resource whatever the filter says, so a filter bug becomes a non-production incident instead of a production one. ## What it costs, honestly The split is not free, and an answer that claims it is has not done it: - The network layout, the baseline settings and the logging and monitoring wiring exist twice, and every later change to them is made twice. - Access is granted again rather than copied, because the grants should be broad in non-production and narrow in production. - Some services carry a fixed charge per account, so a second account has a floor even when little runs in it. - Engineers switch context, and a change that used to be one action becomes two. That price is why the split people actually defend is **production against everything else**, rather than an account per environment per team. ## What the boundary does not do - It does not design the permissions *inside* production; an over-broad grant within the production account is untouched by the split. - It does not stop a mistake made by production's own automation against production. - It does not separate a person who legitimately holds access in both, which is why access review survives the split. - It does not make labels useless. Inside each account they still answer which team, which service and which cost centre — they simply stop being the isolation story.

  • If the two environments genuinely have to stay in one account for now, what is the strongest separation still available, and what stays shared?
    Issue a distinct credential per environment and scope each one as narrowly as the platform allows, so at least the automation cannot name the other side by accident. What you cannot split is structural: one quota pool, one bill to attribute by tagging, one audit trail, and administrators of the account who can widen any scope. Treat it as harm reduction with a dated plan to split, not as an equivalent.
  • When you do split, do you move production out or move the other environments out?
    Usually move the non-production environments, because the migration is the risky part and you would rather rehearse it where downtime is cheap. Either way it is a rebuild, not a transfer: on most platforms live resources cannot be relocated between accounts, so you recreate the setup in the new account and migrate data into it. Whichever account keeps production must be the one you are willing to lock down.

A label on a shelf tells staff which boxes hold live stock; a separate locked room means the courier who holds a key to one cannot walk into the other by mistake. The label informs; the door decides.

saying these in an interview costs you the question

  • Says a naming convention is enough as long as the whole team follows it.
  • Believes a label restricts which resources a credential can act on.
  • Treats a separate private network inside the same account as the same isolation.
  • Claims the split removes the need to design permissions inside production.
  • Assumes code review catches the mis-scoped job that runs unattended every night.
open as a page

What do you actually pay twice when production moves into an account of its own, and what stays single?

level: middleimportance: should knowfreq 54%

basics

~20 s

Paid twice: the network and baseline setup, the logging and monitoring wiring, every access grant, some per-account fixed charges, and every later change to all of it. Still single: the built artifact, the repository, the pipeline definition and the team.

open as a page

A nightly cleanup job meant for staging deleted live ledger data because both environments sit in one account — what allowed it, and which fix closes the route?

level: seniorimportance: should knowfreq 46%

basics

~20 s

One account held both environments, so the job's filter rather than the platform decided what it could delete, and resources predating the label escaped that filter. Only a credential that cannot name production closes the route; an alert merely reports it.

open as a page

Your artifact registry and log destination serve both environments after the account split — where should each live, and what keeps the split real?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A service both environments genuinely share belongs in neither: put it in a third account with every grant pointing inward — environments pull artifacts and append logs they cannot delete. A grant back out rejoins what you split.

open as a page