skip to content

You are breaking a large infrastructure estate into separately applied components. What criteria decide where the boundaries go?

level: seniorimportance: should knowfreq 55%

answer

  1. cut along how things change
  2. cadence, ownership, consequence, run time
  3. a plan can only destroy what is inside
  4. every seam is an interface plus manual ordering
  5. dependencies flow one way, never cycle

basics

~20 s

Cut along change cadence, ownership and consequence of failure: things that change together and are owned by the same people belong in one unit, and the seam goes where you want a mistake to stop. Keep cross-boundary dependencies few and pointing one way.

solid answer

~50 s

I use four criteria. **Change cadence** — networking and accounts change a few times a year, services change daily; putting them in one unit means the daily change re-reads and risks the yearly one. **Blast radius** — the seam belongs where a bad change should stop, which usually means stateful things (databases, buckets, DNS) live in a unit no routine deploy ever touches. **Ownership and approval** — if two groups need different reviewers, that is a boundary, because approval granularity cannot be finer than the apply unit. **Run time and contention** — a unit that takes fifteen minutes and holds a lock blocks everyone behind it. Then I check the cost side: every boundary becomes an interface plus a cross-unit dependency someone must order by hand, so I prefer a few well-placed seams over many, and split further only when a real pain shows up.

go deeper

for a junior

Know that a big estate is not applied as one unit, and that pieces are grouped so a change touches only what it needs to. Being able to name environment and layer as two different kinds of split is enough here.

for a middle

Explain the layering — foundation, platform, services — and why change cadence drives it. Be able to say that a plan can only propose changes to what is inside its own unit, and what that implies for stateful resources.

for a senior

Demonstrate judgement: give your criteria, apply them to a concrete estate, and be explicit about the cost of each seam you add. Expect to be pushed on how a downstream unit consumes upstream values and what happens when that contract changes.

for a principal

Own the organisational half. Decide the layering that matches team boundaries and approval authority, set the rule for who may create a new unit, and weigh the operational cost of many records, pipelines and credentials against the containment they buy.

## Why the seam matters more than the size Asked to split an estate, most people reach for a taxonomy — one unit per service, or one per resource type. Both produce boundaries that cut *across* how the estate actually changes, and you feel it immediately: a routine deploy has to touch three units in order, and a network change has to be coordinated with twelve teams. The useful criteria are behavioural, not categorical. ## Criterion 1: change cadence Group things that change at the same rate. A landing-zone layer — accounts, DNS zones, core network, organisational guardrails — changes a handful of times per year and is nearly always reviewed by a specialist. A platform layer — clusters, shared databases, message brokers — changes monthly. Application infrastructure changes several times a day. Merging a daily-changing thing with a yearly-changing thing means every deploy re-reads and re-diffs the stable layer, exposes it to every mistake, and makes the fast layer as slow and as scary as the slow one. ## Criterion 2: blast radius, meaning what a mistake can reach A plan can only propose destroying what is inside the unit it runs against. That single sentence is the strongest architectural lever in IaC. So put the irreplaceable things — the primary database, the object store holding customer data, the DNS zone the whole company resolves through — in a unit that routine changes never run against. Then the worst plausible outcome of a bad service deploy is a broken service, not a deleted database. This is separate from, and stronger than, any deletion-protection flag on the individual resource, because it removes the resource from consideration entirely rather than adding a guard someone can override. ## Criterion 3: ownership and approval granularity **Approval cannot be finer-grained than the apply unit.** If the network team must sign off on network changes and the product team owns its own services, and both live in one unit, then either every change waits for both reviewers or one group is rubber-stamping the other's work. Making the seam match the ownership line is how you get a review process people actually follow. ## Criterion 4: run time and lock contention Every unit serialises on its own record. A unit spanning nine hundred resources takes many minutes to refresh and diff, and while it runs, everyone else who needs it waits. Teams feel this long before they feel anything architectural: the queue at 5pm on a Friday is the symptom that a unit is too coarse. ## The cost of a seam — cutting is not free Every boundary you add creates two obligations. First, **an interface**. The downstream unit needs values from the upstream one — subnet identifiers, cluster endpoints, role names. However that value travels (reading the upstream unit's published outputs, looking the resource up by a stable tag or name, or a parameter store), it is now a contract. Renaming an upstream output becomes a breaking change with consumers you must find. Second, **manual ordering**. Inside one unit the tool derives the dependency order for you. Across units, a human or a pipeline knows that network applies before platform applies before services. Creating a brand-new environment from scratch becomes a choreographed sequence rather than one command, and that sequence is the part that rots when it is only exercised once a year. So the seams should be few and deliberate. A good default estate has three or four layers, not thirty: ``` foundation accounts, DNS, core network (yearly, specialist review) platform clusters, shared data stores (monthly) services per-service infrastructure (daily, team-owned) ``` ## Keep dependencies one-directional Dependencies should flow strictly downward: services read from platform, platform reads from foundation, and nothing reads upward. A cycle between two units cannot be resolved by the tool the way an in-unit cycle can — you resolve it by hand, applying half of one unit, then the other, then coming back. If you find yourself wanting a cycle, the boundary is in the wrong place and the two pieces probably want to be one unit. ## Splitting an existing monolith Do it once, deliberately, along the highest-value seam — usually pulling the stateful and foundational resources out of the everyday unit. Moving resources between units means moving their entries between recorded states without recreating anything, which is the delicate part and worth rehearsing in a copy of the environment first. Because it is delicate, split for a reason you can name, not preemptively. ## What interviewers listen for That you have criteria at all, that blast radius and ownership appear among them, and that you can articulate the cost of over-splitting. Someone who answers "one unit per microservice" without mentioning what happens to shared networking has not run this at scale.

  • Where does a shared production database belong in that layering, and why?
    In a slow-moving unit that routine deploys never run against, together with the other irreplaceable resources. The goal is that no everyday change even has the database in scope, so it can never appear in a destroy or replace action. Deletion guards on the resource help, but exclusion from the apply unit is the stronger control.
  • What is the practical cost of splitting an estate into too many units?
    Ordering and interfaces. The tool no longer derives the dependency order for you, so creating or rebuilding an environment becomes a hand-maintained sequence that rots between uses. Each cross-unit value becomes a contract you cannot rename freely, and the number of pipelines, records and credentials to operate multiplies.
  • Two units end up needing values from each other. How do you resolve it?
    Treat the cycle as evidence the boundary is wrong and merge them, or break it by introducing a stable, independently-created identifier both sides can reference — a fixed name or tag — rather than a live output. Manually applying half of each unit in turn works once and becomes an operational trap you rediscover during an incident.

saying these in an interview costs you the question

  • Split by resource type: one unit for all databases, one for all networks.
  • One unit per microservice, with shared networking left unassigned.
  • Smaller units are always better, split as far as possible.
  • Cross-unit dependencies are free because the tool orders them.
  • Blast radius is handled by deletion protection, so boundaries do not matter.

context