skip to content

Your fleet's GitOps definitions for 50 clusters are generated from one template plus an inventory list rather than 50 hand-written files. What does that generation buy you, and what failure mode does it introduce?

level: seniorimportance: should knowfreq 40%

answer

  1. template plus inventory equals N units
  2. onboarding is a list entry
  3. the reviewed diff hides the expansion
  4. no natural staging window
  5. canary ring before the fleet

basics

~20 s

Generation removes copy-paste drift and makes onboarding a cluster a one-line commit against an inventory that becomes the source of truth. The cost is blast radius: one template edit changes all 50 clusters at once, and the reviewed diff shows one line rather than fifty rendered manifests.

solid answer

~50 s

Fan-out generation means the fleet's definitions are a function: template plus inventory equals N reconciled units. That buys real things — clusters cannot drift apart because there is only one template, adding or retiring a cluster is an entry in a list, and the inventory becomes a queryable record of what exists. The failure mode is that the review no longer shows what ships. A one-line template change expands into fifty changed clusters, and because reconciliation is continuous there is no natural staging window: every cluster picks it up within an interval. So fleets that do this add back the staging that generation removed — render the manifests in CI and review the rendered diff, carry a ring or wave field in the inventory so a canary cluster and then a tranche take the change first, and keep an explicit exclusion path so one cluster can hold a version back without leaving the pattern.

code

yaml · 19 lines
yaml
# inventory.yaml — the source of truth a fan-out generator expands.
# One entry per cluster; ring drives progressive rollout of template changes.
clusters:
  - name: eu-prod-1
    region: eu-west-1
    ring: canary
    tier: production
    replicas: 3
  - name: eu-prod-2
    region: eu-west-1
    ring: main
    tier: production
    replicas: 6
  - name: ap-prod-1
    region: ap-southeast-1
    ring: late
    tier: production
    replicas: 6
    pinnedRevision: v2.3.1   # sanctioned exception, still in the inventory

go deeper

for a junior

Know that a fleet is described once as a template plus a list of clusters, and that adding a cluster means adding a list entry instead of copying a folder.

for a middle

Explain both directions of the trade: generation removes copy drift and makes onboarding trivial, but the reviewed diff no longer shows how many clusters change or what the manifests will actually say.

for a senior

Describe the controls you would add — render the manifests in CI and review the rendered diff, carry a rollout ring in the inventory, guard against a shrinking expansion — and explain why continuous reconciliation leaves no natural staging window.

for a principal

Own the fleet rollout policy: how many rings exist, what soak each gets, who may grant an exception, and when a cluster's difference is structural enough to deserve its own template rather than another parameter.

## Generation as a function Once the same platform runs on more than a handful of clusters, hand-written per-cluster definitions stop working: fifty near-identical files diverge, because someone fixes a bug in the one they were looking at. The alternative is to treat the fleet as a function — one template, one inventory of clusters or tenants, and a generator that expands them into N reconciled units. ```yaml # fleet inventory: one entry per cluster, consumed by the generator clusters: - name: eu-prod-1 region: eu-west-1 ring: canary replicas: 3 - name: eu-prod-2 region: eu-west-1 ring: main replicas: 6 - name: us-prod-1 region: us-east-1 ring: main replicas: 6 ``` What comes out is not different in kind from a hand-written definition; the controller cannot tell. The change is in who wrote it and what a human reviews. ## What it genuinely buys **No copy drift.** There is exactly one description of what a cluster runs. Fixing it once fixes it everywhere, which is the whole point. **Cheap onboarding and offboarding.** A new cluster is a list entry; a retired cluster is a deleted entry, and — with pruning — its managed objects go with it. Cluster count stops being a cost per change. **The inventory becomes the record.** The list is the answer to "which clusters exist, in which region, at which tier". Other things hang off the same list: dashboards, alert routing, cost attribution, upgrade tracking. Where the list is populated from a live source rather than by hand, the fleet definition follows cluster registration automatically. **Tenant fan-out is the same trick.** Substitute tenants for clusters and the mechanics are identical: one template, one list of tenants, N isolated sets of resources. ## The failure mode: the diff lies Here is the sharp edge. In the hand-written world, a change touching fifty clusters was a fifty-file diff — ugly, but the reviewer could see the size of what they were approving. With generation, that same change is one line in a template, and the reviewer sees one line. The expansion happens after the merge, inside a controller. Continuous reconciliation makes it worse than in a push-based pipeline. There is no deploy step to stagger, no "we will roll it out to the second region tomorrow": every cluster is polling, and within one interval all fifty have converged on the new template. A subtly wrong value — a bad image tag, a resource limit that makes pods unschedulable, a removed key — reaches production everywhere before anyone reads the alert. ## Putting the staging back Mature fleets deliberately re-introduce what the abstraction removed: **Render in CI and review the rendered output.** The pull request shows the expanded manifests, or at least the diff for a representative cluster, so the reviewer approves what will actually run. This is the single highest-value control, and it also catches template errors — an empty expansion, a lost entry — before they reach a controller that would happily converge to them. **Rings or waves in the inventory.** Carry a field like `ring: canary | main | late` and let the generator resolve which template revision each ring uses. The canary cluster picks up the change, you watch it, then the rest follow. This turns "fifty at once" back into a progressive rollout without abandoning generation. **A version per ring, not per cluster.** Pinning individual clusters produces fifty snowflakes again. Pinning a small number of rings keeps the fleet describable. **A sanctioned exception path.** There will be a cluster that must hold a version back — a regulated region, a customer-managed cluster, a broken node pool. Support that as an inventory field, not as a hand-edited copy of the generated output, or the copy becomes permanent. **Guardrails on expansion size.** A check that fails the pipeline when the rendered set shrinks unexpectedly protects against the empty-expansion accident, which under pruning is a fleet-wide deletion rather than a fleet-wide bad deploy. ## Where over-parameterisation shows up The other slow failure is a template that grows a flag for every per-cluster difference until it is unreadable and no combination is tested. When a cluster's needs genuinely differ in structure rather than in values, a second template is usually cheaper than a fifteenth conditional. "One template, many values" is the goal; "one template, many branches" is how it degrades. ## The interview answer Generation converts fifty maintenance problems into one, and converts one review problem into fifty. You keep the first trade and pay down the second with rendered-diff review and rings.

  • How would you stop a single template change from reaching every cluster at once?
    Carry a rollout ring in the inventory and resolve the template revision per ring, so a canary cluster takes the change first, then a tranche, then the rest. Pin rings rather than individual clusters, or you recreate snowflakes. Pair it with a soak period long enough for the failure you actually fear — pod scheduling shows up in minutes, a slow leak does not.
  • What check protects against a generator that renders an empty or truncated set?
    Render in CI and compare the expanded output against the previous render, failing the pipeline when the number of generated units drops unexpectedly. Without it, an empty expansion is valid desired state, and a controller with pruning enabled converges the fleet toward nothing. Conservative deletion behaviour on the generated units is a useful second layer.
  • When should a cluster get its own template instead of another parameter?
    When its difference is structural rather than a value — different components, a different topology, a different dependency layer. Values belong in the inventory; a template accumulating conditionals for each special case becomes untested and unreadable. Two clear templates beat one template with fifteen branches whose combinations nobody exercises.

saying these in an interview costs you the question

  • Assumes a one-line template diff is a one-cluster change
  • Reviews the template but never the rendered manifests
  • Pins template versions per cluster and recreates snowflakes
  • Adds a boolean parameter for every per-cluster difference
  • Believes continuous reconciliation gives a natural rollout window

context