skip to content

You run a Kubernetes developer platform: how do you draw the line between platform-owned and workload-owned configuration, and which escape hatches do you allow?

level: principalimportance: should knowfreq 35%

answer

  1. ownership follows the pager
  2. harms neighbours means platform bounds
  3. admission is law, template is advice
  4. graded hatches, lowest that works
  5. hatch usage feeds the roadmap

basics

~20 s

Draw the line at a versioned contract: the platform owns the cluster, shared services and non-negotiable guardrails; teams own their images, sizing, scaling and service behaviour. Escape hatches are graded, stay under admission policy, and every use is tracked.

solid answer

~50 s

I draw the line by asking who gets paged and who can safely change something. The platform owns the cluster, nodes, networking, shared gateways, observability plumbing, namespace provisioning and the guardrails that protect other tenants — enforced at admission, not only in templates. Teams own the image, configuration, resource sizing, replica and scaling choices, probes, and alerts on their own service levels. The golden-path template is the versioned interface between the two. Escape hatches come in grades: extra template parameters, a validated patch or overlay field, then fully custom manifests in the team's namespace — all still subject to the same admission rules — and finally a time-boxed exception with an owner. I track how often each hatch is used: frequent use of one hatch means the paved road is missing a feature, and that becomes platform roadmap work, not a support queue.

go deeper

for a junior

Recall the basic split: the platform team runs the cluster and shared services, and your team owns its own service's configuration and behaviour.

for a middle

Explain which configuration a team sets through template values and which rules the cluster enforces regardless of how manifests are written.

for a senior

Show how you would operate the split: admission policy for guardrails, graded escape hatches, and template upgrades that do not strand teams.

for a principal

Argue the tradeoffs — thick versus thin abstraction, mandate versus earned adoption — and name the metrics that tell you when to move the line.

## Why the line matters When one platform team runs the clusters and many product teams run workloads on them, every field in a manifest has an implicit owner. If ownership is unclear, two failure modes follow: the platform team becomes a ticket queue for every change, or product teams change things that break their neighbours. The job is to make ownership **explicit, enforceable and cheap to change**. ## A test for where each concern belongs For each concern, ask: 1. **Who gets paged when it is wrong?** Ownership follows the pager. 2. **Does a wrong value hurt other tenants?** If yes, the platform sets the bounds. 3. **Does the team need to change it weekly?** If yes, the team must be able to change it without a ticket. ## A typical split | Concern | Owner | Mechanism | |---|---|---| | Cluster version, nodes, CNI, shared gateways | Platform | Platform repositories, no tenant access | | Namespace creation and its standard bundle | Platform | Provisioning controller | | Pod security, required labels, image sources | Platform | Admission policy on every namespace | | Container image, configuration, feature flags | Team | Template values | | CPU and memory sizing, replicas, scaling | Team, within bounds | Template values plus namespace limits | | Probes, graceful shutdown behaviour | Team | Template values with platform defaults | | Service-level alerts and on-call | Team | Shared alerting with team routing | Take a **fraud-rules engine**: the team decides it needs 7 replicas and a large memory request for its rule cache; the platform decides that it must run non-root, carry cost labels, and pull only from approved registries. Neither side needs the other's permission for its own half. ## Guardrails belong at admission A template is advice; the API server's admission chain is law. Anything whose violation harms other tenants should be enforced where every create and update passes — an admission policy — and not only as a template default. That single choice makes escape hatches safe: a team that leaves the template still cannot run a privileged pod. ## Grading the escape hatches - **Grade 1 — parameters.** The template exposes a documented value (tolerations, extra environment variables, a relaxed spread rule). - **Grade 2 — structured patches.** The template accepts a validated patch or overlay for fields it does not model. - **Grade 3 — own manifests.** The team writes raw objects in its own namespace, still under admission policy and the namespace's limits. - **Grade 4 — exceptions.** A policy exemption with a named owner, a reason and an expiry date, reviewed when it lapses. Each grade gives more freedom and costs more support. Teams should reach for the lowest grade that works. ## Tradeoffs to own - **Thick versus thin abstraction.** A thick custom resource hides Kubernetes almost entirely and leaks badly when it fails; a thin template leaves developers reading Kubernetes objects daily. Most platforms land on a thin template with strong defaults. - **Mandate versus earn.** Mandating the path fills the escape-hatch queue with resentment; earning adoption by being the easiest route produces real feedback. - **Template versioning.** Teams pin versions; old versions accumulate. You need an upgrade story — automated pull requests, deprecation windows, a supported-version window. - **Organisational fit.** The split mirrors the team structure; if the platform team is small, keep the platform-owned surface small. ## Measuring whether the line is right - share of workloads on the current template version - number of escape hatches used, by grade and by reason - time from repository creation to first production deploy - tickets to the platform team per team per month A reason that recurs across hatches is a missing feature. Promote it into the template, and the line moves — deliberately, with a changelog, rather than by drift.

  • A team keeps using raw manifests because the template cannot express an extra sidecar their fraud-rules engine needs. What do you do?
    Treat it as a product signal. Check whether other teams need the same thing; if so, add a supported parameter to the template and migrate the team back. If it is genuinely unique, leave it on the raw-manifest grade, confirm admission policy still covers it, and record the reason so the next review can revisit it.
  • How do you roll a breaking template change out across two hundred services without a flag day?
    Version the template and publish the change as a new major version with a changelog. Support the previous version for a stated window, open automated pull requests that bump teams onto the new version with the rendered diff attached, and track adoption. Anything that must change everywhere at once belongs in admission policy, not the template.
  • When would you choose a thick custom-resource abstraction over a thin template?
    When the workload shape is narrow and repeated — many near-identical services — and the platform team can afford to own debugging for it. The payoff is a small, stable interface. The cost is that every failure surfaces as a problem in objects developers never see, so you need strong status reporting and a way to inspect the generated objects.

saying these in an interview costs you the question

  • The platform team should approve every workload change to stay safe.
  • If the template is good enough, no escape hatch is needed.
  • Guardrails set only as template defaults are enough to protect other tenants.
  • Teams that leave the golden path should lose platform support entirely.
  • The ownership split is fixed once and never needs revisiting.