skip to content

A retailer plans a 5-node edge Kubernetes cluster in each of 400 stores. Would you use provider-managed control planes or run your own, and how would you decide?

level: principalimportance: should knowfreq 34%

answer

  1. four axes, not one
  2. fee per cluster times 400
  3. WAN link joins the control plane
  4. local control plane eats node memory
  5. fleet automation or nothing

basics

~20 s

Weigh the per-cluster fee times 400, the store's dependence on a WAN link to a regional control plane, and the memory a local control plane takes from five small nodes. Many fleets self-run a lightweight distribution with strong fleet automation instead.

solid answer

~50 s

I would decide on four axes. **Cost**: managed control planes are usually billed per cluster, so 400 small clusters multiply a fee that is negligible for one large cluster. **Connectivity**: a managed control plane runs in the provider's region, so each store depends on its WAN link; during an outage kubelets keep running existing Pods, but nothing can be scheduled, rescheduled or reconfigured. **Local capacity**: a self-run control plane takes memory and CPU from the five nodes that should run the inference server. **Operations**: self-running 400 control planes means owning patching, certificates, etcd and upgrades for all of them, which is only viable with fleet automation. For stores that must keep working offline, I would usually self-run a lightweight distribution with a single declarative fleet pipeline, and keep managed control planes for the central regional clusters.

go deeper

for a junior

Recall that a managed control plane runs in the provider's region and that providers usually charge a fee for every cluster.

for a middle

Explain what still works on a node that loses contact with the API server and what stops, such as scheduling and configuration changes.

for a senior

Size the store nodes against the real workload with and without a local control plane, and describe the eviction behaviour during partial outages.

for a principal

Decide on cost, connectivity, local capacity and operations together, and commit to the fleet automation that self-running at scale demands.

## The decision in one sentence A managed control plane trades configuration freedom and a per-cluster fee for someone else running the API server and etcd. At the edge, two more things matter: **where** that control plane runs and **how many** of them you need. ## Axis 1: cost scales with cluster count Managed services commonly charge a flat fee per cluster, independent of its size. That is negligible for a few large clusters and significant for a fleet of small ones. An illustrative calculation with a hypothetical fee of $0.10 per cluster-hour: - One cluster: 0.10 x 730 hours = **$73 per month**. - 400 store clusters: 73 x 400 = **$29,200 per month**, before any node cost. The same 400 stores would pay the fee once if they were nodes in one central cluster, which is why the choice of **cluster count** and the choice of **managed versus self-run** cannot be made separately. ## Axis 2: the WAN link becomes part of the control plane A managed control plane runs in the provider's region. Each store's kubelets talk to it over the store's uplink. When that link fails: 1. **Running Pods keep running.** The kubelet does not stop containers because it lost the API server. 2. **Nothing new happens.** No scheduling, no replacement of a crashed node's Pods on the other four nodes, no ConfigMap or Deployment changes, no `kubectl` from the store. 3. **Nodes go `Unknown`.** Their Lease objects stop renewing, and after the node lifecycle controller's grace period (`--node-monitor-grace-period`, 50 seconds by default) their `Ready` condition becomes `Unknown`. 4. **Eviction is dampened, not guaranteed.** Upstream `kube-controller-manager` stops evictions when *every* node in the cluster is not ready, treating it as a control-plane connectivity problem. A partial outage, for example two of five nodes on a flaky switch, can still lead to Pods being deleted after the default 300-second `tolerationSeconds` for the `node.kubernetes.io/unreachable` taint. A provider may also tune these settings, and you cannot see or change them. 5. **Recovery depends on reconciliation.** When the link returns, the kubelets resync with whatever the API server now says. If a store must sell while offline, a remote control plane is a design risk, not an implementation detail. ## Axis 3: a local control plane costs local capacity Self-running puts the control plane on the store's own five nodes. That has a price: | Choice | Survives a node loss | Local cost | |---|---|---| | One server node with an embedded datastore | No control plane until repaired | Lowest | | Three control-plane nodes with an etcd quorum | Yes, one node | Highest; three nodes carry control-plane load | | Remote managed control plane | Not applicable | None locally, WAN dependency instead | An illustrative sizing: each node has about 6.1 GiB allocatable, and each replica of the recommendation-model inference server has a **2.6 GiB memory limit** (and an equal request). Without a local control plane, two replicas fit per node: **10 replicas**. If control-plane components take about 1.2 GiB on three of the nodes, those nodes have 4.9 GiB left and fit only one replica each: 2 x 2 + 3 x 1 = **7 replicas**. That is a 30% capacity loss for local autonomy. Lightweight distributions such as k3s exist to shrink that footprint, which is why they are common at the edge. ## Axis 4: who operates 400 control planes Self-running moves every day-2 task back to you, multiplied by 400: - patch releases and minor upgrades for every store; - control-plane certificate renewal; - datastore backups and recovery; - hardware failures handled by store staff with no Kubernetes knowledge. This is only viable with **fleet automation**: identical, declaratively built clusters, a management layer such as Cluster API or a vendor fleet manager, and a GitOps controller delivering workloads. Without that, managed control planes are often cheaper in practice, even at $29,200 per month. ## Putting it together | Requirement | Leans towards | |---|---| | Store must operate fully offline | Self-run, local control plane | | Reliable uplink, small team, few clusters | Managed | | Hundreds of clusters, strong platform team | Self-run with fleet automation | | Tight node memory | Remote control plane or a single-server local one | | Strict audit or custom API server flags | Self-run | A common answer is a **hybrid**: managed control planes for central regional clusters that run shared services, and self-run lightweight clusters in stores, all built from one template and upgraded by one pipeline. Some providers also offer managed control planes that extend to on-premises hardware; evaluate those on the same four axes rather than on the label. ## What a strong answer shows It asks for the offline requirement first, does the fee arithmetic, sizes the nodes against the real workload and admits that self-running 400 control planes is a platform-team commitment, not a configuration choice.

  • Why does upstream Kubernetes stop evicting Pods when every node goes not ready at once?
    The node lifecycle controller in `kube-controller-manager` treats a cluster where all nodes are not ready as a sign that the control plane has lost contact, not that every machine died. It enters a full-disruption mode that stops evictions, so a network cut does not delete every Pod. A partial outage does not trigger this, so some nodes' Pods can still be evicted.
  • How would you keep 400 self-run store clusters identical over time?
    Build them from one declarative definition, never by hand: a management layer such as Cluster API or a fleet manager for the clusters themselves, and a GitOps controller for everything inside them. Roll changes in waves (a few pilot stores first), report each cluster's version centrally, and treat any drift as a defect to reconcile.
  • When would you choose one central cluster with store machines as nodes instead of 400 clusters?
    When stores have reliable links and the team is small: one control plane, one fee and one upgrade. The cost is a single, very large blast radius, kubelets permanently depending on the WAN, and scheduling constraints to keep each store's Pods on its own nodes. For offline-critical stores it is usually the wrong trade.

It is like choosing between a central dispatch office and a manager in every shop: central dispatch is cheaper per decision but useless when the phone line is down, while local managers cost a salary each and need a common playbook.

saying these in an interview costs you the question

  • A managed control plane keeps a store fully working during a WAN outage.
  • The per-cluster fee is irrelevant because node cost dominates.
  • Self-running a control plane is free because the nodes are already paid for.
  • Pods on disconnected nodes are always evicted after five minutes.
  • 400 clusters can be kept consistent by hand with good documentation.