A retailer plans a 5-node edge Kubernetes cluster in each of 400 stores. Would you use provider-managed control planes or run your own, and how would you decide?
answer
- four axes, not one
- fee per cluster times 400
- WAN link joins the control plane
- local control plane eats node memory
- fleet automation or nothing
basics
~20 sWeigh the per-cluster fee times 400, the store's dependence on a WAN link to a regional control plane, and the memory a local control plane takes from five small nodes. Many fleets self-run a lightweight distribution with strong fleet automation instead.
solid answer
~50 sI would decide on four axes. **Cost**: managed control planes are usually billed per cluster, so 400 small clusters multiply a fee that is negligible for one large cluster. **Connectivity**: a managed control plane runs in the provider's region, so each store depends on its WAN link; during an outage kubelets keep running existing Pods, but nothing can be scheduled, rescheduled or reconfigured. **Local capacity**: a self-run control plane takes memory and CPU from the five nodes that should run the inference server. **Operations**: self-running 400 control planes means owning patching, certificates, etcd and upgrades for all of them, which is only viable with fleet automation. For stores that must keep working offline, I would usually self-run a lightweight distribution with a single declarative fleet pipeline, and keep managed control planes for the central regional clusters.
go deeper
Recall that a managed control plane runs in the provider's region and that providers usually charge a fee for every cluster.
Explain what still works on a node that loses contact with the API server and what stops, such as scheduling and configuration changes.
Size the store nodes against the real workload with and without a local control plane, and describe the eviction behaviour during partial outages.
Decide on cost, connectivity, local capacity and operations together, and commit to the fleet automation that self-running at scale demands.
## The decision in one sentence A managed control plane trades configuration freedom and a per-cluster fee for someone else running the API server and etcd. At the edge, two more things matter: **where** that control plane runs and **how many** of them you need. ## Axis 1: cost scales with cluster count Managed services commonly charge a flat fee per cluster, independent of its size. That is negligible for a few large clusters and significant for a fleet of small ones. An illustrative calculation with a hypothetical fee of $0.10 per cluster-hour: - One cluster: 0.10 x 730 hours = **$73 per month**. - 400 store clusters: 73 x 400 = **$29,200 per month**, before any node cost. The same 400 stores would pay the fee once if they were nodes in one central cluster, which is why the choice of **cluster count** and the choice of **managed versus self-run** cannot be made separately. ## Axis 2: the WAN link becomes part of the control plane A managed control plane runs in the provider's region. Each store's kubelets talk to it over the store's uplink. When that link fails: 1. **Running Pods keep running.** The kubelet does not stop containers because it lost the API server. 2. **Nothing new happens.** No scheduling, no replacement of a crashed node's Pods on the other four nodes, no ConfigMap or Deployment changes, no `kubectl` from the store. 3. **Nodes go `Unknown`.** Their Lease objects stop renewing, and after the node lifecycle controller's grace period (`--node-monitor-grace-period`, 50 seconds by default) their `Ready` condition becomes `Unknown`. 4. **Eviction is dampened, not guaranteed.** Upstream `kube-controller-manager` stops evictions when *every* node in the cluster is not ready, treating it as a control-plane connectivity problem. A partial outage, for example two of five nodes on a flaky switch, can still lead to Pods being deleted after the default 300-second `tolerationSeconds` for the `node.kubernetes.io/unreachable` taint. A provider may also tune these settings, and you cannot see or change them. 5. **Recovery depends on reconciliation.** When the link returns, the kubelets resync with whatever the API server now says. If a store must sell while offline, a remote control plane is a design risk, not an implementation detail. ## Axis 3: a local control plane costs local capacity Self-running puts the control plane on the store's own five nodes. That has a price: | Choice | Survives a node loss | Local cost | |---|---|---| | One server node with an embedded datastore | No control plane until repaired | Lowest | | Three control-plane nodes with an etcd quorum | Yes, one node | Highest; three nodes carry control-plane load | | Remote managed control plane | Not applicable | None locally, WAN dependency instead | An illustrative sizing: each node has about 6.1 GiB allocatable, and each replica of the recommendation-model inference server has a **2.6 GiB memory limit** (and an equal request). Without a local control plane, two replicas fit per node: **10 replicas**. If control-plane components take about 1.2 GiB on three of the nodes, those nodes have 4.9 GiB left and fit only one replica each: 2 x 2 + 3 x 1 = **7 replicas**. That is a 30% capacity loss for local autonomy. Lightweight distributions such as k3s exist to shrink that footprint, which is why they are common at the edge. ## Axis 4: who operates 400 control planes Self-running moves every day-2 task back to you, multiplied by 400: - patch releases and minor upgrades for every store; - control-plane certificate renewal; - datastore backups and recovery; - hardware failures handled by store staff with no Kubernetes knowledge. This is only viable with **fleet automation**: identical, declaratively built clusters, a management layer such as Cluster API or a vendor fleet manager, and a GitOps controller delivering workloads. Without that, managed control planes are often cheaper in practice, even at $29,200 per month. ## Putting it together | Requirement | Leans towards | |---|---| | Store must operate fully offline | Self-run, local control plane | | Reliable uplink, small team, few clusters | Managed | | Hundreds of clusters, strong platform team | Self-run with fleet automation | | Tight node memory | Remote control plane or a single-server local one | | Strict audit or custom API server flags | Self-run | A common answer is a **hybrid**: managed control planes for central regional clusters that run shared services, and self-run lightweight clusters in stores, all built from one template and upgraded by one pipeline. Some providers also offer managed control planes that extend to on-premises hardware; evaluate those on the same four axes rather than on the label. ## What a strong answer shows It asks for the offline requirement first, does the fee arithmetic, sizes the nodes against the real workload and admits that self-running 400 control planes is a platform-team commitment, not a configuration choice.
- Why does upstream Kubernetes stop evicting Pods when every node goes not ready at once?The node lifecycle controller in `kube-controller-manager` treats a cluster where all nodes are not ready as a sign that the control plane has lost contact, not that every machine died. It enters a full-disruption mode that stops evictions, so a network cut does not delete every Pod. A partial outage does not trigger this, so some nodes' Pods can still be evicted.
- How would you keep 400 self-run store clusters identical over time?Build them from one declarative definition, never by hand: a management layer such as Cluster API or a fleet manager for the clusters themselves, and a GitOps controller for everything inside them. Roll changes in waves (a few pilot stores first), report each cluster's version centrally, and treat any drift as a defect to reconcile.
- When would you choose one central cluster with store machines as nodes instead of 400 clusters?When stores have reliable links and the team is small: one control plane, one fee and one upgrade. The cost is a single, very large blast radius, kubelets permanently depending on the WAN, and scheduling constraints to keep each store's Pods on its own nodes. For offline-critical stores it is usually the wrong trade.
It is like choosing between a central dispatch office and a manager in every shop: central dispatch is cheaper per decision but useless when the phone line is down, while local managers cost a salary each and need a common playbook.
saying these in an interview costs you the question
- A managed control plane keeps a store fully working during a WAN outage.
- The per-cluster fee is irrelevant because node cost dominates.
- Self-running a control plane is free because the nodes are already paid for.
- Pods on disconnected nodes are always evicted after five minutes.
- 400 clusters can be kept consistent by hand with good documentation.