In Cluster API, what is a management cluster, and how does it create and upgrade workload Kubernetes clusters declaratively?
answer
- clusters as custom resources
- infrastructure, bootstrap, control-plane providers
- kubeconfig Secret in management cluster
- upgrade = replace machines
- clusterctl move to pivot
basics
~20 sA Cluster API management cluster is a Kubernetes cluster that runs Cluster API controllers. You declare workload clusters there as custom resources, and the controllers provision machines, bootstrap them with kubeadm and replace them to upgrade.
solid answer
~50 sCluster API (CAPI) represents a Kubernetes cluster as Kubernetes objects. The **management cluster** runs the core controllers plus three kinds of provider: an **infrastructure provider** (VMs, networks, load balancers), a **bootstrap provider** (the kubeadm configuration each machine boots with) and a **control-plane provider** (`KubeadmControlPlane`). You apply a `Cluster`, a `KubeadmControlPlane` and a `MachineDeployment`. The controllers then create `Machine` objects, the machines join the new cluster, and its admin kubeconfig is stored in a `<cluster-name>-kubeconfig` Secret in the management cluster. To upgrade, you change `version` on the `KubeadmControlPlane` and in the `MachineDeployment` template, and by default CAPI rolls out new machines and deletes the old ones. `clusterctl init` installs the providers, and `clusterctl move` pivots the objects to another management cluster. If the management cluster is lost, the workload clusters keep serving, but scaling, upgrades and machine remediation stop.
code
bash · 8 linesclusterctl init --infrastructure docker
clusterctl generate cluster accrual-eu-2 --flavor development \
--kubernetes-version v1.37.0 \
--control-plane-machine-count 3 \
--worker-machine-count 38 > accrual-eu-2.yaml
kubectl apply -f accrual-eu-2.yaml
clusterctl describe cluster accrual-eu-2
clusterctl get kubeconfig accrual-eu-2 > accrual-eu-2.kubeconfiggo deeper
Remember the two roles: the management cluster holds the objects and controllers, and workload clusters run the applications.
Walk through the objects (Cluster, KubeadmControlPlane, MachineDeployment, Machine) and the three provider kinds, and explain that an upgrade is a rolling replacement of machines.
Explain how you operate the management cluster as tier-zero: pivoting, backups, access control, and what breaks if it is lost.
Decide how many management clusters the fleet needs, and weigh the blast radius of one central control point against running many management planes.
## Why a fleet needs a management plane Creating one cluster by hand with `kubeadm init` and `kubeadm join` works. Doing it for fifteen clusters, and then upgrading all fifteen every few months, does not scale. **Cluster API** (CAPI) is a SIG Cluster Lifecycle project that applies the Kubernetes model to clusters themselves: you write the desired cluster as objects, and controllers reconcile real infrastructure to match. The cluster where those objects and controllers live is the **management cluster**. The clusters it creates, which run your applications, are **workload clusters**. ## The object model | Object (API group) | What it describes | |---|---| | `Cluster` (`cluster.x-k8s.io`) | The cluster as a whole, with references to its control plane and infrastructure | | `KubeadmControlPlane` (`controlplane.cluster.x-k8s.io`) | Control-plane replica count, Kubernetes `version`, and the kubeadm configuration | | `MachineDeployment` | A group of worker machines, rolled out like a Deployment rolls out pods | | `MachineSet` / `Machine` | One revision of workers, and one node's machine | | Infrastructure machine objects | The provider-specific VM or bare-metal host behind each `Machine` | | `MachineHealthCheck` | Rules for replacing unhealthy machines (remediation) | | `ClusterClass` | A reusable template, so many clusters can be declared with only a few parameters | Providers are pluggable: - **Infrastructure providers** talk to a cloud, a virtualisation platform or bare metal. - **Bootstrap providers** produce the data a machine boots with. The kubeadm provider makes each machine run `kubeadm init` or `kubeadm join`. - **Control-plane providers** manage the control-plane machines as a group. ## Bootstrap and pivot A management cluster has to exist before it can create clusters, so the usual sequence is: 1. Start a temporary **bootstrap cluster** (often kind) and run `clusterctl init` to install the providers. 2. Declare the long-lived management cluster there and let CAPI create it. 3. Run `clusterctl init` on the new cluster, then `clusterctl move --to-kubeconfig ...` to **pivot**. The move pauses reconciliation on the source (`spec.paused` on each `Cluster`), copies the objects and their dependent Secrets, and resumes reconciliation on the target. 4. Delete the bootstrap cluster. The management cluster can now be **self-hosted**, managing its own machines. ```bash clusterctl init --infrastructure docker clusterctl generate cluster accrual-eu-2 --flavor development \ --kubernetes-version v1.37.0 \ --control-plane-machine-count 3 \ --worker-machine-count 38 > accrual-eu-2.yaml kubectl apply -f accrual-eu-2.yaml clusterctl describe cluster accrual-eu-2 clusterctl get kubeconfig accrual-eu-2 > accrual-eu-2.kubeconfig ``` ## Declarative upgrades To upgrade a workload cluster, you change the Kubernetes `version` on the `KubeadmControlPlane` and in the `MachineDeployment`'s machine template. By default CAPI treats machines as **immutable**: - It creates a control-plane machine at the new version, waits for it to join, and removes an old one, repeating until all are replaced. - The `MachineDeployment` then rolls its workers the same way, draining each old node before deleting it. The version-skew and ordering rules still apply (control plane first, one minor version at a time). CAPI automates the replacement but does not remove those rules. `clusterctl upgrade apply` is a different operation: it upgrades the **providers** installed in the management cluster, not the workload clusters. ## When the management cluster is gone Each workload cluster has its own control plane, so losing the management cluster does **not** stop workload clusters or their pods. What you lose: - Scaling or creating clusters and machine groups. - Declarative upgrades. - `MachineHealthCheck` remediation of failed nodes. - Access to the stored `<cluster-name>-kubeconfig` Secrets. This makes the management cluster a **tier-zero** system: run it highly available, back up its objects (`clusterctl move --to-directory` can export them), and keep it separate from workload traffic. ## Fleet shape One management cluster per environment or region is common. It limits blast radius without adding many more management planes. Delivering applications onto the workload clusters is a separate concern, usually handled by a GitOps controller.
- What is the difference between a Cluster API bootstrap cluster and a management cluster?A bootstrap cluster is temporary, often kind on a laptop or CI runner, and exists only to create the first real management cluster. Once that cluster is up, `clusterctl move` pivots the Cluster API objects and their Secrets to it, and the bootstrap cluster is deleted. The management cluster is long-lived and keeps reconciling the workload clusters.
- What does ClusterClass add when a fleet grows past a handful of clusters?`ClusterClass` stores the control-plane and worker templates once. Each `Cluster` then sets `spec.topology` with a class reference, a version, worker groups and variables, instead of carrying full copies of every template. Changing the class or bumping a cluster's topology version is reconciled consistently, so fifty clusters do not drift into fifty hand-edited variants.
- How would you protect a Cluster API management cluster as a single point of control?Run its control plane highly available, keep it off workload traffic, restrict who can edit Cluster API objects, and back up those objects regularly. `clusterctl move --to-directory` exports them, and an etcd snapshot also works. Have a rehearsed plan to restore or re-pivot into a replacement management cluster, since remediation and upgrades stop until one exists.
saying these in an interview costs you the question
- Every workload cluster must run its own Cluster API controllers.
- Losing the management cluster takes all workload clusters down.
- Cluster API upgrades nodes by patching the kubelet in place by default.
- clusterctl move copies only the Cluster object, not machines or Secrets.
- Cluster API replaces kubeadm instead of building on it.
- clusterctl upgrade apply upgrades the workload clusters' Kubernetes version.