In a managed Kubernetes service, which parts of the cluster does the provider operate, and what stays the cluster owner's responsibility?
answer
- split at the API endpoint
- behind it: apiserver, etcd, controllers
- in front of it: nodes, workloads, RBAC
- no SSH, no flags, no control-plane nodes
- managed node groups automate, not decide
basics
~10 sThe provider runs the control plane: kube-apiserver, etcd, the scheduler and controller managers, including their hosts, patching and availability. You still own worker nodes, workloads, RBAC, add-ons, network choices and deciding when to upgrade.
solid answer
~40 sA managed Kubernetes service splits the cluster at the API endpoint. The provider operates everything behind it: the `kube-apiserver`, `etcd`, `kube-scheduler`, `kube-controller-manager` and usually the `cloud-controller-manager`, plus the machines they run on, their certificates, patch releases, etcd operation and a multi-replica, highly available setup backed by an SLA. Everything you *put into* the cluster is still yours: worker nodes (their OS image, size and replacement, unless you pay for a managed node group, which automates provisioning but not your capacity choices), workloads and their resource settings, RBAC, admission webhooks, CNI and add-on configuration, and scheduling the minor-version upgrades within the provider's window. The control-plane machines never appear in `kubectl get nodes`, and you cannot SSH to them or change their flags.
code
bash · 8 lines# Only worker nodes are listed; control-plane hosts are hidden
kubectl get nodes -o wide
# The API server still reports its own health
kubectl get --raw='/readyz?verbose'
# Server version is the provider-run control plane's version
kubectl versiongo deeper
Name the control-plane components the provider runs and list three things that stay yours: nodes, workloads and access control.
Explain where the line sits (the API endpoint), what a managed node group automates and why control-plane nodes are invisible to kubectl.
Show you can place an incident on the right side of the line quickly, such as a failing webhook versus an etcd problem, and know what evidence you can still gather.
Frame the split as a trade: the provider absorbs control-plane toil and risk, and you give up configuration freedom and some diagnostic depth.
## What a managed control plane is A Kubernetes cluster has two halves. The **control plane** is the set of processes that store and reconcile the desired state: `kube-apiserver` (the only entry point), `etcd` (the key-value store holding every object), `kube-scheduler` (picks a node for each new Pod), `kube-controller-manager` (runs the built-in reconcile loops) and, on cloud infrastructure, `cloud-controller-manager` (load balancers, node addresses, routes). The **data plane** is the set of worker nodes running the kubelet, a container runtime and kube-proxy or its CNI replacement. In a **managed Kubernetes service**, the provider runs the control plane as a product. You receive an API endpoint and a kubeconfig; the processes behind that endpoint run on machines you never see. ## What the provider takes over | Area | Provider's job | |---|---| | Control-plane hosts | Provisioning, OS patching, replacement on failure | | `kube-apiserver` | Replicas behind a load balancer, TLS serving certificates, flag configuration | | `etcd` | Quorum sizing, disk, compaction and defragmentation, its own backups | | Patch releases | Rolling new patch versions onto the control plane, usually automatically | | Availability | Spreading replicas across failure domains, an SLA on the API endpoint | | Cluster PKI | Issuing and rotating the control-plane certificates | The practical effect is that the hardest day-2 work of a self-run cluster (etcd quorum loss, expiring control-plane certificates, a failed apiserver upgrade) becomes the provider's incident, not yours. ## What stays with you - **Workloads**: manifests, `resources.requests` and `resources.limits`, probes, PodDisruptionBudgets. - **Access control**: RBAC Roles and bindings, ServiceAccounts, and the mapping of cloud identities to Kubernetes users. - **Extensions you install**: admission webhooks, CRDs and operators, ingress or Gateway API implementations, a metrics pipeline. - **Networking choices**: the CNI configuration, NetworkPolicies, which Services are exposed. - **Upgrade timing**: the provider offers new minor versions, but you choose when to move within its support window. - **Worker nodes**, fully or partly, depending on how you run them. A useful rule: the provider guarantees that the API server answers; it does not guarantee that what you configured through it is correct. A validating admission webhook with `failurePolicy: Fail` whose backend is down will block writes, and that outage is yours. ## Managed node groups versus self-managed nodes | Aspect | Managed node group | Self-managed nodes | |---|---|---| | Provisioning and joining | Provider creates machines and joins them | You build the image and bootstrap the kubelet | | Node OS image | Provider-curated, updated on request | You patch and rebuild | | Node version upgrades | Provider drives a rolling replacement | You cordon, drain and replace | | Kubelet flags and config | Limited, through exposed options | Full control of `KubeletConfiguration` | | Capacity and instance choice | Still yours | Yours | A managed node group removes toil but not decisions: you still choose machine sizes, how many nodes, taints and labels, and whether your Pods tolerate a rolling replacement. Some services go further and hide nodes entirely, billing per Pod, which removes node choice as well. ## What you can no longer see or touch 1. **Control-plane nodes are absent** from `kubectl get nodes`, and there are no `kube-apiserver` static pods in `kube-system`. 2. **No SSH** to control-plane hosts and no direct `etcdctl` access. 3. **No component flags**: you cannot add `--feature-gates` or change `--audit-policy-file`; only the options the provider exposes. 4. **Logs arrive through the provider's logging integration**, if you enable it, rather than from files on disk. ## A worked example A team runs a recommendation-model inference server, each replica capped at a 2.6 GiB memory limit. When a replica is OOM-killed, that is the team's problem: the limit, the model size and the node size are all on their side of the line. When `kubectl apply` times out because an etcd member lost its disk, that is the provider's problem, and the team's job is to check the provider's status page and its support channel rather than log on to a machine they do not have. ## Why interviewers ask The question checks whether a candidate knows that "managed" describes one half of the cluster. Candidates who think a managed service also patches their images, fixes their RBAC or upgrades them forever without a decision usually have not run one.
- Why does `kubectl get nodes` on a managed cluster list no control-plane nodes?The control-plane processes run on the provider's machines, usually outside your network, and those machines never register a Node object with your API server. Because they are not Nodes in your cluster, the scheduler never places your Pods on them. You see the control plane only through its API endpoint and whatever logs or metrics the provider forwards.
- Your validating admission webhook's backend crashes and every Deployment update fails. Is that the provider's outage?No. The API server is healthy and doing exactly what your `ValidatingWebhookConfiguration` told it to do: with `failurePolicy: Fail`, an unreachable webhook rejects the request. The provider's SLA covers the endpoint answering, not the extensions you registered. Fixing it means restoring the webhook backend, narrowing its rules or selectors, or, as a last resort, deleting the configuration.
- What does a managed node group change compared with nodes you bootstrap yourself?It automates creating machines from a provider-curated image, joining them to the cluster and replacing them during version upgrades. You still pick machine sizes, counts, labels and taints, and your workloads still need to survive the rolling replacement. In exchange you usually get less control over kubelet configuration and the node OS.
It is like renting a flat in a serviced building: the landlord keeps the lifts, boiler and wiring running, but what you plug in, who you give keys to and how you furnish the rooms are still your business.
saying these in an interview costs you the question
- A managed service also patches the container images my workloads run.
- The provider is responsible for my RBAC bindings being correct.
- With a managed control plane I never have to plan a version upgrade.
- Worker nodes are always part of what the provider manages.
- I can SSH to the control-plane nodes to debug the API server.
- If my admission webhook breaks the API, the provider's SLA covers it.