In a container orchestration cluster, what does the control plane own and what does each host's node agent own?
answer
- two halves, one cluster name
- who decides against who executes
- API and state store on one side
- node agent enforces what it was assigned
- the halves fail independently
basics
~20 sThe control plane decides: an API, the state store behind it, and the scheduler and controllers that act on declared state. Each host's node agent runs and supervises the containers assigned to that host, alongside the workloads themselves serving traffic.
solid answer
~40 sA cluster is built as two halves. The **control plane** is the deciding half: an API that accepts declared state and answers reads, a durable state store holding every declared object and its last reported status, a scheduler that picks a host for work that has none, and controllers that act on the difference between declared and reported. The **data plane** is the serving half: on every host a node agent receives the assignments meant for that host, drives the local container runtime, supervises the containers and reports status back; beside it sit the containers themselves and the traffic path into them. The rule worth memorising: anything that changes what should exist, or where it lives, is a control-plane act; carrying out an assignment already delivered is a data-plane act.
go deeper
Be able to name both halves and one thing each contains: an API with a state store behind it on the deciding side, a node agent supervising containers on the serving side.
Explain the line rather than the lists: a decision that changes what exists or where it lives is control plane, and enforcing an assignment already delivered is data plane.
Show you have felt the consequence — the halves fail independently, so an unavailable deciding half freezes change while traffic keeps flowing, and your status view is the first thing you lose.
Frame it as where the estate's decision capacity lives: how many clusters carry their own deciding half, who operates it, and what each team is signing up to run when the answer is different per cluster.
## One cluster, two halves A container cluster looks like one system from the outside, but it is built as two halves with different jobs and different failure modes. The **control plane** is the deciding half. It is made of an **API** that accepts declared state and answers reads, a durable **state store** behind that API holding every declared object and the last status reported about it, a **scheduler** that picks a host for work that has none, and a set of **controllers** that watch the record and issue the changes that close the gap between what was declared and what was reported. The **data plane** is the serving half. On every host there is a **node agent** that receives the assignments meant for that host, asks the local container runtime to start them, supervises them, and reports back what it observes. Beside the agent sit the container runtime, the running containers, and the per-host network path that carries packets into them. The short version: **the control plane decides what should exist and where; the data plane makes it exist and keeps it serving.** ## What each half actually holds - **Control plane** — the API front door, the state store, the scheduler, the built-in controllers, and any custom controller added on the same contract. - **Data plane** — the node agent on each host, the container runtime it drives, the running containers, and the local path packets take to reach them. - **Shared in practice** — hosts. A control-plane component is itself a process on a machine, and on a small cluster that machine may also run ordinary workloads. The split is a split of *roles*, not necessarily of hardware. ## Where the line falls, question by question | The question being asked | Answered by | Needs the other half? | |---|---|---| | Should six copies of this workload exist? | Control plane — it is declared in the state store | No | | Which host should the next copy go to? | Control plane — the scheduler | No | | Is the process on this host alive right now? | Data plane — the node agent watching it | No | | This container exited; start it here again | Data plane — the agent already holds the assignment | No | | That host is gone; run its copies somewhere else | Control plane — a new placement decision | Yes | | What is this workload's live status? | Control plane — status is read back through the API | Yes | Read in one direction, that table gives the sentence interviewers are listening for: **anything that changes the set of assignments, or where they live, is a control-plane act; anything that carries out an assignment already delivered is a data-plane act.** ## Why the split is built this way 1. **Decisions want one place.** Copy counts, placement and rollouts need a single consistent view of the whole cluster. Duplicating that view on every host would mean every host disagreeing slightly, and two hosts both concluding they own the same copy. 2. **Execution must be local.** Starting a process, supervising it and forwarding a packet cannot wait on a round trip to a central service for every event. The decision is pushed down once and enforced on the host from then on. The second reason is the one most interview follow-ups are aiming at. Because the halves are separate, they fail separately: the deciding half can be entirely unavailable while every workload in the data plane keeps serving traffic. *The control plane is down* and *the cluster is down* are different sentences, and treating them as the same one is the classic error in this material. ## What the split is not - **Not a hardware boundary.** Larger estates put the deciding half on dedicated machines so workload pressure cannot starve it, but that is an operational choice, not the definition. - **Not a security boundary on its own.** The two halves hold different privileges, but what actually enforces a boundary is the identity and permission rules applied at the API plus the isolation applied to each container on the host. - **Not the only use of those two words.** Other systems borrow *control plane* and *data plane* for their own architectures, so when the term is ambiguous in a conversation, say which system's planes you mean. ## What to say when asked Name the four pieces of the deciding half and the three of the serving half, then give the consequence rather than stopping at the list: the halves are separated precisely so that losing the ability to decide does not stop the cluster from serving. That consequence is the whole reason the distinction is on the interview sheet.
- Does a host that runs control-plane components ever run ordinary workloads too?It can, and on small clusters it commonly does — the plane split is a split of roles, not of machines. Larger estates separate them so a noisy workload cannot starve the deciding half of CPU, memory or disk, but that separation is an operational choice rather than part of the definition.
- Which half does an operator touch when they read a workload's live status?The control plane. Node agents report what they observe, the state store keeps the last report, and the API serves it back. That is why status can go stale or unreadable while the workloads themselves are perfectly healthy: the reading path runs through the deciding half, and the serving path does not.
saying these in an interview costs you the question
- Calls the control plane the place where workloads actually run.
- Thinks each node agent decides which host a copy lands on.
- Says the cluster is down whenever the API is unreachable.
- Treats the state store as a cache of what agents report.
- Assumes the two halves must sit on separate hardware to be separate planes.