How does a Kubernetes node register itself with the cluster and prove it is still alive, and what would you check when a node shows status NotReady while its workloads are still serving traffic?
answer
- bootstrap token -> CSR -> system:node:<name> cert
- Lease in kube-node-lease renewed ~10s
- status update ~5min; Lease is the cheap heartbeat
- ~40s grace -> Ready=Unknown -> unreachable taint -> 300s eviction
- NotReady = reporting path broken, containers may still run
basics
~20 sOn startup the kubelet authenticates (usually a bootstrap token plus a CSR) and creates the Node object with its capacity and labels. It then renews a Lease in the kube-node-lease namespace every ~10s as a heartbeat and updates node conditions less often. NotReady means the kubelet stopped reporting or reported a problem — not that containers died.
solid answer
~50 s**Registration:** the kubelet starts with a bootstrap kubeconfig, submits a CertificateSigningRequest, and once approved gets a client certificate identifying it as `system:node:<name>` in group `system:nodes`. It then creates the Node object carrying capacity, allocatable, labels, taints and runtime version. NodeRestriction admission limits a kubelet to modifying only its own Node and the Pods bound to it. **Heartbeats:** two mechanisms. The kubelet renews a **Lease** object in `kube-node-lease` roughly every 10 seconds — cheap and scalable — and separately updates **Node conditions** (Ready, MemoryPressure, DiskPressure, PIDPressure) every few minutes or on change. The node-lifecycle controller marks the node `Ready=Unknown` (NotReady) after the node-monitor grace period (~40s) without a heartbeat, then applies `node.kubernetes.io/unreachable:NoExecute`, which evicts Pods after their toleration seconds (default 300s). **Debugging NotReady:** running Pods prove the runtime is fine, so suspect the reporting path — kubelet down or wedged, expired kubelet client certificate, node-to-API-server connectivity, container runtime down, or a real condition such as DiskPressure. `kubectl describe node` shows which condition and its message.
code
bash · 8 lineskubectl get nodes
kubectl describe node node-1 | sed -n '/Conditions/,/Addresses/p'
kubectl get lease -n kube-node-lease node-1 -o yaml # renewTime
# on the node
systemctl status kubelet
journalctl -u kubelet -n 200 --no-pager
openssl x509 -in /var/lib/kubelet/pki/kubelet-client-current.pem -noout -datesgo deeper
Know that the kubelet registers the Node object and sends heartbeats, and that NotReady means the control plane stopped hearing from it.
Distinguish the Lease heartbeat from node status updates and outline the grace period, unreachable taint and eviction timeline.
Drive an actual diagnosis — describe node, kubelet journal, certificate expiry, API connectivity, runtime health, disk and PID pressure — and explain why running Pods are unaffected.
Reason about fencing and split-brain for stateful workloads, eviction rate limits during correlated failures, heartbeat load on etcd at scale, and the Node authorizer and NodeRestriction trust boundary.
## Registration A node joins by convincing the API server it is who it claims to be: 1. The kubelet starts with a **bootstrap kubeconfig** holding a short-lived bootstrap token (what `kubeadm join` hands you). 2. It submits a **CertificateSigningRequest** for a client certificate. With TLS bootstrapping enabled, the CSR approver auto-approves node CSRs backed by valid bootstrap credentials. 3. It receives a certificate with subject `system:node:<nodeName>` in group `system:nodes`, which the built-in Node authorizer understands. 4. With node registration enabled (the default) it **creates the Node object**, populating capacity (CPU, memory, pods, ephemeral storage), allocatable (capacity minus system and kube reserved), labels (hostname, OS, instance type), any taints it was told to register with, and node info such as kernel, OS image, container runtime version and kubelet version. 5. Certificates rotate automatically when rotation is enabled — a very common cause of a node going NotReady months later is exactly this failing. Two security controls matter here: the **Node authorizer**, restricting what a kubelet may read (only the Secrets, ConfigMaps and PVCs referenced by Pods on its node), and the **NodeRestriction** admission plugin, which stops a kubelet editing other nodes' objects or granting itself privileged labels. ## Heartbeats: why there are two Originally the kubelet proved liveness by updating `.status` on its own Node object every few seconds. In a large cluster each of those is a full object write to etcd — thousands of nodes writing every 10 seconds became a serious source of etcd churn. So the signal was split: - **Lease objects** (`coordination.k8s.io/v1`) in the `kube-node-lease` namespace, one per node, holding little more than `renewTime`. The kubelet renews roughly every 10 seconds. Tiny object, tiny write. - **Node status updates**, carrying the real detail (conditions, capacity, images), computed frequently but pushed to the API server only every 5 minutes unless something actually changed. The node-lifecycle controller in kube-controller-manager watches both: a fresh Lease is enough to consider the node alive. ## What happens when heartbeats stop 1. After the node-monitor grace period (~40s) with no Lease renewal, the controller sets condition `Ready=Unknown`; `kubectl get nodes` shows **NotReady**. 2. The controller applies the taint `node.kubernetes.io/unreachable` with effect `NoExecute` (or `not-ready` if the kubelet reported itself unhealthy). 3. Pods carry a default toleration for those taints with `tolerationSeconds: 300`, so after five minutes they are marked for deletion and their controllers create replacements elsewhere. 4. StatefulSet Pods are **not** force-deleted — the API object may linger `Terminating` because the control plane cannot confirm the container is gone, and a replacement with the same identity could mean two writers to one volume. An unreachable node holding a StatefulSet replica needs deliberate operator action. 5. If a large fraction of nodes go unready at once, the controller enters a rate-limited mode and slows or stops evictions, on the theory that the network, not the nodes, is broken. All of this is control-plane bookkeeping. The containers on that node keep running until the kubelet returns and reconciles, or the machine actually dies. That is exactly why a NotReady node can still serve traffic — and why fencing matters for stateful systems. ## Debugging a NotReady node Work from the reporting path inward: 1. `kubectl describe node <name>` — which condition is false or unknown, and what message? `KubeletNotReady`, container runtime down, DiskPressure, PIDPressure, or a CNI message like network plugin not ready. 2. On the node: `systemctl status kubelet`, `journalctl -u kubelet -n 200`. A crash-looping kubelet, a config parse error, or a wedged process shows here. 3. **Certificates** — an expired kubelet client cert means it cannot talk to the API server at all; check the kubelet PKI directory and the API server's CSR list for pending rotation requests. 4. **Connectivity** — can the node reach the API server endpoint on 6443? Security group, route or DNS changes commonly break only the control-plane path while workload traffic is fine. 5. **Runtime** — `crictl info`; a dead containerd makes the kubelet report NotReady even though existing containers survive. 6. **Resources** — a full disk trips DiskPressure; memory pressure triggers eviction; PIDPressure indicates a fork storm. 7. `kubectl get lease -n kube-node-lease <name> -o yaml` to see whether renewals stopped and exactly when. The distinction to voice in an interview: **NotReady is a statement about the kubelet's reporting, not about your containers.**
- Why did Kubernetes introduce Lease objects instead of keeping node status updates as the heartbeat?Every node-status update is a write of a fairly large object to etcd, and at thousands of nodes heartbeating every few seconds that became a dominant source of etcd write load and database growth. Leases are tiny objects containing essentially a renew timestamp, so frequent renewals are cheap. Full node status is then written only every few minutes or when something actually changes.
- A node is unreachable and holds a StatefulSet Pod. Why does that Pod stay Terminating, and what is the risk of force-deleting it?The control plane cannot confirm the container has stopped, and StatefulSet identity guarantees at most one Pod per ordinal. Force-deleting removes the API object without proof the process is dead, so if the node is merely partitioned you can end up with two instances writing to the same volume or claiming the same identity. The safe order is to fence the node — power it off or detach the volume — before forcing deletion.
- What stops a compromised kubelet from reading every Secret in the cluster?The Node authorizer restricts a kubelet's reads to objects referenced by Pods scheduled on its own node, and the NodeRestriction admission plugin prevents it from modifying other Node objects or setting privileged labels on itself. Together they bound the blast radius of a single compromised node to the workloads that already run there.
saying these in an interview costs you the question
- Saying NotReady means the node's containers have stopped
- Not knowing about Lease objects and claiming status updates are the only heartbeat
- Force-deleting StatefulSet Pods from an unreachable node without fencing
- Assuming eviction is immediate rather than after the 300s toleration
- Forgetting expired kubelet client certificates as a cause of sudden NotReady