How does a Kubernetes device plugin make a node's GPUs visible to kube-scheduler and wire an allocated GPU into a container?
answer
- kubelet directory with Unix sockets
- Register, then a watched stream
- healthy devices become allocatable
- kubelet picks IDs, plugin returns injections
- checkpoint survives kubelet restarts
basics
~20 sThe plugin registers with the kubelet over a gRPC socket and streams its healthy devices through ListAndWatch. The kubelet publishes the count in Node allocatable for the scheduler, then calls Allocate to get the device files, mounts and environment variables to inject.
solid answer
~40 sA device plugin, usually a DaemonSet pod, serves gRPC on a Unix socket under `/var/lib/kubelet/device-plugins/`. It calls `Register` on the kubelet's `kubelet.sock` with its resource name, such as `nvidia.com/gpu`, and API version `v1beta1`. The kubelet opens a `ListAndWatch` stream to the plugin, which sends device IDs, each marked `Healthy` or `Unhealthy`, and optional NUMA topology. The kubelet writes healthy devices into the Node's `status.allocatable`, and kube-scheduler only counts them. When a pod is admitted, the kubelet's device manager picks free device IDs, optionally asking `GetPreferredAllocation`, and calls `Allocate`. The response lists the environment variables, mounts, device specs, annotations or CDI device names the runtime adds to the container. The kubelet checkpoints which IDs belong to which container, so a restart does not hand the same GPU out twice.
code
bash · 2 lineskubectl get node gpu-node-07 -o jsonpath='{.status.capacity.nvidia\.com/gpu}{" "}{.status.allocatable.nvidia\.com/gpu}{"\n"}'
kubectl describe node gpu-node-07 | grep -A8 'Allocated resources'go deeper
Recall that a DaemonSet-run device plugin makes GPUs appear as a counted resource on the Node, and that pods then request it by name.
Walk the call sequence: Register, ListAndWatch, allocatable, Allocate. Say what the Allocate response injects into the container.
Show operational awareness: plugin re-registration after kubelet restarts, the unhealthy-device gap between capacity and allocatable, and why the scheduler cannot choose specific devices.
Judge the limits of a count-only interface for a heterogeneous accelerator fleet, and when node-local choice and topology rejection justify a richer allocation API.
## The pieces - **Device plugin**: a vendor-supplied process, typically run as a DaemonSet on accelerator nodes. It knows how to discover and health-check the hardware. - **Kubelet device manager**: the part of the kubelet that talks to plugins, tracks which device IDs are allocated, and reports counts on the Node. - **kube-scheduler**: sees only the counts in `status.allocatable`. It never talks to a plugin. - **Container runtime**: receives the instructions the plugin returned and makes the device appear inside the container. The contract is the device-plugin gRPC API, version `v1beta1`. ## Step by step 1. **Registration.** The kubelet serves a `Registration` service at `/var/lib/kubelet/device-plugins/kubelet.sock`. The plugin starts its own gRPC server on another socket in the same directory. It then calls `Register` with three things: its endpoint (socket name), its **resource name** (`nvidia.com/gpu`), and the API version. 2. **Options.** The kubelet calls `GetDevicePluginOptions`. The plugin says whether it needs `PreStartContainer` calls (`pre_start_required`) and whether it implements `GetPreferredAllocation` (`get_preferred_allocation_available`). 3. **Inventory.** The kubelet calls `ListAndWatch`, a server stream. The plugin sends the full device list whenever anything changes. Each device has an ID, a health value (`Healthy` or `Unhealthy`), and optionally `TopologyInfo` naming its NUMA node. 4. **Advertising.** On its next node status update, the kubelet sets `status.capacity` to all known devices and `status.allocatable` to the healthy ones. For example, a node with one failed card out of four shows capacity 4 and allocatable 3. 5. **Scheduling.** kube-scheduler's resource fit check subtracts the GPU requests of pods already bound to the node from allocatable. A pod needing `nvidia.com/gpu: 1` is bound only where at least one is free. 6. **Allocation.** When the kubelet admits the pod, the device manager chooses concrete IDs. If the plugin supports it, the kubelet first asks `GetPreferredAllocation` for a preferred set, for example GPUs on the same interconnect. It then calls `Allocate` with those IDs. 7. **Injection.** `Allocate` returns a `ContainerAllocateResponse` per container, with: - `envs`: environment variables, such as a visible-devices list; - `mounts`: host paths for driver libraries and tools; - `devices`: host device nodes plus cgroup permissions (`r`, `w`, `m`); - `annotations`: passed through to the runtime; - `cdi_devices`: fully qualified Container Device Interface names the runtime resolves itself. 8. **Optional pre-start hook.** If requested, the kubelet calls `PreStartContainer` just before the container starts, for example to reset the device. ## Bookkeeping and failure behaviour | Event | What the kubelet does | |---|---| | Device turns `Unhealthy` | Drops it from allocatable; capacity still counts it | | Plugin socket disappears | Stops advertising after a grace period; the resource goes to zero | | Kubelet restarts | Recreates its socket, so plugins must detect this and register again; allocations are restored from the `kubelet_internal_checkpoint` file | | `Allocate` fails | Pod admission fails on that node; the pod does not start there | The checkpoint matters. Without it, a kubelet restart could forget that GPU 2 belongs to the running transcoding pod and give it to the next pod. ## Why the design looks like this - **Vendor-neutral core.** The kubelet has no GPU code. Anything that can be listed, health-checked and exposed as files fits: GPUs, SR-IOV virtual functions, FPGAs. - **Counting keeps the scheduler simple.** The price is that the scheduler cannot pick *which* device, and cannot filter on device properties. That limitation is what Dynamic Resource Allocation later addressed. - **Node-local choice.** The kubelet, not the scheduler, picks the IDs, so node-local topology can shape the choice. ## Inspecting it - `kubectl get node <name> -o jsonpath='{.status.allocatable}'` shows the advertised count. - `kubectl describe node <name>` shows `Allocated resources`, including GPU requests. - The kubelet's pod-resources gRPC endpoint lets monitoring agents map device IDs to pods.
- The video-transcoding pods run slowly because each one's GPU and memory sit on different NUMA nodes. What kubelet feature addresses that, and what can it not fix here?The kubelet's Topology Manager. With `topologyManagerPolicy` set to `restricted` or `single-numa-node`, it merges NUMA hints from the device manager and other hint providers, and rejects the pod with `TopologyAffinityError` if they cannot be aligned. It cannot pin CPUs for a 0.35-core request, because the static CPU manager gives exclusive CPUs only to Guaranteed pods with integer CPU. The scheduler is also not NUMA-aware, so rejected pods can keep landing on the same node.
- Where does the vendor's GPU Operator fit relative to the device plugin?It is packaging and lifecycle automation around the same interface. It installs the driver, the container-runtime hooks and the device plugin DaemonSet, so the plugin can register with the kubelet as described. Kubernetes still sees only the extended resource the plugin advertises. The Operator's own components, and GPU sharing modes, belong to the GPU-fleet topic, not to the core device-plugin mechanism.
- A GPU fails while a pod is using it. Does Kubernetes move the pod?No. The plugin reports the device `Unhealthy`, so the kubelet removes it from allocatable and no new pod gets it. The running container is not evicted by the device manager. The workload fails or keeps running depending on the driver, and an operator or controller has to act.
saying these in an interview costs you the question
- kube-scheduler calls the device plugin to find a free GPU
- The device plugin edits the Node object directly to add GPUs
- Unhealthy GPUs disappear from both capacity and allocatable
- The kubelet needs vendor-specific GPU code compiled in
- After a kubelet restart all GPU assignments are forgotten