What does Kubernetes Dynamic Resource Allocation add over the device-plugin extended-resource model for GPUs, and when would you move to it?
answer
- count versus described inventory
- slice, class, claim, template
- scheduler records the chosen device
- CEL selectors and fallback lists
- bridge from plain extended resources
basics
~20 sDynamic Resource Allocation replaces an anonymous device count with published device inventories (ResourceSlices) and claims (ResourceClaims) that select devices by attributes using CEL. The scheduler picks specific devices, and claims can be shared or given fallbacks. It is GA since Kubernetes 1.34.
solid answer
~50 sA device plugin gives the scheduler only a number: `nvidia.com/gpu: 4`. **Dynamic Resource Allocation** (`resource.k8s.io/v1`, GA in 1.34) publishes each device with attributes and capacities in **ResourceSlice** objects, written by a DRA driver. Workloads create a **ResourceClaim**, usually from a **ResourceClaimTemplate** referenced in `spec.resourceClaims`. The claim asks a **DeviceClass** for devices, filtered by CEL `selectors`. The scheduler's `DynamicResources` plugin then chooses concrete devices and records the allocation in the claim's status, before the pod binds. That makes things possible that a count cannot express: selecting by memory or model, a prioritised `firstAvailable` fallback list, and one claim referenced by several pods. Move to it when you need that expressiveness. Stay with device plugins when whole, interchangeable GPUs are enough, because DRA needs a DRA-capable driver and new objects to operate. `DeviceClass.extendedResourceName` (GA in 1.37) lets existing `nvidia.com/gpu`-style requests be served by DRA devices.
code
yaml · 30 linesapiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: transcode-gpu
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.example.com
allocationMode: ExactCount
count: 1
---
apiVersion: v1
kind: Pod
metadata:
name: transcode-worker
spec:
resourceClaims:
- name: gpu
resourceClaimTemplateName: transcode-gpu
containers:
- name: ffmpeg
image: registry.example.com/transcoder:2.4.1
resources:
requests:
cpu: 350m
claims:
- name: gpugo deeper
Recall that DRA is a newer way to request devices through claim objects instead of a plain resource count.
Name the four objects and say who writes each: ResourceSlice, DeviceClass, ResourceClaim, ResourceClaimTemplate.
Explain that the scheduler allocates specific devices and records them in the claim, and judge when attribute selection justifies moving a pool.
Weigh the expressiveness DRA brings against driver maturity, operational surface and a mixed-model migration across a heterogeneous accelerator fleet.
## The limitation DRA addresses With a device plugin, a Node advertises `nvidia.com/gpu: 4` and a container requests `1`. That is all the scheduler knows: - it cannot tell a 24 GiB card from an 80 GiB one, except through node labels you maintain by hand; - it does not choose *which* device; the kubelet does, after the pod is already bound; - the only unit is a whole integer, with no way to say "this, or failing that, that"; - a device cannot be referenced as an object that several pods share. ## The DRA object model **Dynamic Resource Allocation** lives in the `resource.k8s.io` API group, served as `v1` in current releases. | Object | Written by | Purpose | |---|---|---| | `ResourceSlice` | the DRA driver on each node, or centrally | Publishes devices with a driver name, pool, attributes and capacities | | `DeviceClass` | cluster admin | A named category of devices, with CEL `selectors` and optional driver config | | `ResourceClaim` | user or template controller | A request for devices: which class, how many, which selectors | | `ResourceClaimTemplate` | user | A blueprint; a new claim is generated for each pod that references it | A pod lists claims under `spec.resourceClaims`, each pointing at `resourceClaimName` or `resourceClaimTemplateName`. Containers reference them by name under `resources.claims`. ## How allocation works 1. A DRA driver publishes ResourceSlices for the devices it manages. 2. A pod references a ResourceClaimTemplate, and the resource-claim controller creates a claim for that pod. 3. kube-scheduler's `DynamicResources` plugin evaluates the claim's requests against the ResourceSlices on candidate nodes. Each request can be: - `exactly`: a `deviceClassName`, CEL `selectors`, and an `allocationMode` with a `count`; - `firstAvailable`: an ordered list of sub-requests; the first one that can be satisfied wins. 4. The scheduler writes the chosen devices into the claim's `status.allocation`, then binds the pod to a node where those devices live. 5. The kubelet asks the DRA driver on that node to prepare the devices, and the container gets them, typically through CDI device names. The key shift: **allocation is a scheduling decision recorded in the API**, not a node-local pick after binding. ## Capabilities that follow - **Attribute-based selection.** A CEL expression such as `device.driver == "gpu.example.com"`, combined with attribute or capacity checks, replaces hand-maintained node labels. - **Prioritised alternatives.** `firstAvailable` expresses "a large-memory GPU, otherwise two smaller ones". The feature (`DRAPrioritizedList`) is GA in 1.36. - **Shared claims.** Several pods can reference one named ResourceClaim, and the API tracks who reserved it. Vendor sharing modes build on this, and they are a separate topic. - **Admin and driver configuration.** Configuration can be attached to a DeviceClass or a claim and is passed to the driver. - **Extended-resource bridge.** `DeviceClass.extendedResourceName` lets a plain `resources.limits` request be satisfied from DRA-managed devices. The `DRAExtendedResource` gate is GA and locked on in 1.37. If the field is unset, the implicit name is `deviceclass.resource.kubernetes.io/<class name>`. ## When to move, and when not to Move when: 1. the fleet mixes accelerator models or sizes, and label-based steering has become a maintenance burden; 2. workloads need fallback choices or several devices with constraints between them; 3. the vendor ships a supported DRA driver for your hardware. Stay on device plugins when: - every GPU in a pool is interchangeable, as in a video-transcoding pool where any single card will do; - the vendor driver for DRA is not mature for your stack; - you do not want to operate ResourceSlice publication, claim lifecycle and new RBAC. **Migration risk** is mostly operational. Both models can coexist on a cluster, but a given device should be managed by only one of them. Otherwise the kubelet's device-plugin count and DRA's allocation would both hand it out.
- In DRA, who decides which physical device a pod gets, and where is that recorded?kube-scheduler's `DynamicResources` plugin chooses the devices while it is scheduling the pod, and writes them into the ResourceClaim's `status.allocation` before binding. That is the reverse of device plugins, where the scheduler only counts and the kubelet chooses IDs after binding.
- Can a cluster run the device plugin and a DRA driver for GPUs at the same time?Yes, but not for the same devices. Each physical device should be managed by one model, for example DRA on a new node pool while older pools keep the device plugin. If one card were advertised by both, the extended-resource count and DRA allocations could hand it out twice.
saying these in an interview costs you the question
- DRA is a newer name for the same device-plugin gRPC API
- With DRA the kubelet still picks the device after binding
- DRA lets a container request nvidia.com/gpu: 0.5 directly
- Adopting DRA requires no vendor driver changes
- DRA is still alpha and unsuitable for any production cluster