skip to content

In Kubernetes, how does a pod request a GPU, and why can it request 0.35 CPU but not 0.35 GPU?

level: juniorimportance: must knowfreq 68%

answer

  1. a counted name, not a share
  2. domain-prefixed resource name
  3. limit alone is enough
  4. integers only, request equals limit
  5. kernel meters CPU, nothing meters GPU

basics

~20 s

A pod requests a GPU as an extended resource, such as nvidia.com/gpu: 1, in its container limits. Extended resources are whole units the scheduler counts, not shares. So Kubernetes rejects fractions, requires the request to equal the limit, and never overcommits them.

solid answer

~50 s

A device plugin on each GPU node advertises the GPUs as an **extended resource**, for example `nvidia.com/gpu`, in the Node's capacity and allocatable. A container asks for one under `resources.limits`. If you set only the limit, the request defaults to it. If you set both, the API server requires them to be **equal**. If you set only a request, the pod is rejected with `Limit must be set for non overcommitable resources`. Values must be integers, so `0.5` fails with `must be an integer`. The CPU case is different because the kernel can time-slice a core: a `0.35` CPU request is a scheduling weight, and the CFS quota enforces the limit. Kubernetes has no equivalent meter for a GPU. It hands a container a whole device, and the scheduler just subtracts counted devices from the node's allocatable. So there is no fraction and no overcommit.

code

yaml · 15 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: transcode-worker
spec:
  containers:
  - name: ffmpeg
    image: registry.example.com/transcoder:2.4.1
    resources:
      requests:
        cpu: 350m
        memory: 1536Mi
      limits:
        memory: 1536Mi
        nvidia.com/gpu: 1

go deeper

for a junior

Recall the resource name format, that GPUs go under resources.limits, and the three rules: integers only, request equals limit, no overcommit.

for a middle

Explain why CPU can be fractional: the kernel's CFS quota meters it. Then explain that nothing below Kubernetes meters a GPU, so the scheduler can only count whole devices.

for a senior

Connect the counting model to its cost: idle but allocated GPUs, Pending pods on an underused fleet, and placement that needs labels because the count has no model information.

for a principal

Weigh whole-device allocation against the utilisation it wastes, and decide when sharing tooling or a claim-based allocation model is worth its operational complexity.

## What an extended resource is Kubernetes schedules against **resources** listed on each Node object. The built-in ones are `cpu`, `memory`, `ephemeral-storage` and `hugepages-<size>`. Anything else a node can offer is an **extended resource**: a name with a domain prefix outside `kubernetes.io`, such as `nvidia.com/gpu`, a vendor NIC's virtual-function resource, or an FPGA. A GPU becomes one when a **device plugin** runs on the node and registers with the kubelet. The kubelet then reports the device count in two places: - `status.capacity`: every device the plugin knows about; - `status.allocatable`: the healthy devices pods may use. kube-scheduler knows nothing about GPUs as hardware. To the scheduler, `nvidia.com/gpu: 4` is just a number, and each pod's request is subtracted from it. ## How a pod asks for one A container declares the resource in `resources`, the same way it declares CPU and memory: ```yaml apiVersion: v1 kind: Pod metadata: name: transcode-worker spec: containers: - name: ffmpeg image: registry.example.com/transcoder:2.4.1 resources: requests: cpu: 350m memory: 1536Mi limits: memory: 1536Mi nvidia.com/gpu: 1 ``` The **API server** validates this and applies three rules to every extended resource: 1. **Integer only.** `nvidia.com/gpu: 0.5` is rejected with `must be an integer`. 2. **Request equals limit.** If both are present and differ, validation fails (`must be equal to nvidia.com/gpu limit of 1`). If only the limit is present, defaulting copies it into the request, which is the idiomatic way to write it. 3. **A limit is required.** A request with no limit is rejected with `Limit must be set for non overcommitable resources`. The validation code decides this with a single check: overcommit is allowed only for native resources that are not hugepages. Extended resources and `hugepages-*` both fail it. So hugepages follow the same equal-request rule, even though they are a native resource and not an extended one. ## Why a CPU core can be shared and a GPU cannot | | `cpu` | `nvidia.com/gpu` | |---|---|---| | Unit | millicores; `350m` is valid | whole devices only | | What the request does | scheduling reservation plus a CFS weight | scheduling reservation only | | What enforces the limit | Linux CFS quota in the container's cgroup | nothing meters it; the container gets whole device files | | Overcommit | allowed (sum of limits may exceed the node) | forbidden (request equals limit) | | Sharing between containers | kernel time-slicing | none by default; each allocated device belongs to one container | For CPU, the Linux kernel already knows how to split a core in time and to throttle a cgroup that exceeds its quota. A request of `0.35` cores is meaningful because something below Kubernetes enforces it. For a GPU, Kubernetes has no such meter. The device plugin gives the container device IDs, and the kubelet mounts the matching device nodes and libraries. From then on the process owns the whole card. Allowing `limit > request` would let the scheduler place more GPU consumers than there are devices, which would fail at container start rather than degrade gracefully. So the API forbids it. ## Consequences you will meet - **Idle GPUs are still allocated.** A pod that holds `nvidia.com/gpu: 1` and uses it 5% of the time still blocks that device for everyone else. - **Scheduling is by count, not by model.** The scheduler cannot tell one card model from another. You steer placement with node labels and a `nodeSelector` or node affinity. - **Init containers count too.** An init container's GPU request is part of the pod's effective request. - **Sharing is a separate topic.** Splitting one physical GPU into several advertised units (partitioning, time-slicing) is done by the vendor's tooling, which changes what the plugin advertises. The core rules above do not change: each advertised unit is still a whole integer. ## A worked example Take a node advertising `nvidia.com/gpu: 4`, with three transcoding pods each holding one GPU. The scheduler sees `4 - 3 = 1` free. A fourth pod fits. A fifth stays `Pending` with an `Insufficient nvidia.com/gpu` reason, no matter how idle the four cards are.

  • What happens if the pod sets `requests: {nvidia.com/gpu: 1}` and no GPU limit?
    The API server rejects it during validation with `Limit must be set for non overcommitable resources`. Extended resources cannot be overcommitted, so Kubernetes insists the limit is stated, and stated equal to the request. The reverse, a limit alone, is accepted, because defaulting copies the limit into the request.
  • Does requesting a GPU change the pod's QoS class?
    No. The QoS class is derived only from CPU and memory requests and limits. A pod that holds one GPU but sets a 0.35-core CPU request with no CPU limit is still Burstable, even though its GPU request equals its limit.
  • Why does a pod with `nvidia.com/gpu: 1` stay Pending on a cluster whose GPUs sit at 3% utilisation?
    The scheduler compares requests with allocatable device counts, not with real usage. If every advertised device is already held by some pod, the node has zero free `nvidia.com/gpu` however idle the cards are, so the new pod cannot fit anywhere.

CPU is like a shared meeting room booked in time slots: several teams can use it, and a clock enforces each slot. A GPU is like a key to a private office: you either hold the key or you do not, and handing out more keys than there are offices only creates arguments at the door.

saying these in an interview costs you the question

  • You can request nvidia.com/gpu: 0.5 to share a card between two pods
  • A GPU request lower than its limit lets pods burst onto idle GPUs
  • kube-scheduler looks at GPU utilisation to decide where pods fit
  • GPUs are requested with a special gpu field, not under resources
  • Requesting a GPU makes the pod Guaranteed QoS automatically