Your team wants `kubectl scale` and a Kubernetes HorizontalPodAutoscaler to drive a custom resource. What must its CustomResourceDefinition's scale subresource declare, and what must the controller keep populated?
answer
- a uniform Scale view
- three JSON paths into the object
- selector must be a string
- controller writes actual count and selector
- HPA on the custom kind only
basics
~10 sThe CRD declares subresources.scale with specReplicasPath, statusReplicasPath and, for an HPA, labelSelectorPath. The controller applies the desired replicas and writes the actual replica count and a string label selector for the Pods it owns.
solid answer
~40 s`subresources.scale` makes the API server serve `/scale`, which returns an `autoscaling/v1` `Scale` built from three JSON paths. `specReplicasPath` (required, under `.spec`) is the desired count, and if no value is there, GET `/scale` fails. `statusReplicasPath` (required, under `.status`) is the actual count and reads as 0 when empty. `labelSelectorPath` (optional, under `.spec` or `.status`) must point at a **string** selector such as `app=ledger-worker`, and the HPA needs it: an empty selector makes the HPA emit a `SelectorRequired` event and set `ScalingActive=False`. The controller has to turn the desired count into real Pods, for example by setting its Deployment's replicas, and keep the actual count and selector current in status. The HPA controller already has `get`/`update` on `*/scale`, but people running `kubectl scale` need RBAC on `<plural>/scale`. Nothing else should autoscale the inner Deployment.
code
yaml · 18 linesapiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: nightly-ledger-workers
spec:
scaleTargetRef:
apiVersion: ledger.example.com/v1
kind: LedgerWorkerPool
name: nightly
minReplicas: 7
maxReplicas: 11
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65go deeper
Remember that kubectl scale and the HPA work through a /scale endpoint, and that a CRD opts into it under subresources.scale.
Name the three paths, which are required, where each must live, and what the Scale shows when a value is missing.
Diagnose an HPA stuck on an invalid selector or zero replicas, sort out RBAC on the scale resource, and keep a single owner for the inner workload's replica count.
Decide which custom kinds deserve a uniform scale contract and which hide topology choices behind a single replica number.
## What the scale subresource is Built-in workloads such as Deployments and StatefulSets expose a **`/scale` subresource**, a small uniform view of "how many replicas do you want, how many do you have, and which Pods are yours". Generic clients use that view instead of each kind's full schema: `kubectl scale` and the **HorizontalPodAutoscaler** (HPA) controller both work through it. A custom resource gets the same endpoint when its CustomResourceDefinition sets `subresources.scale`. The API server then serves an `autoscaling/v1` `Scale` object whose fields are mapped onto the custom resource by JSON paths. ## The three paths | CRD field | Required | Must live under | Maps to `Scale` field | If the value is missing | |---|---|---|---|---| | `specReplicasPath` | yes | `.spec` | `spec.replicas` | GET `/scale` returns an error | | `statusReplicasPath` | yes | `.status` | `status.replicas` | reported as 0 | | `labelSelectorPath` | no | `.spec` or `.status` | `status.selector` | reported as an empty string | Rules to remember: - Paths are simple JSON paths without array notation, such as `.spec.replicas`. - `labelSelectorPath` must point at a **string** holding a serialized selector (`app.kubernetes.io/name=ledger-worker,ledger.example.com/pool=nightly`), not at a structured `matchLabels` object. - The fields the paths point at must also exist in the structural schema, or pruning removes them. Declaring them is the schema's job. - A write to `/scale` changes the custom resource's `spec` replicas field, and that bumps `metadata.generation` like any other spec change. ## The CRD and the HPA together Take a `LedgerWorkerPool` custom resource whose controller runs the workers for the nightly ledger-reconciliation batch as a 7-replica Deployment on a 27-worker cluster. ```yaml apiVersion: apiextensions.k8s.io/v1 kind: CustomResourceDefinition metadata: name: ledgerworkerpools.ledger.example.com spec: group: ledger.example.com scope: Namespaced names: plural: ledgerworkerpools singular: ledgerworkerpool kind: LedgerWorkerPool versions: - name: v1 served: true storage: true subresources: status: {} scale: specReplicasPath: .spec.replicas statusReplicasPath: .status.replicas labelSelectorPath: .status.selector schema: openAPIV3Schema: type: object x-kubernetes-preserve-unknown-fields: true ``` The HPA then names the custom kind directly: `scaleTargetRef` with `apiVersion: ledger.example.com/v1`, `kind: LedgerWorkerPool` and the object's name. How the HPA turns metrics into a replica count is a separate subject. What matters here is the contract it relies on. ## What the controller owes The API server only **maps** fields. It does not create a single Pod. The controller must: 1. Read `spec.replicas` and apply it to what it manages, for example by setting its Deployment's `spec.replicas` from 7 to 11 when the window opens. 2. Write `status.replicas` with the number that actually exists, so the HPA and `kubectl get` see reality. 3. Write `status.selector` as the string form of the selector matching **exactly** the worker Pods. The HPA uses it to find Pods whose metrics it averages. 4. Write status through `/status`, which also needs the status subresource. With the Kubebuilder markers, a single line generates the scale block: `+kubebuilder:subresource:scale:specpath=.spec.replicas,statuspath=.status.replicas,selectorpath=.status.selector`. ## Failure modes you will actually meet - **HPA never scales and reports an invalid selector.** The controller never wrote `status.selector`, or `labelSelectorPath` is not set. The HPA controller emits a `SelectorRequired` warning event ("selector is required") and sets its `ScalingActive` condition to `False` with reason `InvalidSelector`. - **HPA and `kubectl get` show 0 current replicas.** Nothing is at `statusReplicasPath`, so the Scale reports 0. - **`kubectl scale` errors on a new object.** `spec.replicas` is unset and no default fills it, so GET `/scale` fails. Default the field in the schema. - **`kubectl scale ledgerworkerpool/nightly --replicas=11` is forbidden.** The user has rights on `ledgerworkerpools` but not on `ledgerworkerpools/scale`. The HPA controller is unaffected, because its bootstrap role already grants `get` and `update` on `*/scale` in every group. - **Replica count flaps.** Someone also pointed an HPA at the inner Deployment. Now two writers own its replicas: the operator resets the count from the custom resource while the second HPA changes it. Autoscale only the custom resource. - **Selector too broad.** If the selector also matches unrelated Pods, their metrics skew the average, and the HPA refuses to act when another HPA's Pods overlap. ## When it is worth it Add the scale subresource when a custom resource really represents a horizontally scalable set and users or autoscalers should change its size without knowing its schema. Skip it for singleton or topology-driven resources, where a single replica number would hide important choices.
- Why must `labelSelectorPath` point at a string instead of a `matchLabels` object?The `Scale` object's `status.selector` is a string holding a serialized label selector, and the API server copies the value at the path into it as-is. A structured selector cannot be copied into that field. Controllers usually keep a structured selector in spec for their own use and also write its string form, such as `app.kubernetes.io/name=ledger-worker`, into `status.selector`.
- Your operator creates a Deployment for each LedgerWorkerPool. Why not put the HPA on that Deployment instead?The operator sets the Deployment's replicas from the custom resource on every reconcile, so an HPA writing the same field gives two owners that keep undoing each other. Targeting the custom resource keeps one path of intent: HPA to `/scale`, `/scale` to the custom resource's spec, and the operator to the Deployment.
saying these in an interview costs you the question
- Believing the API server creates or removes Pods when /scale is written
- Pointing labelSelectorPath at a structured matchLabels object
- Thinking statusReplicasPath may point anywhere under spec
- Also attaching an HPA to the operator's inner Deployment
- Assuming the HPA needs a custom RBAC grant for each new CRD
- Leaving status.selector empty and expecting the HPA to infer Pods