CPU is a poor proxy for load on a queue-consuming service. How do you drive a Kubernetes HorizontalPodAutoscaler from a non-CPU signal such as queue depth or requests per second, and what does that require in the cluster?
answer
- Pods / Object / External / ContainerResource
- custom.metrics.k8s.io + external.metrics.k8s.io need an adapter
- Prometheus Adapter vs KEDA (scale-to-zero)
- AverageValue = per-pod backlog
- don't autoscale on latency
basics
~20 sUse the autoscaling/v2 metric types Pods, Object or External instead of Resource. Those are served by custom.metrics.k8s.io and external.metrics.k8s.io, which core Kubernetes does not implement — you install an adapter (Prometheus Adapter, KEDA, a cloud provider's) that translates a query into that API. Use External with AverageValue for queue depth so the target means backlog per pod.
solid answer
~50 s`autoscaling/v2` has four metric sources beyond `Resource`: - **Pods** — a per-pod metric averaged across pods (e.g. `http_requests_per_second`), served by `custom.metrics.k8s.io`. - **Object** — a metric describing one Kubernetes object, e.g. requests per second on an Ingress. - **ContainerResource** — CPU/memory scoped to a single named container, which stops a sidecar from diluting the app's utilization. - **External** — something outside the cluster entirely: SQS backlog, Kafka consumer lag, a Cloud Monitoring series. Served by `external.metrics.k8s.io`. Core Kubernetes ships **no implementation** of those two API groups. You register an adapter — **Prometheus Adapter** (rules mapping PromQL to metric names) or **KEDA** (ScaledObject with per-source scalers, which also gives scale-to-zero) — as an APIService, and the HPA controller consumes it exactly like metrics-server. For backlog, prefer `type: AverageValue`: the reported value is divided by the replica count, so `averageValue: 30` means "30 messages per consumer", and the replica math stays proportional. `type: Value` compares the raw number and does not scale sensibly with replicas.
code
yaml · 28 linesapiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: worker
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: worker
minReplicas: 2
maxReplicas: 50
metrics:
- type: External
external:
metric:
name: queue_messages_visible
selector:
matchLabels:
queue: orders
target:
type: AverageValue
averageValue: "30"
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 75go deeper
Know that CPU is not the only option and that non-CPU metrics require installing something extra.
Name the metric types and the two API groups, and know that Prometheus Adapter or KEDA provides them.
Pick the right target type, justify the signal choice, and account for adapter staleness and availability in the design.
Weigh coupling autoscaling to the observability stack, standardize which signals teams may scale on, and decide where scale-to-zero is worth its cold-start cost.
## Why CPU is often the wrong signal A worker pulling from a queue may sit near-idle on CPU while a million-message backlog accumulates — the bottleneck is I/O wait or downstream latency, not compute. An HTTP service may be latency-bound on a database. Scaling on CPU in either case reacts late or not at all. The fix is to autoscale on a metric that actually correlates with the user-visible symptom: backlog, request rate, or concurrency. ## The three metric APIs Kubernetes deliberately separates the *consumer* (the HPA controller) from the *producer* via three aggregated API groups: 1. **`metrics.k8s.io`** — resource metrics, CPU and memory only, served by metrics-server. Used by `Resource` and `ContainerResource` metric specs. 2. **`custom.metrics.k8s.io`** — metrics *associated with a Kubernetes object*. Used by `Pods` and `Object` specs. 3. **`external.metrics.k8s.io`** — metrics with no Kubernetes object behind them. Used by `External` specs. Core Kubernetes implements only the first. The other two are contracts; you install something that registers an `APIService` for them. ## The two common adapters **Prometheus Adapter** exposes PromQL results through the custom and external metrics APIs. You write rules that select a series, associate it with Kubernetes objects via label matching, name it, and define the query the HPA will call. It is the right choice when Prometheus is already the source of truth for the metric. **KEDA** takes a different shape: you write a `ScaledObject` naming a scaler (Kafka, SQS, RabbitMQ, Azure Service Bus, Prometheus, cron, …), and KEDA *generates and manages an HPA for you* while serving the external metric itself. Two capabilities make it attractive: a catalogue of ready-made source integrations, and **scale-to-zero**, which a plain HPA cannot do because `minReplicas` must be at least 1 unless the alpha `HPAScaleToZero` feature gate is enabled. KEDA activates the workload from zero itself and hands over to the HPA above one replica. ## Value versus AverageValue — the decision that matters For `External` and `Object` metrics you choose the target type: - **`Value`** compares the raw metric to the target and the replica math becomes `ceil(replicas × metric/target)` on an aggregate number. For a backlog that keeps growing, this can demand enormous replica counts. - **`AverageValue`** divides the metric by the current replica count first. `averageValue: 30` on a queue depth means "aim for 30 messages per consumer pod" — a backlog of 600 asks for 20 pods, and the relationship stays intuitive as the fleet grows. For backlog-style signals `AverageValue` is nearly always what you want. For a metric that is already a per-unit rate, either can work. ## Choosing the metric itself Good autoscaling signals share properties: they lead the symptom rather than lag it, they are roughly linear in replica count, and they are cheap to collect. Backlog depth and consumer lag are excellent — they rise before customers notice. In-flight request concurrency is good. **Latency percentiles are a trap**: latency is not linear in replicas, and scaling on it creates a feedback loop where a slow dependency triggers a stampede of new pods that make the dependency slower. Error rate is worse still. Prefer a saturation signal over a symptom signal. ## Operational consequences - **A new dependency in the availability path.** If the adapter or Prometheus is down, the metric goes `<unknown>`; scale-down is suppressed and scale-up on that metric stops. Autoscaling now depends on your monitoring stack's uptime, which is an architectural decision worth stating explicitly. - **Staleness.** Prometheus scrape interval plus rule evaluation plus the HPA's 15s sync means the control loop can easily be a minute behind reality. Size `minReplicas` so the workload survives that lag. - **Cardinality and cost.** Adapter rules that match high-cardinality series can be expensive to evaluate on every HPA sync. - **Debuggability.** `kubectl get --raw "/apis/external.metrics.k8s.io/v1beta1/namespaces/<ns>/<metric>"` returns exactly what the HPA sees; make that your first check rather than reasoning about adapter config. ## Combining with CPU Listing a queue metric *and* CPU in one HPA is common and safe: the controller computes a desired count per metric and takes the maximum, so CPU acts as a backstop when the custom metric is wrong or stale, while the queue metric drives normal operation.
- Why is `AverageValue` usually the right target type for scaling on queue depth?With `AverageValue` the HPA divides the reported metric by the current replica count before comparing, so the target reads as "backlog per consumer pod". A backlog of 600 with `averageValue: 30` asks for 20 pods, and the relationship stays stable as the fleet grows. With `Value` the raw aggregate is compared, which produces runaway replica demands as the backlog grows.
- What does KEDA give you that a plain HorizontalPodAutoscaler cannot?Two things: a large catalogue of prebuilt scalers for external event sources, so you don't hand-write adapter rules, and genuine scale-to-zero — KEDA activates the workload from zero replicas on the first event, then delegates to a generated HPA above one replica. Vanilla HPA requires minReplicas of at least 1 unless the alpha HPAScaleToZero gate is on.
saying these in an interview costs you the question
- Believing Kubernetes ships an implementation of custom.metrics.k8s.io out of the box
- Using target type Value for a growing backlog and being surprised by huge replica demands
- Autoscaling on latency percentiles and creating a retry/stampede feedback loop
- Assuming custom metrics react instantly, ignoring scrape plus sync lag of tens of seconds
- Not realizing the monitoring stack becomes a dependency of the autoscaling path