skip to content

Event-Driven Autoscaling

A queue consumer's CPU says nothing about how far behind it is, so KEDA scales on the backlog instead - Kafka lag, SQS depth, a PromQL result - and can drop to zero replicas when the queue drains. Scale-to-zero and its cold start are the usual follow-up.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

4

What is KEDA in a Kubernetes cluster, and what does a KEDA ScaledObject declare about the workload it scales?

level: juniorimportance: must knowfreq 62%

answer

  1. backlog, not CPU
  2. operator plus metrics adapter
  3. external.metrics.k8s.io APIService
  4. keda-hpa-<name> owned object
  5. authenticationRef to TriggerAuthentication

basics

~20 s

KEDA is a Kubernetes operator that scales workloads on external event signals such as Kafka lag, queue depth or a Prometheus query. A ScaledObject names the target workload, its replica bounds and its triggers. KEDA turns that into an HPA it owns.

solid answer

~30 s

KEDA is an operator plus a metrics adapter. A `ScaledObject` (`keda.sh/v1alpha1`) points `scaleTargetRef` at a Deployment or StatefulSet and sets `minReplicaCount` (default `0`) and `maxReplicaCount` (default `100`). It lists `triggers`: scalers such as `kafka`, `aws-sqs-queue` or `prometheus`, each with its own metadata like `lagThreshold`. KEDA creates and owns an HPA named `keda-hpa-<name>` with External metrics. The HPA controller fetches those metrics through `external.metrics.k8s.io`, which KEDA's adapter serves. KEDA itself handles scaling between zero and one replica. Credentials come from a `TriggerAuthentication` referenced by `authenticationRef`, not from inline metadata.

code

yaml · 36 lines
yaml
apiVersion: keda.sh/v1alpha1
kind: TriggerAuthentication
metadata:
  name: fraud-kafka-auth
  namespace: fraud
spec:
  secretTargetRef:
    - parameter: sasl
      name: fraud-kafka-creds
      key: sasl
    - parameter: username
      name: fraud-kafka-creds
      key: username
    - parameter: password
      name: fraud-kafka-creds
      key: password
---
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: fraud-rules
  namespace: fraud
spec:
  scaleTargetRef:
    name: fraud-rules-engine
  minReplicaCount: 1
  maxReplicaCount: 18
  triggers:
    - type: kafka
      metadata:
        bootstrapServers: kafka-bootstrap.streaming.svc:9092
        consumerGroup: fraud-rules
        topic: card-auth-events
        lagThreshold: "350"
      authenticationRef:
        name: fraud-kafka-auth

go deeper

for a junior

Say what KEDA is for (scaling on backlog, not CPU) and name the ScaledObject fields: scaleTargetRef, min and max replicas, and triggers with metadata.

for a middle

Explain the pipeline: KEDA generates an owned HPA with External metrics, and the HPA fetches them through external.metrics.k8s.io from KEDA's adapter.

for a senior

Show you know the ownership rules: never hand-edit keda-hpa objects, one autoscaler per workload, and credentials through TriggerAuthentication with RBAC for KEDA itself.

for a principal

Frame KEDA as a platform component: who may create ScaledObjects, which scalers and identity providers you allow, and how its adapter's availability affects every HPA using External metrics.

## What KEDA is **KEDA** (Kubernetes Event-Driven Autoscaling) is an operator you install into a cluster. It adds custom resources, a controller, and a **metrics adapter**. Together these let a workload scale on signals that live *outside* the cluster: the lag of a Kafka consumer group, the depth of an SQS queue, or the result of a Prometheus query. KEDA exists because CPU tells you very little about a queue consumer. Take a fraud-rules engine that reads card-authorisation events. It can sit at 20% CPU while 6,000 events pile up behind it, because each event spends most of its time waiting on I/O. What matters is **how far behind** the consumer is, and that number lives in the broker, not in the pod's cgroup. KEDA does **not** replace the HorizontalPodAutoscaler (HPA). It feeds it. Kubernetes' own replica arithmetic still runs, so KEDA's job is to deliver the right number to it. ## The moving parts - **The KEDA operator** watches `ScaledObject` and `ScaledJob` resources. For each ScaledObject it creates and reconciles an HPA. - **The KEDA metrics adapter** is an extension API server. It registers the `v1beta1.external.metrics.k8s.io` APIService with the aggregation layer. When the HPA controller in kube-controller-manager asks for an External metric, the API server proxies the request to KEDA. KEDA then queries the real source through a **scaler**. - **Scalers** are KEDA's plug-ins, one per event source. Each one turns a source into two answers: a number (the metric) and a yes/no (is the source *active*?). - **Admission webhooks** reject invalid objects. One example is a ScaledObject aimed at a workload that another HPA already manages. ## What a ScaledObject declares A `ScaledObject` (API group `keda.sh`, version `v1alpha1`) is a contract about one scalable workload: | Field | Meaning | Default | |---|---|---| | `scaleTargetRef` | The Deployment, StatefulSet, or any resource with a `/scale` subresource | required | | `minReplicaCount` | The floor. `0` allows scale-to-zero | `0` | | `maxReplicaCount` | The ceiling written into the generated HPA | `100` | | `pollingInterval` | How often KEDA's own loop checks the triggers, in seconds | `30` | | `cooldownPeriod` | How long all triggers must stay inactive before scaling to zero, in seconds | `300` | | `triggers[]` | One or more scalers, each with `type`, `metadata` and an optional `authenticationRef` | at least one | | `advanced.horizontalPodAutoscalerConfig` | Name and `behavior` to pass through to the generated HPA | none | | `fallback` | Replica count to hold if a scaler keeps failing | none | Each trigger's `metadata` is scaler-specific. For example: 1. A **Kafka** trigger names `bootstrapServers`, `consumerGroup`, `topic` and `lagThreshold` (the target lag per replica, default `10`). 2. An **SQS** trigger names `queueURL`, `awsRegion` and `queueLength` (the target messages per replica, default `5`). 3. A **Prometheus** trigger names `serverAddress`, `query` and a required `threshold`. ## The generated HPA and who owns it When you apply a ScaledObject named `fraud-rules`, KEDA creates an HPA named `keda-hpa-fraud-rules`. It carries a controller owner reference back to the ScaledObject and one External metric per trigger. Consequences: - **Do not edit the generated HPA by hand.** KEDA reconciles it back to what the ScaledObject says. Put HPA `behavior` under `spec.advanced.horizontalPodAutoscalerConfig` instead. - **Delete the ScaledObject and the HPA goes with it**, through garbage collection of owned objects. - **One autoscaler per workload.** KEDA's webhook refuses a ScaledObject whose target already has an HPA it does not own. Two HPA-style controllers writing one `/scale` subresource would fight. - **The HPA's `minReplicas` is never 0.** KEDA sets it to `minReplicaCount`, or to `1` when you asked for zero. KEDA's own loop handles the step between zero and one. ## Credentials: TriggerAuthentication Scalers often need secrets such as SASL credentials or cloud credentials. Instead of putting them inline, a trigger points `authenticationRef` at a **`TriggerAuthentication`** (namespaced) or a **`ClusterTriggerAuthentication`**. The object can: - map a scaler parameter to a Secret key with `secretTargetRef` (`parameter`, `name`, `key`); - read a variable from the target's container with `env`; - use workload identity through `podIdentity`, with a `provider` such as `aws` or `gcp`; - or pull the value from an external secret store. Credentials are read by KEDA's own pods, not by your workload. That means the KEDA operator needs permission to read the Secret you reference. ## Where KEDA stops KEDA decides **how many replicas** a workload should have. It does not place pods, add nodes, or make your consumer faster. When the new pods do not fit, a node autoscaler must add capacity. The generated HPA still does its own replica arithmetic and stabilisation.

  • Why can't you just edit the HPA that KEDA generated to change its scale-down behaviour?
    KEDA owns the HPA through a controller owner reference and reconciles it back to what the ScaledObject says, so a hand edit gets overwritten. HPA `behavior`, and a custom HPA name, go under `spec.advanced.horizontalPodAutoscalerConfig` in the ScaledObject. KEDA copies them into the HPA on every reconcile, so the ScaledObject stays the single source of truth.
  • Your Deployment already has a hand-written HorizontalPodAutoscaler. What happens when you apply a KEDA ScaledObject that targets the same Deployment?
    KEDA's validating webhook rejects the ScaledObject with an error saying the workload is already managed by that HPA. Two autoscalers writing one `/scale` subresource would overwrite each other's replica counts. Delete the hand-written HPA first, or move its metrics into the ScaledObject; a `cpu` or `memory` trigger covers resource metrics. KEDA also supports an annotation that adopts an existing HPA.
  • Why does KEDA ask you to reference a TriggerAuthentication instead of putting a password in trigger metadata?
    Trigger metadata is a plain part of the ScaledObject, so anyone who can read that object sees the value. A `TriggerAuthentication` maps scaler parameters to Secret keys, container env, or a workload identity provider, and it can be shared across ScaledObjects. KEDA's own pods do the reading, so the KEDA operator needs RBAC to read the referenced Secret.

KEDA is like a dispatcher who watches the queue at the counter and tells the shift manager, the HPA, how many clerks to staff. The manager still does the rostering; the dispatcher only supplies the number.

saying these in an interview costs you the question

  • KEDA replaces the HorizontalPodAutoscaler with its own replica controller
  • KEDA scales pods by editing their CPU requests
  • You tune scaling by editing the generated keda-hpa object directly
  • KEDA and a separate hand-written HPA can safely target one Deployment
  • Scaler passwords belong in the ScaledObject trigger metadata
  • KEDA adds nodes when the new replicas do not fit
open as a page

How does a KEDA ScaledObject take a Kubernetes Deployment to zero replicas and back, and which fields govern each transition?

level: middleimportance: must knowfreq 58%

basics

~20 s

KEDA's own loop handles the zero boundary. It activates the target to one replica when a trigger exceeds its activation threshold, and returns it to zero after cooldownPeriod with every trigger inactive. The generated HPA scales between one and maxReplicaCount.

open as a page

A KEDA-scaled fraud-rules engine on Kubernetes sits at zero replicas, and after a Kafka burst the first event waits 47 seconds before processing. Where does that time go, and what would you change?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The delay is a chain: up to one pollingInterval before KEDA notices, then scheduling, start-up and readiness, then the consumer joining its group. Measure each stage; for a tight budget, keep minReplicaCount at 1 instead.

open as a page

When would you use a KEDA ScaledJob instead of a ScaledObject to process a queue on Kubernetes, and how do they scale differently?

level: middleimportance: nice to knowfreq 28%

basics

~10 s

A ScaledObject scales a long-running Deployment through a generated HPA. A ScaledJob creates Kubernetes Jobs that run to completion. Use a ScaledJob for long tasks that scale-in must never interrupt.

open as a page