skip to content

How does a KEDA ScaledObject take a Kubernetes Deployment to zero replicas and back, and which fields govern each transition?

level: middleimportance: must knowfreq 58%

answer

  1. two owners, split at one
  2. activation versus target threshold
  3. strictly greater than activation value
  4. 30s poll, 300s cooldown
  5. HPA minReplicas never zero

basics

~20 s

KEDA's own loop handles the zero boundary. It activates the target to one replica when a trigger exceeds its activation threshold, and returns it to zero after cooldownPeriod with every trigger inactive. The generated HPA scales between one and maxReplicaCount.

solid answer

~40 s

With `minReplicaCount: 0`, the generated HPA still gets `minReplicas: 1`, because an HPA's ratio arithmetic cannot start from zero pods. KEDA polls each trigger every `pollingInterval` (default 30s). When a value is strictly above its activation threshold (such as `activationLagThreshold` or `activationQueueLength`, default 0), KEDA scales the target from 0 to 1 and records `lastActiveTime`. From 1 to N, the HPA drives scaling using the scaler's target value, like `lagThreshold`, on its own sync period. When no trigger has been active for `cooldownPeriod` (default 300s), KEDA writes zero, or `idleReplicaCount`, to the scale subresource. `cooldownPeriod` does not slow ordinary scale-in; that is HPA `behavior`.

code

yaml · 20 lines
yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: fraud-rules
  namespace: fraud
spec:
  scaleTargetRef:
    name: fraud-rules-engine
  minReplicaCount: 0
  maxReplicaCount: 18
  pollingInterval: 10
  cooldownPeriod: 420
  triggers:
    - type: kafka
      metadata:
        bootstrapServers: kafka-bootstrap.streaming.svc:9092
        consumerGroup: fraud-rules
        topic: card-auth-events
        lagThreshold: "350"
        activationLagThreshold: "40"

go deeper

for a junior

Remember the split: KEDA handles zero to one and back, and the HPA handles one to many. Know the defaults: 30-second polling and 300-second cooldown.

for a middle

Explain why the HPA cannot leave zero, and the difference between activation thresholds and scaling targets. Also explain why cooldownPeriod governs only the final step to zero.

for a senior

Reason about when pollingInterval matters at all and what lowering it costs in source queries. Use pause annotations and fallback to keep control during incidents.

for a principal

Decide per workload whether idle savings justify cold-start latency, and set a default minReplicaCount policy that product owners have agreed to.

## Two controllers, two ranges A KEDA `ScaledObject` with `minReplicaCount: 0` splits scaling across two controllers. | Range | Who acts | What it decides with | |---|---|---| | **0 -> 1** (activation) | KEDA's operator loop | each scaler's *IsActive* answer | | **1 -> N -> 1** (scaling) | the generated HPA in kube-controller-manager | the scaler's *metric value* through `external.metrics.k8s.io` | | **1 -> 0** (deactivation) | KEDA's operator loop | all triggers inactive for `cooldownPeriod` | The split exists because a plain HPA stops at one replica. Its ratio arithmetic has nothing to multiply when zero pods exist. KEDA therefore writes the generated HPA's `minReplicas` as `max(minReplicaCount, 1)` and handles the zero boundary itself through the target's `/scale` subresource. ## Activation: leaving zero Every **`pollingInterval`** seconds (default **30**), KEDA's loop asks each scaler for its value and its *activity*. A scaler is **active** when its value is **strictly greater than** its activation threshold: - Kafka: `activationLagThreshold`, default `0`; - SQS: `activationQueueLength`, default `0`; - Prometheus: `activationThreshold`, default `0`. With the defaults, a single unconsumed message activates the workload. When any trigger is active and the target sits at zero, KEDA scales it to `minReplicaCount` if that is above zero, otherwise to **1**. KEDA also stamps `status.lastActiveTime` and records an event. The activation threshold is separate from the **scaling target** (`lagThreshold`, `queueLength`, `threshold`), and the two are easy to confuse: 1. The **activation** value answers "should anything run at all?" 2. The **target** value answers "how much backlog should each replica carry?" Example: a fraud-rules engine with `lagThreshold: 350` and `activationLagThreshold: 40`. A lag of 37 leaves it at zero. A lag of 41 activates one pod. A lag of 5,250 drives the HPA towards 15 replicas. ## Scaling between 1 and N Once one replica runs, the HPA takes over. On its own sync period (kube-controller-manager's `--horizontal-pod-autoscaler-sync-period`, default 15 seconds), it asks the metrics API for the External metric. KEDA's adapter answers by querying the scaler. `pollingInterval` has **no** say in this range. In fact, when `minReplicaCount` is above zero, no `idleReplicaCount` is set, and no trigger uses cached metrics, KEDA's webhook warns that `pollingInterval` has no effect on scaling. It still sets how often KEDA refreshes the ScaledObject's status conditions and events. The HPA's stabilisation windows and tolerance still apply here. Tune them through `spec.advanced.horizontalPodAutoscalerConfig.behavior`. ## Deactivation: back to zero On each poll where **no** trigger is active, KEDA compares the current time with `lastActiveTime + cooldownPeriod` (default **300** seconds): - **Still inside the window:** KEDA leaves the replicas alone and sets the Active condition to false with reason `ScalerCooldown`. - **Past the window:** KEDA writes `minReplicaCount` (0) to the scale subresource. If `idleReplicaCount` is set, which must be less than `minReplicaCount`, KEDA writes that value instead. Two details are easy to get wrong: - **`cooldownPeriod` only governs the step to zero** (or to idle). It does not slow scale-in from 12 replicas to 3. That is HPA `behavior.scaleDown`. - **`initialCooldownPeriod`** (default `0`) delays the first scale-to-zero after the ScaledObject is created. It protects a freshly deployed workload that has not yet been active. ## Timing arithmetic The worst-case delay before KEDA even notices a burst at zero is about one `pollingInterval`. After that come scheduling, image pull, container start, readiness, and the consumer's own start-up (for Kafka, joining the consumer group). Lowering `pollingInterval` to 5 seconds cuts the first term. The cost is six times as many broker or API queries per ScaledObject. ## Pausing and failure - **Pause annotations.** `autoscaling.keda.sh/paused-replicas: "<n>"` pins the workload at `n` replicas. `autoscaling.keda.sh/paused` freezes scaling. `paused-scale-in` and `paused-scale-out` block only one direction. The scale-in pause also blocks the step to zero. - **`fallback`.** Set `failureThreshold` and `replicas` so that a scaler that keeps erroring holds a known replica count, instead of freezing at whatever count was last computed. ## Choosing zero Scale-to-zero suits bursty, delay-tolerant consumers: nightly reconciliations and low-traffic tenants. It fits badly when the first event after an idle spell has a latency budget smaller than your cold start. In that case keep `minReplicaCount: 1` and let KEDA scale from there.

  • A KEDA ScaledObject has minReplicaCount: 3 and pollingInterval: 5, and the team expects faster scale-up. Why does nothing change?
    With `minReplicaCount` above zero, no `idleReplicaCount` and no cached metrics, KEDA's loop never moves the replica count. The generated HPA does all the scaling on its own sync period, 15 seconds by default, and asks KEDA's adapter for the value on demand. KEDA's webhook warns that `pollingInterval` is not relevant here. Faster scale-up comes from HPA `behavior`.
  • What does the KEDA ScaledObject field idleReplicaCount change about scale-in?
    Instead of dropping to `minReplicaCount` when every trigger has been inactive for `cooldownPeriod`, KEDA sets the target to `idleReplicaCount`. That value must be below `minReplicaCount`. In practice it is used with `idleReplicaCount: 0` and a `minReplicaCount` above 1: the workload idles at zero, and on activation it jumps straight to the minimum rather than to one.
  • How do you stop KEDA from scaling a Deployment to zero during an incident without deleting the ScaledObject?
    Annotate the ScaledObject. `autoscaling.keda.sh/paused-replicas: "4"` pins the target at four replicas. `autoscaling.keda.sh/paused-scale-in: "true"` blocks scale-in, including the step to zero, while still allowing scale-out. Remove the annotation to resume. Deleting the ScaledObject would also delete its generated HPA.

It is like a shop that is closed overnight: a guard glances at the door every half hour and opens up when someone is waiting, while the floor manager handles how many tills are open during the day.

saying these in an interview costs you the question

  • The generated HPA itself scales the Deployment from zero to one
  • cooldownPeriod slows every scale-in step, like an HPA stabilization window
  • pollingInterval controls how fast the HPA reacts between 1 and N replicas
  • The activation threshold and lagThreshold are the same setting
  • A trigger exactly at its activation threshold counts as active
  • KEDA writes minReplicas: 0 into the generated HPA