skip to content

When would you use a KEDA ScaledJob instead of a ScaledObject to process a queue on Kubernetes, and how do they scale differently?

level: middleimportance: nice to knowfreq 28%

answer

  1. what gets scaled
  2. scale-in kills pods, not Jobs
  3. no HPA for ScaledJob
  4. running Jobs deducted from max
  5. rollout immediate deletes Jobs

basics

~10 s

A ScaledObject scales a long-running Deployment through a generated HPA. A ScaledJob creates Kubernetes Jobs that run to completion. Use a ScaledJob for long tasks that scale-in must never interrupt.

solid answer

~30 s

A `ScaledObject` resizes an existing workload: the generated HPA scales it between 1 and N, and KEDA handles zero. On scale-in the HPA removes pods, so an in-flight task can be cut short. A `ScaledJob` has no HPA. Every `pollingInterval`, KEDA reads the backlog and creates new Jobs from `jobTargetRef`, up to `maxReplicaCount` (default 100). The `scalingStrategy` decides how running Jobs are deducted. Jobs are never scaled in; they run until they complete. So I use a ScaledJob for long, independent units, such as a 25-minute re-scoring task, and write the worker to take one item, finish it and exit.

code

yaml · 29 lines
yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledJob
metadata:
  name: fraud-rescore
  namespace: fraud
spec:
  jobTargetRef:
    backoffLimit: 2
    activeDeadlineSeconds: 2700
    template:
      spec:
        restartPolicy: Never
        containers:
          - name: rescore
            image: registry.example.com/fraud/rescore:4.12.3
  pollingInterval: 15
  maxReplicaCount: 23
  successfulJobsHistoryLimit: 5
  failedJobsHistoryLimit: 9
  rollout:
    strategy: gradual
  scalingStrategy:
    strategy: default
  triggers:
    - type: prometheus
      metadata:
        serverAddress: http://prometheus-operated.monitoring.svc:9090
        query: sum(fraud_rescore_pending_tasks)
        threshold: "1"

go deeper

for a junior

Remember the distinction: a ScaledObject resizes a Deployment, and a ScaledJob creates Jobs that run to completion.

for a middle

Explain why scale-in can interrupt a Deployment's work but never a ScaledJob's, and how maxReplicaCount and running Jobs bound new Job creation.

for a senior

Pick the scalingStrategy by what the scaler actually counts, set rollout strategy to gradual for long tasks, and bound Jobs with deadlines and history limits.

for a principal

Decide which workloads may run as bursty per-task Jobs on shared nodes, and how their peak count fits quota and node headroom.

## Two ways to turn a backlog into work KEDA offers two custom resources for event-driven work. Both use the same scalers and the same `triggers` block. What differs is **what gets scaled** and **how work ends**. | | `ScaledObject` | `ScaledJob` | |---|---|---| | Scales | an existing Deployment, StatefulSet or any `/scale` resource | Kubernetes **Jobs** that KEDA creates from `jobTargetRef` | | Worker lifetime | long-running; consumes message after message | usually one Job per unit of work, runs to completion | | Replica arithmetic | a generated HPA (`keda-hpa-<name>`) plus KEDA's 0<->1 step | KEDA's own loop, every `pollingInterval`; no HPA | | Scale-in | HPA removes pods, which receive SIGTERM | nothing is removed; Jobs finish on their own | | Main risk | a pod killed mid-message during scale-in | too many Jobs, or Jobs deleted by a spec update | ## When a ScaledObject fits Choose a `ScaledObject` when the worker is a normal service loop: it fetches, processes, commits, and repeats. Each message takes milliseconds to seconds, and the process can stop cleanly within its termination grace period. A fraud-rules engine scoring card authorisations in about 40 ms each fits well. Losing a pod to scale-in costs at most a few in-flight events, which are redelivered. ## When a ScaledJob fits Choose a `ScaledJob` when **one unit of work is long and must not be interrupted by scaling**. Examples: re-scoring a merchant's 90-day history against a new rule set, rendering a video, or running a batch model. A scale-in decision can kill a Deployment pod halfway through a 25-minute task. A Job is not a scale target that anything shrinks, so it keeps running until it completes or fails. Other reasons to use a ScaledJob: - **Per-task isolation.** Each Job gets a clean pod, with its own resource requests and restart and backoff rules from the Job spec. - **Burst without a standing Deployment.** Nothing exists between bursts; KEDA creates Jobs only while the backlog exists. ## How a ScaledJob scales On every `pollingInterval` (default 30 seconds), KEDA asks the scalers for the queue length and decides how many **new** Jobs to create. It is bounded by `maxReplicaCount` (default 100), which caps the number of Jobs KEDA runs at once. `spec.scalingStrategy.strategy` decides how running and pending Jobs count against that: 1. **`default`**: the effective ceiling is `maxReplicaCount` minus the number of running Jobs. 2. **`custom`**: adds `customScalingQueueLengthDeduction` and `customScalingRunningJobPercentage` to tune that deduction. 3. **`accurate`** and **`eager`**: other deduction rules for queues where the scaler's count does or does not already exclude in-flight work. Choose one only after checking what your scaler reports. The strategy matters because of double counting. Some queues still count messages that a running Job has taken. SQS, for example, counts in-flight messages by default through the scaler's `scaleOnInFlight`. Without the right strategy, KEDA may create a new Job for work that is already being done. Other ScaledJob fields to know: - `jobTargetRef` is a full **JobSpec**: `template`, `parallelism`, `completions`, `backoffLimit`, `activeDeadlineSeconds`. - `successfulJobsHistoryLimit` and `failedJobsHistoryLimit` keep finished Jobs from piling up in the namespace. - `rollout.strategy` is `immediate` or `gradual`. With `immediate`, updating the ScaledJob **deletes the running Jobs** it created under the previous spec. With `gradual`, they run to completion and only new Jobs use the new spec. ## The worker's contract changes too A ScaledObject worker loops forever. A ScaledJob worker should **take one unit of work, finish it, and exit 0**. A Job pod that loops forever never completes, so the ScaledJob keeps counting it as running, and the backlog stalls behind `maxReplicaCount`. Its exit code also feeds the Job controller's failure handling (`backoffLimit`). ## Choosing - Short, uniform messages and a long-lived consumer: **ScaledObject**. - Long, independent tasks that must not be cut short: **ScaledJob**. - Long tasks in a Deployment you cannot rewrite: a ScaledObject with a generous `terminationGracePeriodSeconds` and scale-in `behavior` can work. It is a mitigation, not the same guarantee.

  • You update the container image in a KEDA ScaledJob while 11 re-scoring Jobs are running. What happens to them?
    It depends on `spec.rollout.strategy`. With the default, `immediate`, KEDA deletes the Jobs it created under the previous spec, so all 11 tasks die mid-run. With `gradual`, the running Jobs finish on the old image and only new Jobs use the new one. For long tasks, set `gradual` explicitly.
  • Why must a KEDA ScaledJob worker exit after finishing its task instead of looping on the queue?
    A Job counts as running until its pod completes. A worker that loops never completes, so KEDA keeps deducting it from `maxReplicaCount`. Once enough such Jobs exist, KEDA creates no new ones even though the backlog grows. Exiting 0 also lets the Job controller record success, and a non-zero exit feeds `backoffLimit`.

saying these in an interview costs you the question

  • A ScaledJob generates an HPA just like a ScaledObject
  • KEDA deletes running Jobs when the queue drains
  • A ScaledJob worker should loop forever like a Deployment consumer
  • Updating a ScaledJob never affects Jobs already running
  • ScaledObject scale-in waits for in-flight tasks to finish