skip to content

In docker-selenium's Helm chart, when is on-demand node scaling worth what it obliges you to run?

level: principalimportance: nice to knowfreq 42%

answer

  1. the shipped default is a fixed pool
  2. it grows because something counted the queue
  3. jobs by default, deployments by choice
  4. a miscounting scaler runs away
  5. cold start joins the critical path

basics

~20 s

When demand is spiky and idle nodes are waste. The chart ships autoscaling disabled, so a default install is a fixed pool; enabling it installs KEDA and hands you a scaler, a metric to keep correct, and cold starts.

solid answer

~40 s

`docker-selenium`'s chart ships `autoscaling.enabled: false`, so an install you have not touched is a **fixed pool** sized by hand. Turning it on installs KEDA, or attaches to one you already run via `enableWithExistingKEDA`. The default `autoscaling.scalingType` is `job`, which renders a KEDA `ScaledJob` whose `jobTargetRef` sets `parallelism: 1` and `completions: 1` — each scaled unit is a Pod that serves work and exits, the same disposable shape a container grid gets for free. Setting `scalingType: deployment` renders a `ScaledObject` over long-lived node Deployments instead, which is why the chart carries `terminationGracePeriodSeconds` and `deregisterLifecycle` for that mode only: a long-lived node must be drained, a Job simply finishes. Scale on demand when demand is bursty and idle capacity is waste; keep a fixed pool when the shape of the load is known.

code

yaml · 9 lines
yaml
autoscaling:
  enabled: true          # the chart's own default is false
  scalingType: job       # default; renders a KEDA ScaledJob
  scaledJobOptions:
    scalingStrategy:
      strategy: accurate
    jobTargetRef:
      parallelism: 1
      completions: 1

go deeper

for a junior

Know that a grid can be a fixed pool or one that grows with demand, and that the chart's default is fixed. You are not expected to configure a scaler, only to know which kind you are pointing your tests at.

for a middle

Be able to say what has to be true for a grid to grow: something reads how much work is waiting and asks the cluster for more capacity. Name the trade-off against cold start on a session that arrives to an empty pool.

for a senior

Expect to diagnose a scaler behaving badly. Know that a metric which double-counts running work creates nodes without bound, and that trigger metadata has to match the browser and platform you intended to scale.

for a principal

This is your call to make. Justify a fixed pool when demand has a known shape, put a number on what idle capacity costs against what the scaler costs to operate, and say who owns it when it is wrong.

A container grid can be a fixed pool that somebody sized, or a pool that grows because something counted the work waiting. `docker-selenium`'s Helm chart supports both, and the interesting part of the question is not the switch but what the switch obliges you to own. ## What the chart actually ships The chart's `autoscaling.enabled` value defaults to **false**. A plain install is therefore a fixed set of node replicas that you chose; a queue that grows produces waiting, not nodes. Enabling autoscaling installs KEDA as part of the release, or `autoscaling.enableWithExistingKEDA` attaches to an installation you already operate — a distinction that matters the moment a platform team owns KEDA and does not want a test grid installing its own. ## Two scaling shapes, and the default is not the obvious one `autoscaling.scalingType` defaults to `job`. | | `scalingType: job` (default) | `scalingType: deployment` | |---|---|---| | KEDA object rendered | `ScaledJob` | `ScaledObject` | | What is scaled | Kubernetes Jobs | A node Deployment's replicas | | Lifetime of a node | Serves its work, then exits | Long-lived, reused | | Shutdown concern | The Job simply completes | Must be drained before termination | | Chart settings that apply | `scaledJobOptions` | `terminationGracePeriodSeconds`, `deregisterLifecycle` | The chart's `jobTargetRef` sets `parallelism: 1` and `completions: 1`, so a scaled unit in job mode is a Pod that does its work and finishes. That is the same disposable-per-session instinct a container grid has natively, expressed in Kubernetes objects. Deployment mode trades it for warm nodes that must be told to stop accepting work before they are killed, which is exactly what the grace period and the deregister lifecycle hook exist for. Note also that autoscaling and the drain-after-work setting solve adjacent problems: `SE_DRAIN_AFTER_SESSION_COUNT` detaches a node from the grid once it has served its allotted work, so a long-lived node can still be made replaceable. ## What enabling it obliges you to run - **A scaler, and its credentials.** Something has to read the grid's live state to know how much work is waiting, and it needs an address and, where the grid is protected, an authenticated way in. - **A metric that matches your grid.** The scaler's trigger metadata has to line up with the browser and platform the pool actually serves. Get it wrong and you scale the wrong pool, or nothing at all, while the queue looks perfectly healthy from the outside. - **Correct counting of work in progress.** The chart's own documentation records that combining a strategy which already deducts pending and running work with a metric that also counts ongoing sessions double-counts, and the result is unbounded node creation. A scaler that miscounts does not scale badly; it runs away. - **Cold start on the critical path.** A session arriving at an empty pool now waits for a Pod to be scheduled, an image to be present, and a node to register before anything happens. - **A blast radius in the cluster.** The grid can now consume cluster capacity in response to traffic, which makes it a neighbour of every other workload on those machines. ## When it is worth it Scale on demand when the load is **bursty and unpredictable** and idle capacity is genuinely wasted: a grid that sits untouched for most of the day and then absorbs a rush after a merge queue drains. Scale on demand when the peak is far above the median, because that is precisely the ratio a fixed pool has to pay for continuously. Keep the fixed pool when: - The load is a known nightly window of known width. You are not saving anything by discovering its size again every night. - Session start-up latency is already the dominant complaint. Adding a cold pool underneath it makes the worst case worse. - Nobody owns the cluster. An autoscaled grid is a running system with its own failure modes, and it needs someone whose job includes noticing them. - The grid is small enough that the whole pool costs less attention than the scaler would. ## How to decide, concretely 1. Measure the actual shape of demand over a fortnight — peak, median, and how long the peak lasts. 2. Price the fixed pool at the peak, and compare it with the median plus the operational cost of the scaler. 3. Decide whether cold start is acceptable for the slowest tolerable session, and set a floor of warm nodes if it is not. 4. Choose the scaling shape from how your nodes end: disposable work units, or long-lived nodes that must be drained. 5. Write down what happens when the scaler is wrong in both directions, because it will be. ## What an interviewer is listening for - That you know the chart ships autoscaling off, so "we use the chart" does not mean "we autoscale". - That growth is driven by counting waiting work, and that the counter is now a component you operate. - That you can argue for a fixed pool when the demand shape is known, instead of treating autoscaling as strictly better.

  • Why does deployment mode need a grace period and a pre-stop hook when job mode does not?
    Because a long-lived node can be holding a live session when the scaler decides to shrink. The pre-stop hook deregisters it from the grid so no new session lands on it, and the grace period gives the running session time to end before the Pod is killed. A Job in the default mode has no such problem: it finishes its work and exits on its own terms.
  • What is the first thing you check when an autoscaled grid keeps creating nodes and never settles?
    Whether the metric is double-counting work in progress. The chart documents that pairing a deduction strategy which already subtracts pending and running units with a metric that also counts ongoing sessions produces runaway node creation. After that, check that the trigger's browser and platform metadata match the pool you actually meant to grow.

saying these in an interview costs you the question

  • Assumes installing the chart gives you autoscaling
  • Thinks the default scaling shape is a long-lived deployment
  • Cannot say what the scaler counts to decide to grow
  • Ignores cold start when the pool is allowed to empty
  • Treats autoscaling as strictly better than a sized pool