skip to content

What does concurrency scaling do on Amazon Redshift, and what must be true for it to kick in?

level: seniorimportance: should knowfreq 58%

answer

  1. extra compute only during the burst
  2. the trigger is a backlog, not slowness
  3. one parameter is your cost ceiling
  4. enabled per queue, not per cluster
  5. credits accrue, overflow is per-second

basics

~20 s

Concurrency scaling adds transient clusters when queries start queueing, routes eligible queued queries to them, and shuts them down after the burst. It needs concurrency scaling enabled on the WLM queue, actual queueing, and an eligible query — and it never speeds up a single running query.

solid answer

~50 s

When a WLM queue with concurrency scaling set to `auto` starts queueing, Redshift spins up a transient cluster, routes eligible queued queries to it against the same data, and returns results that are consistent with the main cluster; the extra cluster shuts down when the burst passes. Three conditions must all hold: the queue must have concurrency scaling enabled, queries must actually be **queueing** (it is a backlog remedy, not a speed remedy), and the query must be of an eligible shape — read queries qualify and a documented subset of writes does, but not everything. The cluster-level `max_concurrency_scaling_clusters` parameter caps how many transient clusters can run. Billing works on accrued credits that build up as the main cluster runs, with per-second charges beyond them, so an unbounded cap plus a runaway dashboard is a real cost incident. Confirm what actually ran there via `concurrency_scaling_status` on the query record and the `SVCS_` views.

go deeper

for a junior

Know that Redshift can temporarily add extra clusters when too many queries arrive at once, and that this handles bursts of concurrency rather than making one query faster.

for a middle

Explain the trigger and the mechanism: it is enabled per WLM queue, starts when that queue queues, runs eligible queries on transient clusters over the same data, and is capped by a cluster-level parameter.

for a senior

Show that you verify it actually engaged from the query records and cross-cluster views, and that you know when it is the wrong remedy — a slow query or a permanently undersized cluster.

for a principal

Own the cost boundary: which workloads deserve burst capacity, what the maximum-clusters cap should be, when constant burst usage means the cluster should be resized or the workload moved to its own compute.

## The problem it solves A provisioned Amazon Redshift cluster is sized for a steady workload, but analytics traffic is spiky: at 9am every dashboard in the company refreshes at once, and a queue sized for the average starts backing up. The historical fix was to size the cluster for the peak and pay for idle capacity the rest of the day. Concurrency scaling is the alternative — extra compute appears for the burst and disappears afterwards. ## How it works Concurrency scaling is enabled **per WLM queue**, by setting that queue's `concurrency_scaling` property to `auto` (the alternative is `off`). When queries begin queueing in such a queue, Redshift starts one or more transient clusters. Eligible queued queries are routed to a transient cluster, which reads the same data as the main cluster and returns results consistent with it — the user sees an ordinary query result and has no idea where it ran. When the backlog clears, the transient clusters are released. How many can run at once is bounded by the cluster-level parameter `max_concurrency_scaling_clusters`. That parameter is your cost ceiling and it is the one number to set deliberately before enabling the feature broadly. ## The three preconditions Candidates most often get this wrong by asserting that concurrency scaling makes queries faster. All three of these must hold: 1. **The queue has it enabled.** A queue left at `off` never uses it, no matter how long its backlog. If BI traffic is landing in the default queue and only the ETL queue has it on, nothing happens. 2. **Queries are queueing.** The trigger is a backlog. A single long-running query on an otherwise idle cluster will never be moved to a transient cluster, and it would not run any faster if it were — a transient cluster adds *concurrency*, not per-query speed. 3. **The query is eligible.** Read queries are the core case; AWS has extended eligibility over time to cover a documented set of write operations as well. Not every statement qualifies, and rather than reciting a list from memory it is better to say that eligibility is documented and that you verify empirically from the query records which queries actually ran on a transient cluster. ## Confirming it actually happened The query record carries a `concurrency_scaling_status` field indicating whether a query ran on a concurrency scaling cluster, and the `SVCS_` family of system views (the cross-cluster counterparts of the `SVL_`/`STL_` views, such as `SVCS_QUERY_SUMMARY` and `SVCS_ALERT_EVENT_LOG`) covers activity on both the main and transient clusters. On current clusters the `SYS_` views also span transient-cluster activity. This matters during triage: if you look only at main-cluster STL tables, work that ran on a transient cluster is missing from your analysis and the numbers will not add up. CloudWatch also publishes concurrency-scaling activity and usage metrics. ## The cost model Concurrency scaling usage accrues against credits that build up while the main cluster runs; usage beyond the accrued credits is billed per second of transient cluster time. Prices and credit caps change, so quote the mechanism rather than a number: *credits accrue with main-cluster runtime, are capped, and overflow is billed per second*. The operational consequences are what matter in an interview: - Set `max_concurrency_scaling_clusters` deliberately. Left high, a badly written dashboard that queues constantly can spin up capacity all day. - Enable it on the queues that serve latency-sensitive interactive users, not on the ETL queue where queueing is harmless and delay is free. - Watch the usage metrics and, if cost is a concern, pair the feature with a usage limit on concurrency scaling so the account gets an alarm or a hard stop. ## When it is the wrong tool Concurrency scaling absorbs bursts. It does not fix: - **A slow query.** If the split of elapsed time is nearly all execution and no queue, nothing here helps; fix the plan, the statistics, the distribution and sort keys. - **A structurally undersized cluster.** If the queue is backed up all day rather than for twenty minutes at 9am, you are paying for transient capacity continuously and should resize, or move the workload to its own compute. - **Isolation between teams.** Transient clusters follow a queue's backlog; they do not give a team a guaranteed, independent compute pool. That is a different design conversation. ## The interview shape A strong answer states the trigger (queueing on an enabled queue), the mechanism (transient clusters over the same data, consistent results), the bound (`max_concurrency_scaling_clusters`), the billing shape (credits then per-second), and — most importantly — the negative claim that it does nothing for a single slow query.

  • You enabled concurrency scaling but queries still queue for minutes. What are the likely reasons?
    Check three things in order: whether the traffic is landing in the queue you enabled it on rather than the default queue, whether the cluster-level `max_concurrency_scaling_clusters` cap is already reached, and whether the queries are of an eligible shape. Then confirm empirically with the query records' concurrency scaling status — if nothing ever ran on a transient cluster, it is a routing or eligibility problem, not a capacity one.
  • Why do main-cluster STL views sometimes fail to explain a query's behaviour once concurrency scaling is on?
    Because the query may not have run on the main cluster. Activity on transient clusters is exposed through the `SVCS_` cross-cluster views and the newer `SYS_` views; looking only at `STL_`/`SVL_` tables silently omits it. During triage, start from a view that spans both so your query counts and timings actually reconcile.
  • How would you stop concurrency scaling from becoming an unbounded cost line?
    Bound it deliberately: set `max_concurrency_scaling_clusters` to a number you are willing to pay for, enable the feature only on queues serving latency-sensitive interactive users, add a usage limit on concurrency scaling so the account alerts or stops at a threshold, and monitor usage metrics. A dashboard that queues constantly is a query problem to fix, not capacity to buy indefinitely.

saying these in an interview costs you the question

  • Saying concurrency scaling makes an individual query run faster
  • Assuming it is on cluster-wide rather than per queue
  • Forgetting that queueing is the trigger
  • Ignoring the maximum-clusters cap as a cost control
  • Analysing only main-cluster STL views once it is enabled

context