skip to content

WLM & Concurrency Scaling

Workload management decides which query runs now, which waits, and which gets killed, while concurrency scaling adds transient clusters at peak. Asked to check whether I can keep a mixed BI-plus-ETL cluster responsive under load.

part ofAmazon Redshiftoverview, primer and where to startread it →
on this pageshow

questions

6

What is the difference between manual and automatic WLM in Amazon Redshift?

level: middleimportance: must knowfreq 70%

answer

  1. two modes, only one is recommended
  2. who runs now, who waits
  3. one mode makes you do the memory arithmetic
  4. the other gives you priority instead of numbers
  5. auto_wlm true, priority normal

basics

~20 s

Manual WLM makes you define queues with a fixed concurrency level and a memory percentage each. Automatic WLM lets Redshift decide concurrency and per-query memory from the query's estimated needs, and you steer it only with query priority. AWS recommends automatic.

solid answer

~40 s

Redshift's workload manager decides which query runs now and which waits. With **manual WLM** you define up to eight queues in the `wlm_json_configuration` parameter, each with `query_concurrency` (how many slots) and `memory_percent_to_use` (its share of cluster memory); a queue's memory is split evenly across its slots, so concurrency and per-query memory are locked together and you tune them by hand. With **automatic WLM** (`"auto_wlm": true`) Redshift sets concurrency and allocates memory per query from its predicted resource needs, admitting more small queries and fewer large ones; the only lever you keep is `priority` (lowest through critical) on each queue, plus query monitoring rules. Routing is the same in both: a query matches a queue by the submitting user's group or by its `query_group` label, first match wins, otherwise the default queue.

go deeper

for a junior

Be able to say that Redshift queues queries rather than running everything at once, and that a query is routed to a queue by the user's group or a query_group label.

for a middle

Explain the mechanics: manual WLM splits a queue's memory percentage evenly across its concurrency slots, while automatic WLM sizes memory per query and exposes priority instead. Know where the configuration lives.

for a senior

Show judgment about when manual WLM is still worth its maintenance cost — typically a load or ETL path that must never spill — and how you would migrate an existing manual configuration to automatic without a surprise.

for a principal

Own the position that hand-tuned queue arithmetic is a maintenance liability that drifts as workloads change, and be able to argue when a guaranteed memory reservation justifies keeping it anyway.

## What WLM is actually for An Amazon Redshift cluster has a fixed amount of memory and a fixed number of slices. If every submitted query started immediately, each would get a sliver of memory, spill its hash tables and sorts to disk, and everything would run slowly together. Workload management (WLM) is Redshift's admission-control layer: it decides how many queries run concurrently, how much memory each gets, and which ones wait in line. Nothing about WLM makes an individual query faster — it decides *who runs now*. ## Manual WLM Manual WLM is configured as JSON in the cluster parameter group's `wlm_json_configuration` setting. You declare up to eight user-defined queues (plus a reserved superuser queue for administrative work). Each queue carries: - `query_concurrency` — the number of **slots**, i.e. how many queries may run in that queue at once. - `memory_percent_to_use` — the queue's share of the cluster's query memory. - optional `user_group` / `query_group` lists that decide which queries land there. The critical mechanic is that a queue's memory is divided **evenly** across its slots. A queue with 50% of memory and a concurrency of 10 gives each running query 5% of cluster memory. Raise concurrency and every query gets less; lower it and queries wait longer but each gets more room. That coupling is the whole tuning problem: too much concurrency causes disk spill, too little causes queueing. A session can temporarily claim several slots with `SET wlm_query_slot_count`, which is the standard trick for a big one-off maintenance or ETL statement. ```sql -- route this session's work to a labelled queue SET query_group TO 'etl'; ``` ## Automatic WLM Automatic WLM is enabled by setting `"auto_wlm": true` in the same JSON. Redshift then ignores your concurrency and memory numbers and manages both itself: it estimates each query's memory footprint and admits as many as fit, so a burst of cheap dashboard queries can run at high concurrency while a memory-hungry aggregation is given a large allocation and run at low concurrency. The lever you keep is **query priority**, set per queue as `lowest`, `low`, `normal`, `high`, `highest` or `critical`. Priority influences which queued query is admitted next and how resources are shared between running queries; a low-priority ETL job yields to high-priority dashboard traffic rather than being killed. AWS recommends automatic WLM for most clusters, and it is what new clusters generally start with. Manual WLM survives mainly where a team needs a hard, predictable memory reservation for one workload — for example a load process that must never spill regardless of what else the cluster is doing. ## Routing is identical in both modes A query is assigned to a queue by walking the queue list in order and taking the **first match**: 1. If the submitting user belongs to a queue's `user_group`, it goes there. 2. Otherwise, if the session has run `SET query_group TO 'label'` and that label matches a queue's `query_group`, it goes there. 3. Otherwise it lands in the default queue — the last queue in the list. Order matters, and wildcards (`user_group_wild_card`, `query_group_wild_card`) let you match by prefix. A common misconfiguration is putting a broad matching queue first so that everything falls into it. ## Short query acceleration Both modes can enable short query acceleration (`"short_query_queue": true`), which sends queries predicted to be brief into a dedicated space so they do not sit behind a long-running report. With automatic WLM the prediction and the cut-off are managed for you; with manual WLM you can set a maximum run time for what counts as short. ## What changes with Redshift Serverless Redshift Serverless has no WLM queues at all — there is no `wlm_json_configuration`. Capacity is expressed as base and maximum RPUs that scale automatically, with usage limits and query monitoring rules configured on the workgroup. If an interviewer asks about WLM tuning, they are asking about a provisioned cluster. ## How to answer in an interview Say what WLM controls (admission, concurrency, memory), then contrast the two modes on the one axis that matters: manual couples concurrency and memory in numbers you own, automatic decouples them and hands you priority instead. Finish with the operational point — automatic WLM is the default recommendation, and the reason to keep manual is a workload that needs a guaranteed memory floor.

  • With automatic WLM, what happens to a query submitted at 'lowest' priority when the cluster is busy with 'high' priority work?
    It is admitted later and gets a smaller share of resources while running, but it is not cancelled. Redshift favours higher-priority queries for admission and for resource allocation, so the low-priority query stretches out rather than failing. If you need it stopped outright, that is a query monitoring rule with an abort action, not a priority setting.
  • How does a query end up in a specific WLM queue?
    Redshift walks the queue list in order and takes the first match: the submitting user's membership in a queue's `user_group`, then the session's `SET query_group TO 'label'` against a queue's `query_group`. Anything that matches nothing runs in the default queue, which is the last one in the configuration. Because it is first-match, queue order in the JSON is part of the design.
  • Why can changing WLM configuration require a cluster reboot?
    WLM lives in the cluster parameter group. Several WLM properties are dynamic and apply without a restart, but switching between manual and automatic WLM, and changing some cluster-level parameters such as the concurrency scaling cluster cap, are static and take effect only after the cluster reboots. Plan the switch in a maintenance window rather than mid-peak.

saying these in an interview costs you the question

  • Claiming WLM makes an individual query run faster
  • Thinking automatic WLM still honours your memory percentages
  • Believing query priority exists under manual WLM
  • Assuming queue order does not matter for routing
  • Describing WLM queues on Redshift Serverless, which has none

context

open as a page

How do you prove that slow Amazon Redshift dashboards are queueing rather than executing slowly?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Split each query's elapsed time into wait and run. In Redshift, STL_WLM_QUERY gives total_queue_time and total_exec_time per query, and SYS_QUERY_HISTORY exposes queue and execution time directly. High wait means a WLM problem; high run time means a query or table problem.

open as a page

In Amazon Redshift manual WLM, how much memory does a query get and what happens when it needs more?

level: middleimportance: should knowfreq 55%

basics

~20 s

A query gets one slot's share: the queue's memory percentage divided evenly by its concurrency level. If its hash tables or sorts exceed that, the step spills to disk instead of failing, and the query slows down sharply.

open as a page

What does concurrency scaling do on Amazon Redshift, and what must be true for it to kick in?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Concurrency scaling adds transient clusters when queries start queueing, routes eligible queued queries to them, and shuts them down after the burst. It needs concurrency scaling enabled on the WLM queue, actual queueing, and an eligible query — and it never speeds up a single running query.

open as a page

How would you keep daytime BI responsive on a Redshift cluster that also runs continuous ETL?

level: principalimportance: should knowfreq 38%

basics

~20 s

Route the two workloads to separate WLM queues by user group, run automatic WLM with BI at a higher priority than ETL, enable concurrency scaling on the BI queue and monitoring rules on the ad-hoc one. When contention persists, move BI to its own compute via data sharing.

open as a page

How do query monitoring rules work in Amazon Redshift WLM, and what would you use one for?

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

A query monitoring rule attaches to a WLM queue and combines up to three predicates on runtime metrics such as execution time or rows scanned. When all of them hold for a running query, Redshift performs the rule's action: log, hop, abort, or change its priority.

open as a page