When would you use elastic resize versus classic resize on an Amazon Redshift cluster?
answer
- Node count change means slice count change
- One path keeps the cluster, one builds a new one
- Minutes versus hours
- The fast path has allowed-target limits
- Performance ramps while data rebalances
basics
~20 sElastic resize adds or removes nodes in minutes by remapping data across the new slice count, with a short pause in query processing. Classic resize provisions a new cluster and copies the data, taking far longer with the source read-only. Prefer elastic; use classic when elastic cannot make the change you need.
solid answer
~50 s**Elastic resize** is the default tool: it changes node count (and on some paths node type) in minutes. Sessions are held and queries pause briefly while slices are remapped, then the cluster comes back on the same endpoint. Data is redistributed across the new slice arrangement in the background afterwards, so performance ramps rather than jumping instantly. It is constrained — only certain target node counts and type changes are permitted from a given starting shape. **Classic resize** provisions a whole new cluster and moves the data into it. Historically that meant hours and a read-only source cluster, which is why it is the fallback rather than the default. AWS has since made classic resize substantially faster on RA3, where the data already lives in managed storage, so confirm current behaviour for your node type before planning a maintenance window. A third option is snapshot-and-restore into a differently shaped cluster, then cut the endpoint over — useful when you want the old cluster intact as a rollback.
code
bash · 5 linesaws redshift resize-cluster \
--cluster-identifier prod-dw \
--node-type ra3.4xlarge \
--number-of-nodes 8 \
--no-classicgo deeper
Know that a Redshift cluster can be resized and that there are two mechanisms, one fast and one slow, and that resizing is not instantaneous or invisible to running queries.
Explain why changing node count forces data to be remapped across a new slice count, and describe the availability profile of each path: a brief pause versus a long read-only period.
Plan a real resize: pick the path, size the maintenance window, warn stakeholders about the post-resize warm-up, and distinguish capacity pressure from concurrency pressure before touching node count at all.
Own the policy — scheduled resizes around ETL windows versus steady-state sizing, when a migration deserves snapshot-restore with rollback, and when the answer is separate compute rather than one cluster that keeps growing.
## Why resizing is not just "add a node" An Amazon Redshift cluster's parallelism is nodes multiplied by slices per node, and every table's rows are placed on specific slices. Changing the node count therefore changes the slice count, which means the data placement must change too. That is what makes resizing an operation with a maintenance profile rather than a configuration toggle. ## Elastic resize Elastic resize is the fast path. Redshift keeps the existing cluster and its endpoint, briefly pauses query processing while it remaps data partitions onto the new set of slices, and resumes. Client sessions are held rather than dropped, so well-behaved applications experience a stall rather than an error, though long-running queries in flight can be terminated. Two properties matter operationally: - **It is constrained.** From a given node type and count, only certain targets are allowed; the permitted range depends on the starting shape. If the change you want falls outside it, the console will not offer elastic resize. - **Rebalancing continues afterwards.** Immediately after the resize, some slices carry more than their eventual share, and Redshift redistributes in the background. Queries run throughout, but the cluster does not deliver its final performance the second the resize reports complete. Benchmarking in the first minutes after a resize produces misleading numbers. Elastic resize can also be **scheduled**, which is how teams run a larger cluster during the nightly ETL window and a smaller one during the day. ## Classic resize Classic resize takes the other approach: Redshift provisions a new cluster of the target shape and moves the data across. Historically this is the slow, heavyweight operation — long-running, with the source cluster available for reads but not writes for the duration, and duration scaling with data volume. You reach for it when the change you need is outside elastic resize's allowed transitions, for example a node-type migration elastic cannot perform or a target size far from the current one. On RA3, where the durable data already lives in Redshift Managed Storage rather than on node-local disk, AWS has improved classic resize so far less data movement is required and the operation completes much faster, with redistribution finishing in the background. Because this behaviour has changed over time, the professional answer is to check the current documented behaviour for your node type rather than quoting a fixed duration. ## Snapshot and restore The third path is to take a snapshot and restore it into a new cluster with the shape you want, then repoint applications. Its advantage is that the original cluster survives untouched, so rollback is trivial and you can validate the new shape before cutting over. Its cost is a cutover window, the need to catch up any writes that happened after the snapshot, and paying for two clusters briefly. This is the usual choice for a node-family migration such as DC2 to RA3, where you want a safety net. ## Choosing in an interview scenario Work through it in this order: 1. Is the change within elastic resize's allowed targets? If yes, use it — minutes, same endpoint, no data copy. 2. Is it a node-type or magnitude change elastic cannot do? Then classic resize, or snapshot-restore if you want the old cluster retained as rollback. 3. Is the pressure actually *concurrency* rather than capacity? Then resizing may be the wrong tool entirely; bursts of queued queries are addressed with concurrency features, not by permanently growing the cluster. 4. Is the pressure *storage*? On RA3 that is no longer a reason to resize at all — managed storage grows on its own. That fourth point is worth making explicitly, because a large share of historical resize traffic existed only because DC2 and DS2 coupled disk to node count. ## Aftermath and verification After any resize, expect a warm-up: on RA3 the local SSD cache must refill from managed storage, and background redistribution may still be running. Validate with your own representative queries after the cluster has taken real traffic for a while, not in the first minutes. Also re-examine assumptions that depended on the old slice count — most visibly, the number of files you split bulk loads into, which should track the new total slice count.
- Why does query performance sometimes keep improving for a while after an elastic resize completes?Elastic resize first remaps data partitions onto the new slice set so the cluster can serve queries quickly, then redistributes data evenly in the background. Until that finishes, some slices carry more than their share and pace each step. On RA3 the local SSD cache is also cold and refilling from managed storage. Both effects resolve with time and traffic.
- Your cluster is fine on CPU but queries queue during a morning burst. Is a resize the right response?Usually not. Queueing is a concurrency and admission problem, not a capacity one, and permanently enlarging the cluster pays for peak-shaped hardware all day. Address the burst with Redshift's concurrency features and workload configuration; resize when sustained CPU or memory pressure, not short queues, is the evidence.
- When would you prefer snapshot-and-restore over either resize path?When you want the original cluster left intact as a rollback, or when you are changing node family and want to validate the new shape before committing. You restore a snapshot into a cluster of the target shape, test it, then repoint applications. The cost is a cutover window, catching up post-snapshot writes, and briefly paying for two clusters.
saying these in an interview costs you the question
- Assuming any resize is instant and transparent to clients
- Believing classic resize keeps the cluster fully writable throughout
- Benchmarking immediately after a resize and drawing conclusions
- Resizing to fix query queueing rather than concurrency settings
- Adding nodes on RA3 purely because storage is growing