skip to content

How does an Amazon Redshift RA3 node with managed storage differ from a DC2 node?

level: middleimportance: should knowfreq 62%

answer

  1. One number used to size two things
  2. Local disk stops being the source of truth
  3. Data grows without adding nodes
  4. Node-local SSD becomes a cache
  5. Sharing data between clusters needs it

basics

~20 s

RA3 nodes keep the authoritative copy of data in Redshift Managed Storage on S3 and use node-local SSD as a cache, so compute scales independently of data volume. DC2 stores data on node-local SSD only, so storage capacity dictates how many nodes you buy.

solid answer

~50 s

On the older DC2 node type, a cluster's capacity is whatever its nodes' local SSDs add up to. Once you outgrow that you must add nodes — even if the CPU is idle — and you pay for compute you do not need. RA3 breaks that coupling: the durable copy of every block lives in **Redshift Managed Storage**, an S3-backed layer billed separately per gigabyte, while each node's local SSD acts as a cache for the hot blocks. Storage grows on its own; you size the node count for the CPU and memory the workload needs. RA3 is also the modern feature baseline — cross-cluster data sharing and the AQUA acceleration layer require it, and the DS2 dense-storage type is legacy. DC2 still makes sense only for small datasets that fit comfortably in local SSD, where its price-performance on a tiny hot working set can win.

code

bash · 2 lines
bash
aws redshift describe-clusters \
  --query 'Clusters[].{Id:ClusterIdentifier,Type:NodeType,Nodes:NumberOfNodes}'

go deeper

for a junior

Know that RA3 is the current node family and that its data lives in Redshift Managed Storage rather than only on node disks, so storage size no longer forces the node count.

for a middle

Explain the mechanics: durable blocks in managed storage, local SSD as a hot cache, storage and compute billed and scaled separately, and the cold-cache effect after a restore or resize.

for a senior

Bring the operational consequences — sizing an RA3 cluster on CPU and memory rather than disk, planning a DC2/DS2 migration, and knowing which capabilities such as data sharing depend on the RA3 storage layer.

for a principal

Own the platform argument: what decoupled storage does to capacity planning, cost attribution between storage and compute, and the option to run several compute clusters over one shared copy of the data.

## The coupling RA3 was built to break On the legacy provisioned node types — DC2 (dense compute, local SSD) and the older DS2 (dense storage, magnetic disk) — a cluster's storage capacity is the sum of its nodes' local disks. That makes one number do two jobs. If your data grows, you add nodes to get disk, and you buy their CPU and memory whether or not the workload needs them. If your queries get heavier but the data does not grow, you add nodes and pay for storage you will never fill. Teams routinely ran DC2 clusters at 80–90% disk full and did emergency deletes, because running out of local disk is a hard wall. ## What RA3 changes RA3 node types are backed by **Redshift Managed Storage (RMS)**. The authoritative copy of every block lives in an S3-backed managed layer that Redshift operates for you; each node's large local SSD becomes a **cache** for the blocks that workload actually touches. Redshift tracks block access and keeps hot data local, fetching colder blocks from managed storage when a query needs them. Three things follow: 1. **Storage and compute are sized separately.** You pick node count for CPU and memory; managed storage grows on its own and is billed by the gigabyte-month, separately from node hours. 2. **The disk-full wall mostly disappears.** You are no longer forced to resize a cluster because a table grew. 3. **Scaling for performance is cheap in the data-movement sense.** Because the durable copy is not pinned to node-local disk, changing node count is a remapping and cache-warming exercise rather than a full copy of the dataset. What does *not* change: the leader/compute split, slices, distribution and sort keys, and the whole MPP execution model. RA3 changes where the durable bytes live, not how queries execute. Distribution style still matters exactly as much, because slices still own logical partitions of each table and still exchange rows over the network. ## Features that assume RA3 RA3 became the baseline for newer capabilities. Cross-cluster **data sharing** — letting another Redshift cluster or a serverless workgroup read your data live, without copying it — depends on the shared managed-storage layer, and so requires RA3 (or serverless) on both sides. The AQUA acceleration layer applies to eligible RA3 node types. DS2 is legacy and AWS's guidance has long been to migrate off it; DC2 remains available but is positioned for small workloads. ## When DC2 is still defensible A dataset that comfortably fits in local SSD, with a hot working set the cluster can hold entirely, and no need for data sharing, can be perfectly well served by DC2 — small dev clusters and modest marts are the honest use case. The moment data volume is the thing forcing node count, or you want to read the same data from more than one compute cluster, RA3 is the answer. ## Cold cache behaviour Because local SSD is a cache, a freshly resized or freshly restored RA3 cluster starts cold: the first queries pull blocks from managed storage and run slower than the same queries will an hour later. This is normal and worth calling out in an interview — people who have only read the marketing say "RA3 has unlimited storage with no performance difference" and miss that a cache miss costs a fetch. ## Migration path Moving an existing DC2 or DS2 cluster to RA3 is done either by restoring a snapshot into a new RA3 cluster or by resizing to the RA3 node type. The practical planning question is node count: you are no longer sizing for disk, so the old node count is usually the wrong starting point — you size for the CPU and memory the workload actually consumes and let managed storage take care of the rest. ## What an interviewer is testing Whether you understand that RA3 is a *storage-architecture* change with billing and scaling consequences, not a faster CPU. Strong answers name Redshift Managed Storage, describe local SSD as a cache, separate the two billing dimensions, and note the cold-cache caveat and the data-sharing dependency.

  • On RA3, is the node's local SSD still used at all?
    Yes, as a cache. Redshift keeps the hot blocks a workload touches on local SSD and fetches colder blocks from Redshift Managed Storage on demand. That is why a freshly restored or freshly resized RA3 cluster runs its first queries slower: the cache is cold and has to be warmed by real traffic.
  • Does RA3 make distribution and sort key choices less important?
    No. RA3 changes where the durable copy of data lives, not how queries execute. Slices still own logical partitions, joins still broadcast or redistribute over the network, and zone-map pruning still depends on sort order. A bad distribution key skews slices on RA3 exactly as it did on DC2.
  • Why does cross-cluster data sharing require RA3 rather than DC2?
    Sharing lets a second cluster read the producer's data live, without copying it, which is only possible when the data of record lives in a managed storage layer both compute clusters can reach. DC2 pins the durable copy to node-local SSD inside one cluster, so there is nothing for another cluster to attach to.

DC2 is a van whose cargo space and engine come as one package; RA3 is a van with a warehouse behind it — you upgrade the engine when you need speed and rent more warehouse when you need space.

saying these in an interview costs you the question

  • Calling RA3 simply a faster or newer CPU generation
  • Claiming RA3 makes distribution keys irrelevant
  • Assuming managed storage means queries read S3 every time
  • Expecting full performance immediately after a restore or resize
  • Thinking storage and compute are still billed as one unit on RA3

context