skip to content

What is AQUA in Amazon Redshift, and which kinds of queries does it help?

level: middleimportance: nice to knowfreq 20%

answer

  1. Filter the data before it crosses the network
  2. It lives near managed storage, not on your nodes
  3. Only the RA3 family qualifies
  4. Big scans with selective filters benefit
  5. Not a knob you tune today

basics

~20 s

AQUA, the Advanced Query Accelerator, is a hardware-accelerated caching and compute layer that sits between eligible RA3 compute nodes and managed storage. It pushes scan-time filtering and simple aggregation closer to the data, so fewer rows cross the network.

solid answer

~50 s

AQUA (Advanced Query Accelerator) is an acceleration layer available on eligible RA3 node types. Rather than every compute node pulling blocks over the network and filtering them locally, AQUA nodes sitting next to Redshift Managed Storage apply predicates and some aggregation first, so only the surviving data travels to the cluster. The workloads it helps are the ones dominated by scanning: large table scans with selective filters, including string pattern matching, and simple aggregations over a lot of data. It does nothing meaningful for small queries, join-heavy plans, or anything already served from cache — the bottleneck there is not the scan-to-network path. AQUA's enablement model has changed over Redshift's life; today it is applied automatically on eligible node types rather than being a knob you tune, so treat it as background acceleration, not something you design a schema around.

go deeper

for a junior

You are not expected to know AQUA. If it comes up, it is enough to say it is an Amazon-managed acceleration layer for RA3 clusters that you do not configure yourself.

for a middle

Be able to place it in the architecture: between managed storage and compute nodes, filtering and pre-aggregating so fewer rows cross the network, and name the scan-heavy query shapes it helps.

for a senior

Show judgment about its limits — it does nothing for join-bound plans or small queries, and it never substitutes for pruning through sort order. Be honest that its enablement model has changed over time.

for a principal

Treat it as background acceleration with no architectural leverage: it should never appear in a capacity plan or a platform decision, and citing it as a reason to skip data modelling work is a mistake to challenge.

## The problem AQUA addresses In an RA3 cluster the durable data lives in Redshift Managed Storage and compute nodes cache hot blocks on local SSD. When a query scans far more data than it ultimately returns — a huge fact table filtered down to a small slice, or a pattern match over a wide text column — the naive path moves a lot of bytes from storage to compute only for most of them to be discarded microseconds later. Network bandwidth between storage and compute becomes the limiting resource, not the compute nodes' CPUs. **AQUA (Advanced Query Accelerator)** attacks that path. It is a layer of accelerated nodes positioned close to managed storage that can evaluate certain scan-time work — predicate filtering and some aggregation — before data reaches the cluster's compute nodes. The compute nodes then receive a much smaller stream and finish the query normally. The general idea is the familiar one of pushing computation down to the data rather than pulling data up to the computation; AQUA is Redshift's hardware-assisted implementation of it. ## What it accelerates, and what it does not AQUA helps queries whose cost is dominated by scanning a lot of data and discarding most of it: - large scans with **selective predicates**, where a small fraction of rows survive; - **string matching** predicates over large text columns, which are expensive to evaluate row by row on the compute nodes; - **simple aggregations** over very large inputs. It does not help: - small queries, where the scan is not the bottleneck at all; - **join-heavy** plans, where the cost lives in redistributing and hashing rows between compute nodes rather than in the scan; - queries whose data is already resident in the local SSD cache and whose result is small; - anything served from the result cache, which never executes. ## Where it fits in the architecture It is useful to place AQUA alongside the other layers of an RA3 cluster: the leader node plans and merges; compute nodes hold slices and execute plan segments; managed storage holds the durable blocks; local SSD caches the hot ones; and AQUA sits on the path between managed storage and compute, doing early filtering. Nothing about slices, distribution style, sort order or zone-map pruning changes because AQUA exists — pruning still avoids reading blocks entirely, which is strictly better than filtering them quickly. That ordering matters for an interview answer. The best way to move less data is not to read it: good sort-key design and predicates that let Redshift skip blocks outright dominate any scan-time acceleration. AQUA reduces the cost of the data you *do* read. ## Availability and configuration AQUA applies to eligible RA3 node types — the larger ones — and its availability has varied by region. Its enablement story has also changed over time: it was once an explicit cluster setting, and AWS later moved to applying it automatically where it is beneficial, removing the knob. Because of that history, the safe interview answer states the mechanism confidently and treats the configuration surface as version-dependent: check the current documentation for whether there is anything to turn on for your node type and region. ## Why it is a differentiator question, not a core one Nobody should fail a Redshift interview for not knowing AQUA. It is a background accelerator with no schema or query implications, and it is not something you tune. It appears in interviews as a probe: does the candidate know the RA3 storage path well enough to say where such a layer would even sit, and are they honest about the fact that it does not substitute for distribution, sort-key or query design? A candidate who claims AQUA makes tuning unnecessary, or who attributes it to DC2 clusters or to serverless capacity, has revealed a shakier grasp of the architecture than one who simply says "it is scan-side acceleration on RA3, and I have never had to configure it."

  • Why does AQUA do little for a join-heavy query?
    AQUA reduces the data crossing the path from managed storage to compute by filtering early. A join-heavy plan's cost is elsewhere: redistributing or broadcasting rows between compute nodes and building hash tables. Cutting scan-side bytes barely touches that, so distribution key choice and join strategy remain the levers.
  • If AQUA filters at scan time, does sort-key design still matter?
    Yes, and more. Pruning driven by sort order and block-level min/max metadata avoids reading blocks at all, which beats reading them and filtering quickly. AQUA reduces the cost of data you do read; good sort keys reduce how much you read in the first place. They are complementary, and the pruning win is the larger one.

It is like having a sorter at the warehouse door who discards the wrong parcels before the truck is loaded, instead of driving everything to the depot and sorting there.

saying these in an interview costs you the question

  • Saying AQUA removes the need for distribution and sort key tuning
  • Claiming AQUA works on DC2 clusters
  • Describing AQUA as a result cache for repeated queries
  • Expecting AQUA to speed up join-heavy plans
  • Presenting AQUA as a setting you tune per query

context