A team wants to move an internal Aurora PostgreSQL cluster with spiky, mostly-idle traffic onto Aurora Serverless v2. What is an ACU, how does Serverless v2 change capacity at runtime, and when would you tell them to stay on provisioned instances?
answer
- capacity range, not an instance class
- one unit is about 2 GiB of memory
- scales in place, no dropped connections
- the floor sets your warm cache
- flat busy load loses the bet
basics
~20 sAn Aurora Capacity Unit is roughly 2 GiB of memory plus matching CPU and network. Serverless v2 scales an instance in place between a minimum and maximum ACU setting without dropping connections. It suits spiky or idle workloads; steadily busy clusters cost more than provisioned.
solid answer
~50 sAurora Serverless v2 replaces the fixed instance class with a capacity range. Capacity is measured in Aurora Capacity Units, where one ACU is about 2 GiB of memory with proportional CPU and networking, and you set a minimum and a maximum in fine-grained steps. Aurora then adjusts the running instance in place, in seconds, without dropping connections or restarting the engine — the buffer cache and sessions survive. That is the whole appeal for spiky or overnight-idle workloads: you stop paying for a peak-sized instance twenty-four hours a day. Tell them to stay provisioned when load is steady near the top of the range, because an ACU-hour costs more than the equivalent provisioned capacity and Reserved Instance commitments cover provisioned instances rather than serverless capacity. Also watch the minimum: memory scales with ACU, so a floor set too low leaves a cold, tiny buffer cache to ramp from.
code
bash · 3 linesaws rds modify-db-cluster \
--db-cluster-identifier reporting \
--serverless-v2-scaling-configuration MinCapacity=2,MaxCapacity=32go deeper
Know that Serverless v2 replaces choosing an instance size with setting a capacity range in ACUs, and that roughly one ACU is 2 GiB of memory. Say that it scales automatically with load.
Explain the in-place scaling mechanic — no restart, no dropped connections, seconds to react — and contrast it with the disruptive scaling-point behaviour of the original Serverless v1.
Show the cost and cache reasoning: an ACU-hour carries a premium, so flat high utilization belongs on provisioned instances, and a minimum set too low leaves a cold buffer cache and billed I/O on every morning ramp.
Own it as a portfolio decision. Decide which clusters get committed provisioned capacity and which get elasticity, define how the maximum acts as a spend guardrail, and set the evidence — utilization telemetry over weeks — that moves a cluster between the two.
## What an ACU is Aurora Serverless v2 does not give you an instance class. It gives you a capacity range expressed in **Aurora Capacity Units**. One ACU is approximately 2 GiB of memory together with a corresponding amount of CPU and network throughput. You configure a minimum and a maximum, in fine-grained increments, and Aurora keeps the instance somewhere between them. ```bash aws rds modify-db-cluster \ --db-cluster-identifier reporting \ --serverless-v2-scaling-configuration MinCapacity=2,MaxCapacity=32 ``` Billing is per ACU-second consumed, plus storage and I/O as with any Aurora cluster. ## How the scaling actually behaves This is where v2 differs from the original Serverless v1, and interviewers ask because the two behave nothing alike. Serverless v1 scaled by finding a "scaling point" — a moment with no long-running transactions — and could force connections to drop to get there, which made it unsuitable for anything latency-sensitive. Serverless v2 scales the **running** instance in place: memory and CPU are added or removed underneath a live engine, in seconds, without a restart, without dropping connections, and without discarding the buffer cache. Scale-up is aggressive when demand rises; scale-down is deliberately gradual, because shrinking memory means giving up cache. A Serverless v2 instance is just another member of an Aurora cluster, so you can mix it with provisioned instances — a common pattern is a provisioned writer with serverless readers, or the reverse for a dev cluster. Readers in the highest promotion tiers track the writer's capacity so that a failover does not land on an undersized instance; readers in lower tiers scale on their own demand. On recent engine versions the minimum can be set to zero, which lets an idle cluster pause entirely and pay nothing for compute, waking on the next connection after a short delay. That is genuinely useful for development and demo clusters and genuinely wrong for anything user-facing, where the wake latency lands on a real request. ## Choosing the range The two numbers do more work than people expect. - **The maximum is your blast radius and your budget cap.** It bounds spend, and it is also what connection-related limits are derived from, so setting it absurdly high to "be safe" is not free of consequences. - **The minimum is your warm floor.** Memory scales with ACU, so a low minimum means a small buffer cache. A cluster that sits at the floor overnight and is then hit at 09:00 has to both scale up and re-warm its cache from storage — and every one of those page reads is billed I/O. If your working set is 20 GiB, a minimum of 0.5 ACU guarantees a bad morning. Set the floor to hold the working set you want resident. ## When to stay provisioned Serverless is a utilization bet, and you lose the bet when utilization is high and flat: 1. **Steady load near the top of the range.** An ACU-hour is priced above the equivalent capacity in a provisioned instance. If the cluster runs at 80% of its ceiling all day, provisioned is cheaper before you even negotiate. 2. **Commitment discounts.** Reserved Instance style commitments apply to provisioned Aurora instances. A workload predictable enough to commit to is a workload that should be provisioned. 3. **Predictable large working sets.** If you already know you need 64 GiB resident, you are paying serverless rates for capacity you never let go of. 4. **Latency floors.** Scaling is fast but not instantaneous. A workload that goes from idle to full traffic in one second — a scheduled cron fan-out, a cache-flush stampede — can spend the first part of the spike undersized. Provisioned capacity, or a higher minimum, absorbs it. The honest framing in an interview: serverless converts a capacity-planning problem into a price-per-unit problem. You pay a premium per unit of capacity in exchange for not owning the sizing decision, and that trade is excellent for spiky, idle, or unpredictable clusters and poor for the steady ones. ## How to verify rather than guess Aurora publishes `ServerlessDatabaseCapacity` (the current ACU value) and `ACUUtilization` to CloudWatch. Before migrating, look at the existing provisioned cluster's CPU and memory profile over a couple of weeks; after migrating, alarm on sustained time at the maximum, which means the ceiling is now the constraint rather than a safety net.
- How does Serverless v2 differ from the original Aurora Serverless v1?v1 scaled by pausing at a scaling point and could force connections to drop, jumping between coarse capacity steps, which ruled it out for latency-sensitive work. v2 resizes the running instance in place in fine increments, in seconds, keeping connections and the buffer cache. They are different products, not versions of one behaviour.
- A Serverless v2 cluster with a 0.5 ACU minimum is slow every morning for several minutes. What is happening?Overnight it sat at the floor with a buffer cache sized for 1 GiB of memory. The morning burst forces both a scale-up and a cold cache, so early queries read pages from shared storage — slow and billed as I/O. Raise the minimum so the working set stays resident.
- Can you mix Serverless v2 and provisioned instances in the same Aurora cluster?Yes. A Serverless v2 instance is an ordinary cluster member, so a provisioned writer with serverless readers, or a serverless writer with a provisioned reader for reporting, are both valid. Watch promotion tiers so a failover does not promote an instance sized for a fraction of the writer's load.
saying these in an interview costs you the question
- Says Serverless v2 drops connections to scale, like v1
- Thinks an ACU is a CPU core
- Sets the minimum near zero for a production user-facing cluster
- Assumes serverless is always cheaper than provisioned
- Believes serverless removes the need for a maximum capacity decision