How does Atlas cluster tier and storage auto-scaling decide when to resize a cluster?
answer
- Two separate mechanisms, compute and disk
- Sustained pressure, not an instant spike
- It moves one tier at a time
- Up is eager, down is reluctant
- Disk grows near 90 percent, never shrinks
basics
~20 sAtlas watches sustained CPU and memory utilization and moves the cluster one tier at a time within the minimum and maximum you configure. Scale-down is optional and far more conservative. Storage auto-scaling grows the disk when usage approaches about 90 percent, and only ever grows it.
solid answer
~50 sCompute auto-scaling is available on dedicated tiers. You set a **minimum** and **maximum** instance size and Atlas moves between them based on sustained resource pressure: when average CPU or memory utilization stays high over a rolling window it scales **up one tier**, never several at once. Scaling **down** is opt-in and deliberately slower — Atlas requires a long stretch (a full day of low utilization) before shrinking, so a quiet night does not undo a real capacity increase. Storage auto-scaling is separate: when disk usage approaches roughly 90% Atlas increases the volume, and it only ever increases — reducing disk is a manual operation. Two caveats matter in practice. Each scaling event is a rolling change that ends in a primary election, so it is not free. And auto-scaling reacts over minutes, so it absorbs *trends*, not a sudden spike — and it will happily paper over a missing index while quietly raising the bill.
code
json · 11 lines{
"autoScaling": {
"compute": {
"enabled": true,
"scaleDownEnabled": true,
"minInstanceSize": "M30",
"maxInstanceSize": "M60"
},
"diskGBEnabled": true
}
}go deeper
Know that Atlas can resize a dedicated cluster automatically, that you configure a minimum and maximum tier, and that disk can grow on its own but never shrinks.
Explain the decision mechanics: sustained utilization over a rolling window, one tier per step, an opt-in and much slower scale-down, and a disk threshold around 90 percent.
Demonstrate that you plan for the disruption — a rolling resize ends in an election and a cold cache — and that you treat each scale-up as an investigation trigger rather than a resolution.
Own the ceiling as a budget and architecture control: where the maximum tier sits, when climbing tiers stops being the right answer versus sharding or workload isolation, and who reviews scaling events.
## Two independent mechanisms Atlas offers two things people lump together as "auto-scaling", and they behave differently. **Cluster tier (compute) auto-scaling** changes the instance size — M10 to M20, M20 to M30 — which changes RAM and vCPU together. It is available on dedicated tiers and is configured with a range: a minimum instance size and a maximum instance size. The maximum is your cost ceiling; the minimum is your floor so a quiet period cannot shrink the cluster below what you know you need. **Storage auto-scaling** changes the size of the disk, independently of the instance size. It is enabled by default on dedicated clusters. ## How compute auto-scaling decides Atlas samples resource utilization and looks for *sustained* pressure rather than momentary peaks. When average CPU utilization or memory utilization stays elevated over a rolling window, Atlas scales the cluster up by exactly one tier. It repeats if pressure persists, so a cluster can climb several tiers over time, but never jumps in a single step. That single-step behaviour is deliberate: each change is disruptive, and a one-tier move usually doubles the resources, which is enough to break most feedback loops. Scaling **down** must be explicitly enabled and uses a much longer and more conservative look-back — on the order of a full day of low utilization — before Atlas shrinks the tier. The asymmetry is the point. Capacity problems are urgent and cost problems are not, so Atlas is eager to add and reluctant to remove. Without that asymmetry, a nightly traffic trough would shrink your cluster right before the morning peak. ## How storage auto-scaling decides Storage is simpler: when disk usage climbs to roughly 90% of the provisioned volume, Atlas increases the volume size. This is a one-way ratchet. Atlas will never automatically shrink a disk, because shrinking a volume is a destructive, data-moving operation on every cloud provider. If you delete a large collection and want the money back, that is a manual resize you plan yourself. One subtlety worth knowing: on some cloud providers and volume types the disk's throughput and IOPS allowance are a function of its size. A storage auto-scale event can therefore change your I/O performance as a side effect, usually for the better, and it means disk size is not purely a capacity decision. ## What auto-scaling costs you A tier change is not a slider that takes effect instantly. Atlas applies it as a **rolling** change: secondaries are resized one at a time and resynchronised, then the primary steps down so a resized member can take over. That means an election and a short window where writes must be retried. It also means the new primary starts with a cold cache — the WiredTiger cache on the fresh instance holds nothing — so latency often rises briefly *after* a scale-up before it improves. The practical consequences: - Auto-scaling handles **trends** (a growing dataset, a seasonal ramp, a steadily busier product), not **spikes**. A traffic surge that arrives in thirty seconds will be over, or will have hurt you, long before a new tier is serving traffic. - Your application must tolerate a primary election at an arbitrary time. Retryable writes and retryable reads, an SRV connection string and a sane server-selection timeout are prerequisites, not nice-to-haves. - The maximum instance size is a **budget control**. Leaving it wide open converts a runaway query into a runaway invoice. ## Auto-scaling is not a substitute for engineering The most common failure mode is using auto-scaling to hide a fixable problem. A collection scan that should be an index hit will drive CPU up; Atlas will dutifully buy a bigger machine; the query is still a scan and now costs more per hour. The same is true of a working set that no longer fits in RAM because documents were designed to be huge, or of a connection storm from an over-provisioned application tier. A healthy pattern is: set a min and max that bracket what you have measured, enable scale-down so you actually reclaim the money, then treat every scale-up event as an **alert worth reading** rather than a resolved incident. If the cluster climbs a tier, someone should look at the Performance Advisor and at what changed in the application that week. ## Configuring it deliberately Pick the minimum from your known steady-state working set, not from the smallest tier that runs. Pick the maximum from what you are willing to pay and from what your architecture can actually use — beyond a point, a single replica set stops being the right answer and the conversation turns to sharding. And schedule the risky changes: if you know a large migration or a bulk load is coming, size up ahead of it during a quiet period rather than letting an automatic election land in the middle of the job.
- Why is Atlas's scale-down window so much longer than its scale-up window?Because the costs are asymmetric. Being under-provisioned degrades or breaks the service now; being over-provisioned only wastes money. A short scale-down window would also cause oscillation: every nightly trough would shrink the cluster just before the morning ramp, and each of those changes triggers a rolling resize and an election.
- Your cluster auto-scaled from M30 to M40 overnight. What do you check first?What changed, not whether the new tier is enough. Look at the Performance Advisor and slow-query log for a new or newly unindexed query, check whether data volume pushed the working set out of RAM, and check for a connection or traffic change from the application tier. A scale-up is a symptom report, not a fix.
- Can auto-scaling protect you from a sudden traffic spike?No. It reacts to sustained utilization over minutes and each change requires a rolling resize ending in an election, so the new capacity arrives long after a sharp spike has passed. Spikes are handled by headroom, caching, connection limits and back-pressure in the application — auto-scaling handles growth trends.
- Why can storage auto-scaling change query performance and not just capacity?On several cloud providers and volume types, a disk's IOPS and throughput allowance scale with its provisioned size. Growing the volume therefore raises the I/O ceiling as a side effect. It also means shrinking a disk later to save money can quietly reduce throughput, which is one more reason Atlas never shrinks it automatically.
saying these in an interview costs you the question
- Thinking auto-scaling absorbs sudden traffic spikes
- Believing scaling is instant and non-disruptive
- Assuming Atlas shrinks disks automatically to save money
- Leaving the maximum instance size unbounded
- Treating a scale-up event as the fix rather than a symptom