What happens to a running application while an Atlas cluster tier change is applied?
answer
- It is not an in-place resize
- Members are replaced one at a time
- The primary has to give up its role
- Drivers retry, which is why it looks seamless
- Fresh member means an empty cache
basics
~20 sAtlas applies a tier change as a rolling replacement: secondaries are resized and resynchronised one at a time, then the primary steps down so a resized member takes over. Expect a brief election with no primary, plus higher latency afterwards while caches refill.
solid answer
~50 sA tier change is not an in-place slider. Atlas resizes members **one at a time**, keeping a majority available, and finishes by stepping the primary down so the last member can be replaced. That step-down triggers an election lasting seconds, during which there is no primary: writes fail unless the driver retries them, and clients must rediscover the new topology. Applications survive this cleanly if they use an SRV connection string, keep **retryable writes and retryable reads** enabled, set a server-selection timeout that spans an election, and never cache resolved hosts. After the change, expect a period of *worse* latency: the new primary's WiredTiger cache and the host page cache are cold, so reads hit disk until the working set is paged back in. Consequently: schedule deliberate resizes for low-traffic windows, and never scale down under load — a shrunken cache plus live traffic is how a resize becomes an incident.
code
javascript · 6 lines// SRV string plus retries: an election during a resize is retried, not surfaced
const uri = "mongodb+srv://user:[email protected]/?retryWrites=true&w=majority";
const client = new MongoClient(uri, {
serverSelectionTimeoutMS: 30000,
maxPoolSize: 50
});go deeper
Know that resizing an Atlas cluster causes a brief interruption because the primary changes, and that the driver's automatic retries are what usually hide it.
Explain the rolling sequence — secondaries first, then a step-down and election — and name the client settings that make it survivable, including retryable writes and the SRV connection string.
Demonstrate operational judgment: plan the window, expect a cold-cache latency tail after the change, guard against reconnect storms, and never shrink a tier while the workload is at peak.
Own the policy around unattended resizes: whether auto-scaling may fire during business hours, what the retry and timeout standards are across services, and how tier changes are communicated and reviewed.
## The change is a rolling replacement When you move an Atlas cluster from one tier to another — by hand or via auto-scaling — Atlas does not grow a running machine. It performs a **rolling** change across the replica set: one member at a time is replaced or resized and brought back into sync, while the remaining members continue to serve the workload and preserve a majority. Only when the secondaries are done does Atlas deal with the primary, and the only way to replace a primary is to stop it being the primary. So Atlas issues a **step-down**, an election runs, a resized secondary is elected, and the old primary is replaced last. That sequence has one unavoidable consequence: there is a short interval with **no primary**. Elections complete in seconds on a healthy cluster, but during that interval every write and every primary-targeted read has nowhere to go. ## What the application sees From the driver's point of view, the primary disappears and a new one appears somewhere else in the set. Concretely: - In-flight writes to the old primary fail with a "not primary" style error. - New operations block in server selection until a primary is discovered again. - Connections to the replaced member are closed, so the pool must re-establish them. Modern drivers handle all of this **if you let them**. Retryable writes cause the driver to transparently retry a single-statement write once against the new primary; retryable reads do the same for reads. Both are on by default in current drivers, and the Atlas-provided connection string includes `retryWrites=true`. Server selection has its own timeout — leave enough of it that an election fits inside the window instead of surfacing as an error to your users. An `mongodb+srv://` connection string matters too: it resolves the seed list through DNS and lets the driver discover topology changes, rather than pinning your application to hostnames that a resize may invalidate. ## The cold-cache tail The part people forget is what happens *after* the election succeeds. MongoDB's performance is dominated by whether the working set sits in the WiredTiger cache and the operating system page cache. A newly provisioned or restarted member starts with both empty. So immediately after a tier change, the new primary serves reads from disk, p99 latency climbs, and the cluster looks *worse* than the tier you left — sometimes for minutes on a large working set. This inverts the naive expectation ("I scaled up, so it should be fast now") and it is the reason a scale-up during peak traffic can make an incident worse before it makes it better. It is also why scaling **down** under load is genuinely dangerous: you simultaneously shrink the cache and empty it, while traffic is unchanged. ## Making a resize safe The checklist a senior engineer should produce: 1. **Retries on.** Retryable writes and retryable reads enabled; application-level operations idempotent where a retry could duplicate work. 2. **SRV connection string**, no hard-coded hosts, no cached DNS beyond its TTL. 3. **Timeouts that span an election.** Server-selection timeout comfortably larger than a normal election; sensible client-side operation timeouts so a stalled request fails predictably instead of piling up. 4. **Back-pressure.** A brief period with no primary means requests queue. Bounded pools and fast failure beat unbounded queueing, which turns a five-second election into a minute-long recovery as a thundering herd reconnects. 5. **Timing.** Deliberate resizes go in a low-traffic window. Auto-scaling can fire at any time, which is a reason to test the behaviour rather than hope. 6. **Watch the tail.** Measure latency for a while after the change; do not declare success at the moment the console says the cluster is healthy. ## Multi-region and larger topologies On a multi-region cluster the election has an extra dimension: the member that wins may sit in a different region from the one that just stepped down, which changes the latency your writers see and the round trip that a majority acknowledgement requires. If a specific region should hold the primary, member priorities need to reflect that so the cluster settles back where you intended rather than staying wherever the election landed. ## The interview framing The naive answer is "nothing, Atlas handles it, it's zero downtime". The accurate answer is "it is *near*-zero downtime, and only because the driver retries for you". Saying that out loud — plus the cold-cache tail — is what separates someone who has resized a production cluster from someone who has read the marketing page.
- Which driver settings make a tier change survivable without user-visible errors?An mongodb+srv connection string so topology is rediscovered, retryable writes and retryable reads left enabled, a server-selection timeout comfortably longer than a normal election, and bounded connection pools with client-side operation timeouts so queued work fails fast instead of stampeding when the new primary appears.
- Why is scaling a cluster down under heavy load riskier than scaling up?You shrink the cache and empty it at the same time while traffic is unchanged. The new smaller primary must serve the full working set from disk on less RAM, so latency spikes and may never recover to the old level. Scale down during a trough, and only after confirming the smaller tier can hold the working set.
- Your team says a resize is zero-downtime. How would you correct that precisely?It is near-zero downtime, not zero. Replacing the primary requires a step-down and an election, so there is a short interval with no primary. The gap is invisible only because drivers retry the affected operations. Any non-retryable operation, or a client with a too-short server-selection timeout, will still see errors.
saying these in an interview costs you the question
- Calling a tier change fully zero-downtime
- Not knowing the primary must step down
- Expecting instant improvement right after scaling up
- Scaling down during peak traffic
- Hard-coding replica set hostnames instead of using SRV