What does a replica set's electionTimeoutMillis control, and what breaks if you set it too low?
answer
- A debounce, not an instant trigger
- Measured in milliseconds, defaults to ten seconds
- Heartbeats fire far more often than this
- Too short and a GC pause looks fatal
- Governs the detection phase of failover
basics
~10 selectionTimeoutMillis is how long an eligible secondary waits without reaching the primary before calling an election. It defaults to 10000 ms. Lowering it shortens failover detection but makes ordinary latency spikes trigger needless elections.
solid answer
~40 s`electionTimeoutMillis` is a replica set setting, default `10000`, that bounds how long a secondary goes without successful contact with the primary before it declares the primary gone and stands for election. Members exchange heartbeats every `heartbeatIntervalMillis`, default `2000`, so several missed heartbeats accumulate before the timeout fires. Lowering the value shortens the detection half of a failover, which is usually the dominant part of the write-unavailable window. The cost is false positives: a garbage-collection pause, a saturated network link, or a cross-region latency spike now looks identical to a dead primary, and the set steps down a perfectly good primary. Each spurious election costs another write-unavailable window and risks rolling back writes that were only acknowledged by the old primary. Tune it down only where the network between voters is genuinely stable and low-latency.
code
json · 7 lines{
"settings": {
"heartbeatIntervalMillis": 2000,
"heartbeatTimeoutSecs": 10,
"electionTimeoutMillis": 10000
}
}go deeper
Know that the setting exists, that it is measured in milliseconds with a default of 10000, and that it governs how long a secondary waits before deciding the primary is gone.
Explain the mechanics: heartbeats every two seconds by default, the timeout as a debounce over several of them, and the direct trade between detection speed and false positives.
Demonstrate judgment about tuning it against a measured latency distribution, and be able to name the other phases of a failover so you do not blame the whole outage window on this one setting.
Own the availability budget end to end — what write-unavailability the business actually requires, which phases you can compress, and when client-side retryable writes are the cheaper answer than an aggressive timeout.
## The setting `electionTimeoutMillis` lives in the `settings` sub-document of the replica set configuration, alongside `heartbeatIntervalMillis` and `heartbeatTimeoutSecs`. Its default is `10000` — ten seconds. It answers one question: how long may a secondary go without confirming that a primary is alive and reachable before it concludes that there is no primary and calls for an election? It is a per-set setting, changed with `rs.reconfig()`, and it applies to every member. ## How detection actually works Members of a replica set send heartbeats to each other on an interval governed by `heartbeatIntervalMillis`, whose default is `2000` — every two seconds. A heartbeat that gets no response within `heartbeatTimeoutSecs` marks that member inaccessible in the sender's view of the set. A secondary that stops hearing from the primary does not immediately start an election. It waits out `electionTimeoutMillis`. Only when that window elapses with no successful contact does an eligible secondary — one with `priority` greater than 0 and a sufficiently fresh oplog — become a candidate and solicit votes. MongoDB's protocol also runs a preliminary check with the other members before formally standing, so that a node isolated from everyone does not repeatedly disrupt the set's term with elections it cannot win. So the timeout is a debounce. It is deliberately much longer than the heartbeat interval, giving several heartbeats a chance to succeed before anything drastic happens. ## Where the timeout sits in the failover budget A failover has several sequential phases, and this setting owns the first and usually the largest: 1. **Detection** — up to `electionTimeoutMillis` of silence. 2. **Voting** — a candidate solicits votes from the other voters; this is roughly one round trip to a majority, so it is bounded by the latency to the slowest member of that majority. 3. **Catch-up** — the winner tries to apply any oplog entries that other members have but it does not, before accepting writes. `catchUpTimeoutMillis` bounds this. 4. **Client rediscovery** — drivers monitor the topology and must notice the new primary; until they do, application writes wait on server selection. If your measured failover takes twelve seconds, roughly ten of them are usually phase one. That is why the setting is the first knob people reach for. ## What lowering it costs The timeout is not just a delay to be minimised. It is the margin that separates "the primary is dead" from "the primary is briefly slow". Shrink it and those two conditions become indistinguishable to the rest of the set. Concretely: - A long stop-the-world pause, a swapping host, or a disk stall on the primary makes it miss heartbeat replies. With a short timeout, secondaries call an election, a new primary appears, and the old primary steps down when it learns of the higher term — even though nothing was actually wrong for more than a second. - A transient network event between sites — a link flap, a congested uplink, a route change — produces the same outcome. - Members deployed across regions have baseline round-trip latencies that can approach an aggressive timeout on their own. Every spurious election has real costs. The set is write-unavailable for the duration of the election and catch-up. Writes that had been acknowledged only by the old primary and never replicated to a majority may be rolled back. Connections are dropped and clients see errors or retries. And an unstable primary role makes every other operational signal harder to read. ## What raising it costs The opposite direction is equally real: a longer timeout means a genuinely dead primary is tolerated for longer, and the application's writes fail or block for that whole window. If your service has a hard write-availability target, a very long timeout may simply violate it. ## How to choose The honest procedure is empirical, not theoretical. Measure the actual round-trip latency and its tail between the voting members. Look at how often the primary experiences pauses long enough to miss several heartbeats. Then pick a timeout with comfortable headroom above the worst normal case, and verify by rehearsing failovers with `rs.stepDown()` and measuring the client-visible outage. In practice the default is well chosen for a set whose voters share a low-latency network. If your write-unavailability target is tighter than the default allows, the more productive levers are often elsewhere: enable retryable writes so a single failover looks like a latency blip to the application rather than an error, keep the voters close to each other so voting and catch-up are fast, and use `priority` to steer where the primary lands so it is near the clients. ## Related settings worth naming `heartbeatIntervalMillis` (default `2000`) sets the heartbeat cadence; `heartbeatTimeoutSecs` bounds how long a member waits for a heartbeat response; `catchUpTimeoutMillis` bounds the new primary's catch-up phase. Knowing which of these owns which phase is what distinguishes a tuned deployment from a superstitious one.
- Which other settings determine how quickly a MongoDB replica set failover completes once the election starts?`heartbeatIntervalMillis` sets how often members probe each other. The vote itself costs roughly one round trip to a majority of voters, so it is bounded by the latency to the slowest voter in that majority. `catchUpTimeoutMillis` bounds how long the new primary spends applying oplog entries it was missing before it accepts writes. On the client side, driver server-selection and topology-monitoring settings determine how fast the application notices the new primary.
- Your application sees write errors for a few seconds during every failover even though the election is fast. Where else would you look?The client side. Drivers rediscover the topology on their own monitoring interval, so there is a gap between the new primary being elected and the driver routing to it. Enabling retryable writes lets the driver retry a failed write after re-selecting a server, which converts most single failovers into added latency rather than an error. Server-selection timeout controls how long the driver waits before giving up.
- Is it ever right to raise electionTimeoutMillis above the default?Yes, when the voters sit on a network whose normal tail latency or occasional flapping approaches the default, and when a few extra seconds of write unavailability costs less than repeated spurious failovers. Sets stretched across distant regions are the usual case. Raise it only after measuring the latency distribution between voters, and accept the longer detection window for a genuinely dead primary.
saying these in an interview costs you the question
- Claims a shorter timeout is strictly better
- Confuses the heartbeat interval with the election timeout
- Thinks the setting bounds the whole failover, not detection
- Believes it is a per-client or per-driver option
- Ignores that spurious elections can cause rollbacks