If a MongoDB sharded cluster's config server replica set loses its majority, what stops working?
answer
- metadata lives in its own replica set
- no primary means no metadata writes
- routers keep a cached routing table
- the risk is anything that must refresh
- never restart routers during the incident
basics
~20 sMetadata writes stop: no range splits, no migrations or balancing, no sharding a new collection, no adding or removing shards, no zone changes. Reads and writes to existing sharded collections keep working while routers still hold a usable cached routing table.
solid answer
~50 sThe config server replica set (CSRS) holds the cluster's authoritative metadata in the `config` database: `config.shards`, `config.databases`, `config.collections`, `config.chunks`, `config.tags` and `config.settings`. Metadata changes are written with majority write concern, so without a primary they simply cannot happen. What stops: the balancer, range splits and migrations, `shardCollection`, creating a new database, `addShard` / `removeShard`, and any zone change. What continues: normal reads and writes against already-sharded and unsharded collections, because `mongos` caches the routing table and each shard caches its own ownership metadata. The cached table stays correct precisely because nothing can change it. The exposure is any router that must **refresh** — one starting up, or one that has received a stale-config error — since it cannot rebuild its routing table without reachable config servers. Deploy the CSRS as three members across separate failure domains and monitor it as a first-class production system.
code
javascript · 5 linesuse config
db.settings.find() // balancer state, active window, chunksize
db.shards.find() // shards and their zone tags
db.databases.find() // primary shard per database
db.changelog.find({ what: /^moveChunk/ }).sort({ time: -1 }).limit(5)go deeper
Know that a sharded MongoDB cluster keeps its map of which shard holds which data on a separate set of config servers, and that they are a required component.
Be able to name what the config servers store and explain why routers can keep serving from a cached routing table when metadata writes are unavailable.
Distinguish the operations that halt from the traffic that continues, and know the incident rule: freeze deploys and router restarts until the config servers have a primary again.
Treat the config server replica set as the metadata plane whose availability bounds every sharded cluster: three members across failure domains, monitored and backed up together with the shards, with the balancer's silence wired to an alert.
## What the config servers hold A sharded cluster's config servers form their own replica set (the CSRS) and store the `config` database. That database is small but load-bearing: - `config.shards` — the shards and the zones each belongs to - `config.databases` — which shard is primary for each database - `config.collections` — which collections are sharded and on what shard key - `config.chunks` — one document per range: bounds, owning shard, version - `config.tags` — zone-to-key-range mappings - `config.settings` — balancer state, active window, configured range size - `config.changelog` — the audit trail of splits and migrations Every change here is written with majority write concern, which is what makes the routing table trustworthy. ## What breaks when the majority is lost No majority means no primary, which means no metadata writes. Concretely: - **Balancing stops.** No new migrations start, and an in-flight one cannot commit its ownership change. - **Splits stop.** The balancer cannot carve off a range to move. - **DDL stops.** `sh.shardCollection()` fails; creating a new database fails, because that requires writing a `config.databases` entry naming its primary shard. - **Topology changes stop.** `addShard`, `removeShard` and draining are all metadata operations. - **Zone changes stop.** `addShardToZone` and `updateZoneKeyRange` write to `config.shards` and `config.tags`. ## What keeps working Ordinary traffic. Every `mongos` caches the routing table, and every shard caches the ownership metadata it needs to filter orphans out of results. Reads and writes to existing collections continue against the shards, which are separate replica sets with their own primaries and elections and are unaffected by the config servers' state. There is a pleasing safety property here: the cached routing table cannot go stale, because nothing is able to change the authoritative copy. The cluster is frozen in a consistent shape rather than drifting. ## Where the real exposure is The danger is any component that needs to **build or refresh** its routing table: - A `mongos` that restarts starts with an empty cache and must read the config servers before it can route anything. In an outage this is how a partial failure becomes total: a routine deploy or a crash-loop cycles routers, and each one comes back unable to serve. - A router that receives a stale-config error must refresh; if the config servers are unreachable it cannot, and the operation fails. - Any first access to a collection whose metadata is not yet cached needs a read from the config servers. So the practical rule during a CSRS incident is: **do not restart your routers, and do not deploy**. Freeze changes, restore the config server majority, and let the cluster recover on its own. ## Degrees of failure "Lost majority" is not the same as "unreachable". If members are up but cannot elect a primary, metadata reads may still be served while writes are impossible. If the CSRS is entirely unreachable from a router, that router cannot refresh at all. Total and permanent loss of the config servers is the worst outcome in a sharded deployment: the shards still hold every document, but nothing knows which shard owns which range, and reconstructing that from a stale backup risks contradicting the shards' own filtering metadata. ## Operating the CSRS - Run **three members across separate failure domains** — separate hosts, racks or availability zones. Two members in one rack means a rack loss costs you the majority. - **Monitor it like a production database**: primary presence, election counts, replication lag, disk. - **Back it up** as part of the cluster backup, and test the restore path alongside the shards, since a shard backup without its matching metadata is much harder to use. - **Alert on "balancer has not run"** — a quiet balancer is a common first visible symptom of a config-server problem, well before anyone notices a failed `shardCollection`. - Keep it modest in size but never treat it as unimportant: it is small, low-traffic, and the single component whose loss degrades every other part of the cluster.
- Why is restarting a mongos during a config-server outage a bad idea?A router builds its routing table by reading the config servers at startup and holds it in memory. Restarting throws away the cache that is keeping it useful, and it cannot rebuild it while the config servers are unreachable, so a router that was serving traffic fine comes back unable to route anything. Freeze deploys until the CSRS has a primary again.
- Which config collection would you inspect to see whether the balancer is disabled or windowed?config.settings on the config servers. The document with _id "balancer" carries the stopped flag and any activeWindow, and the document with _id "chunksize" carries the cluster-wide range size. sh.getBalancerState() reads the same information, but knowing where it lives helps when you are diagnosing why nothing is migrating.
- How many config server members would you deploy, and where?Three, in three separate failure domains, so losing any one host, rack or availability zone still leaves a majority able to elect a primary. Co-locating two of them defeats the purpose: the cluster's metadata plane then has the same blast radius as a single rack, and every migration, split and DDL operation depends on it.
saying these in an interview costs you the question
- Says the whole cluster goes read-only or offline immediately
- Thinks mongos stores the routing table durably on its own disk
- Believes shards cannot serve queries without the config servers
- Restarts routers to try to clear a config-server problem
- Runs config server members inside a single failure domain