How would you choose cluster-wide default read and write concerns for a platform with mixed workloads?
answer
- one admin command sets it fleet-wide
- safe by default, downgrade by exception
- an arbiter can move the implicit value
- fix placement before weakening the promise
basics
~20 sSet a safe cluster-wide default with setDefaultRWConcern, keeping majority writes as the floor, and let individual low-value paths opt down explicitly. Tune topology before weakening durability, and watch for arbiters silently lowering the implicit default.
solid answer
~50 sStart from a **safe default and require an explicit downgrade**, not the reverse. Since MongoDB 4.4 the `setDefaultRWConcern` command sets cluster-wide defaults, and since 5.0 the implicit default write concern is already `majority`, so the policy is mostly about which workloads may go below it: telemetry, caches and derived data can justify `w: 1`, but that downgrade should live in the code that owns the collection and be reviewed, not be a global setting that quietly covers payments too. Two traps deserve attention. Arbiters can drop the implicit default to `w: 1` for a whole cluster, so a P-S-A topology is a durability decision disguised as a cost decision. And a cross-region member turns majority latency into every write's latency — the fix is member placement, not a weaker `w`. On the read side, keep `local` as the default and reserve `majority` for reads that gate irreversible actions.
code
javascript · 7 linesdb.adminCommand({
setDefaultRWConcern: 1,
defaultWriteConcern: { w: "majority", wtimeout: 5000 },
defaultReadConcern: { level: "local" }
})
db.adminCommand({ getDefaultRWConcern: 1 })go deeper
Know that the durability of a write is a setting someone chose, that MongoDB has a cluster-wide default, and that your code can override it per operation.
Be able to explain what setDefaultRWConcern controls, how per-operation settings override it, and why the implicit default can differ in a set that contains arbiters.
Argue the tradeoff with evidence: measure what majority acknowledgment actually costs, show why topology and member placement are the first lever, and scope any downgrade to a named collection with a stated loss tolerance.
Own the policy end to end — the workload tiers and their durability contracts, how defaults are asserted and audited, how overrides stay visible in review, and what monitoring proves the promise is being kept.
## What is actually being decided The dials are few — `w`, `j` and `wtimeout` on writes, read concern level and read preference on reads — but a platform decision has three parts: what the **cluster default** is, what an application may **override**, and how you **detect** that the effective settings have drifted from the intent. ## Set defaults explicitly The `setDefaultRWConcern` administrative command (available since MongoDB 4.4) sets a cluster-wide default write concern and default read concern, and `getDefaultRWConcern` reports the current values including whether they came from your setting or from the server's implicit default. Setting them explicitly is worth doing even when it matches the implicit default, because it makes the intent auditable and survives topology changes that would otherwise move the implicit value underneath you. The design principle is **safe by default, downgrade by exception**. A default of `majority` means the failure mode of a forgetful developer is a slightly slower write, not a silently losable one. The reverse — a fast default with the important paths expected to opt up — fails in the direction that costs you data. ## Tiering by workload A workable classification: - **Ledger-like data** (payments, entitlements, audit): `w: "majority"`, always, plus a `wtimeout` so a degraded set fails visibly instead of hanging. Reads that gate an irreversible external action use `readConcern: "majority"`. - **Core application state** (users, orders, content): the `majority` default, `local` reads, primary reads on write-then-read flows. - **High-volume derived or disposable data** (clickstream, metrics, cache-like collections): a documented `w: 1` override on that collection's client, never `w: 0` unless the loss of error reporting is genuinely acceptable. - **Reporting and analytics**: unchanged write concern, reads moved to secondaries with a staleness bound and tags, so the primary is protected. The important organisational move is that the downgrade lives **next to the code that owns the data**, is visible in review, and can be greped for. A global weak default hides the same decision everywhere. ## The arbiter trap Since MongoDB 5.0 the implicit default write concern is `majority`, **except** in sets containing arbiters where the data-bearing voting members do not themselves form a majority of the voting members — there it drops to `w: 1`. A team that adds an arbiter to save the cost of a third data-bearing member has therefore changed the durability posture of every write in the cluster without touching any application code. If you run such a topology deliberately, set the default explicitly so the choice is recorded, and understand that majority writes will also stall whenever the single secondary is down. ## Latency budgets and topology Majority write latency is the round trip to the slowest member of the fastest majority. When that number is unacceptable, the reflex to weaken `w` is usually the wrong lever. Better lever order: 1. Place enough data-bearing members close to the primary that a majority can be formed locally; keep distant members as additional copies beyond the majority. 2. Reduce the write rate per document, batch, or shard the workload. 3. Only then consider a per-collection downgrade, and only for data whose loss you have written down as acceptable. Measure before deciding: track write latency percentiles with and without majority acknowledgment, and track the lag of the majority commit point behind the primary's newest write. If that lag is spiky, majority writes will be spiky no matter how you tune the application. ## What must not be weakened Some features depend on the majority commit point rather than on your preference — change streams only surface events once they are majority-committed, so an unhealthy majority commit point degrades those consumers regardless of the write concern an individual write used. In currently-shipping versions majority read concern is always enabled and cannot be turned off. ## Governance Three practices keep the policy real: - **Assert the defaults.** Have a startup check or a periodic job read the cluster defaults and alarm when they differ from the documented policy. - **Make overrides explicit and searchable.** A weak write concern should appear in code, not in a connection string buried in a deployment template where nobody reads it. - **Alert on write concern errors.** A rising rate is the earliest signal that the durability you promised is not being met, and it is far more actionable than a lagging-secondary dashboard nobody watches. The judgment being tested here is not knowledge of the syntax. It is whether you treat durability as a per-workload product decision with a written rationale, or as a latency knob that gets turned during an incident and never turned back.
- Why is a weak cluster-wide default worse than a weak per-collection override, even if the resulting settings are identical today?A global default applies to code nobody has reviewed yet, including collections created next quarter, and it is invisible at the call site. A per-collection override sits next to the data it affects, appears in code review, and can be searched for during an audit. Same behaviour today, very different failure mode as the system grows.
- What would you monitor to know that your durability policy is actually holding?The rate of write concern errors, the lag of the majority commit point behind the primary's newest write, per-member replication lag, and a periodic assertion that the cluster's configured defaults still match the documented policy. Rising write concern errors are the earliest actionable signal that promised durability is not being met.
- A team wants to move a latency-sensitive service from w: majority to w: 1. What do you ask for before agreeing?A written statement of what data loss is acceptable during a failover, evidence that the latency actually comes from majority acknowledgment rather than from indexing or the workload shape, and confirmation that topology fixes — a closer data-bearing member, batching, sharding — were considered first. If it proceeds, scope it to the specific collection, not the cluster.
saying these in an interview costs you the question
- Treats write concern as a latency knob rather than a data-loss decision
- Sets a weak cluster-wide default and expects services to opt up
- Adds an arbiter without noticing the implicit default drops to w: 1
- Weakens w to hide cross-region latency instead of fixing member placement
- Never verifies the cluster's effective defaults after topology changes