skip to content

How do refineCollectionShardKey and reshardCollection differ when a MongoDB shard key proves wrong?

level: seniorimportance: should knowfreq 45%

answer

  1. One appends, the other replaces
  2. The old key must stay in front
  3. Metadata change versus full data rewrite
  4. Cheap and gradual, or expensive and immediate
  5. Refine for cardinality, reshard for the wrong prefix

basics

~20 s

refineCollectionShardKey only appends suffix fields, keeping the current key as a prefix, and is a metadata change that moves no data. reshardCollection replaces the key entirely by rewriting the collection, at real disk, I/O and cut-over cost.

solid answer

~50 s

`refineCollectionShardKey`, available since MongoDB 4.4, adds one or more fields to the end of the existing shard key; the old key must remain a prefix, and you must create an index on the new longer key first. It changes metadata only — no documents move, existing range boundaries are untouched, and the benefit arrives gradually as future splits and migrations use the finer key. It cures low cardinality and indivisible ranges; it cannot cure a hotspot caused by the leading field, because placement is still driven by that prefix. `reshardCollection`, available since MongoDB 5.0, replaces the key outright, including the leading field, by rewriting the collection under the new key while the workload runs and then cutting over. It needs spare disk and I/O headroom, takes time proportional to data size, and has a short write-blocking critical section at the end. Refine when the prefix is right; reshard when it is not.

code

javascript · 12 lines
javascript
// Cheap: extend the key so ranges become splittable (prefix must be kept)
db.orders.createIndex({ tenantId: 1, orderId: 1 })
db.adminCommand({
  refineCollectionShardKey: "shop.orders",
  key: { tenantId: 1, orderId: 1 }
})

// Expensive: replace the key outright, including its leading field
db.adminCommand({
  reshardCollection: "shop.orders",
  key: { customerId: 1, orderId: 1 }
})

go deeper

for a junior

Know that a MongoDB shard key is not casually changeable, and that there are two distinct operations: one that extends the key and one that replaces it by rewriting the collection.

for a middle

Be able to state the prefix constraint on refinement, the index you must build first, and the fact that refinement moves no data while resharding copies the whole collection.

for a senior

Interviewers expect you to match the remedy to the symptom — refine for indivisible ranges, reshard for a monotonic or unqueried prefix — and to name resharding's disk, I/O and write-blocking cut-over costs before recommending it.

for a principal

Own the planning consequence: because the comprehensive fix costs a full rewrite, candidate keys should be validated against real data and query traffic before sharding, and large collections should have a rehearsed resharding runbook.

## Why two commands exist A shard key can be wrong in two quite different ways. It can be *too coarse* — the leading fields are the right dimensions for your queries, but there are not enough distinct values to divide the data finely, so ranges become indivisible and one shard grows unbounded. Or it can be *the wrong dimension entirely* — the leading field is monotonic, or it is absent from every query filter, so no amount of extra precision at the end will help. MongoDB offers a cheap fix for the first and an expensive one for the second, and choosing correctly is the point of the question. ## refineCollectionShardKey Available since MongoDB 4.4, this command appends fields to an existing shard key. The constraint that defines it is that the current key must remain a *prefix* of the new one: `{ status: 1 }` can become `{ status: 1, orderId: 1 }`, but it cannot become `{ orderId: 1, status: 1 }`. Before running it you must create an index supporting the new, longer key. What it actually does is change the cluster's metadata description of the collection. It does not read or rewrite documents, it does not move ranges, and it does not change existing range boundaries. The effect is forward-looking: once the metadata records the longer key, MongoDB may place new boundaries using the suffix, so a range that previously contained one `status` value and could not be cut is now divisible at `orderId` values. Over time, as splits and migrations occur, the collection redistributes. The consequences follow directly. It is fast and low-risk, because nothing is copied. It does not fix skew immediately — teams who expect the imbalance to vanish the moment the command returns are disappointed. And it is powerless against a monotonic or unqueried prefix, because routing and placement still begin with the leading field. ## reshardCollection Available since MongoDB 5.0, this command changes the shard key to an arbitrary new one, prefix included. Conceptually it builds a new copy of the collection distributed according to the new key, keeps it up to date with ongoing writes as it goes, and then cuts over to it. Because it genuinely rewrites the data, its costs are real. Each shard needs free disk space to hold the incoming copy of what it will own; the operation consumes I/O and CPU that competes with the live workload; and its duration scales with collection size, which for a large collection means hours. Near the end there is a critical section during which writes to the collection are blocked while the cut-over completes — brief, but it must be planned for rather than discovered. On the other hand, the collection remains readable and writable for the bulk of the operation, so it is not the maintenance window a dump-and-reload would be. ## Choosing between them Ask what is broken. If the symptom is an indivisible, oversized range or a shard that keeps growing because a value is dominant, and your queries do carry the leading field, refine. Appending a high-cardinality suffix such as `_id` or a natural unique id gives the range map somewhere to cut, at almost no operational cost. If the symptom is that every insert lands on one shard because the leading field is a timestamp or an `ObjectId`, refinement will not touch the problem — reshard to a key with a non-monotonic prefix. Likewise, if the collection's real queries never filter on the leading field, so every read broadcasts, only a new prefix fixes that. A useful rule of thumb: refinement changes how finely the data can be *divided*; resharding changes where the data *goes*. ## Operating either one safely For refinement, build the supporting index first and verify it, then run the command, then expect gradual convergence — monitor `db.collection.getShardDistribution()` over days, not minutes. For resharding, check free space on every shard against the collection's size before starting, run it in a low-traffic window even though it is online, watch its progress, and make sure the application tolerates a short period where writes to that collection block. Both are worth rehearsing on a copy of production-sized data before doing them for real. ## The framing that scores well The strongest answers do not just list the two commands — they treat them as an escape hatch of last resort and note that neither is free. Refinement is cheap but limited and slow-acting; resharding is comprehensive but expensive and needs planning. That asymmetry is exactly why teams are expected to evaluate candidate shard keys against real data and real query traffic *before* sharding, rather than treating the choice as reversible.

  • Your shard key is { createdAt: 1 } and inserts hotspot. Will refining to { createdAt: 1, deviceId: 1 } help?
    No. Refinement keeps `createdAt` as the prefix, so placement still follows the timestamp and every new document lands in the top range. The refinement only makes that range divisible after the fact. To spread the writes you need a non-monotonic leading field, which means `reshardCollection` to something like `{ deviceId: 1, createdAt: 1 }`.
  • What has to be in place before you can run refineCollectionShardKey?
    An index supporting the new, longer shard key must exist first — MongoDB will not create it for you as part of the refinement. The existing shard key must also remain a prefix of the new key, so you can only append fields, never reorder or replace the leading ones.
  • Is a collection usable while reshardCollection runs?
    Mostly yes. The collection stays readable and writable while the data is copied and kept in sync under the new key, which is the whole point of online resharding. Near the end there is a short critical section where writes block while the cut-over happens. Plan capacity for the copy — spare disk and I/O headroom — and expect duration proportional to collection size.

saying these in an interview costs you the question

  • Thinks refinement can reorder or replace the leading shard-key field
  • Expects refinement to redistribute existing data immediately
  • Describes resharding as a metadata-only or instant operation
  • Assumes resharding requires full cluster downtime
  • Treats the shard key as freely changeable with no planning

context