skip to content

When should you choose a hashed shard key in MongoDB instead of a ranged one, and what do you give up?

level: middleimportance: should knowfreq 68%

answer

  1. Partition on the value, or on something derived from it
  2. Cures the ever-increasing key problem
  3. Order is destroyed, equality is not
  4. One query shape becomes a broadcast
  5. Compound keys may hold one such field

basics

~20 s

Choose hashed sharding when the natural key is monotonic or unevenly distributed and the workload is mostly equality lookups. You give up range targeting: a range filter on that field must be broadcast to every shard.

solid answer

~50 s

With `sh.shardCollection("db.coll", { field: "hashed" })`, MongoDB partitions on the *hash* of the value rather than the value, so neighbouring values scatter across shards. That is the standard cure for a monotonic key like a timestamp or a default `ObjectId`, and for keys whose natural ordering clusters writes. Equality queries still target a single shard, because a given value always hashes to the same number. What you lose is everything that depends on value order: a `$gt`/`$lt` range on the hashed field can no longer be narrowed to a few shards, sorts on it gain nothing from placement, and zone sharding by value range on that field is meaningless. Since MongoDB 4.4 a compound shard key may contain exactly one hashed field, which lets you keep an even spread on the prefix while retaining a second, ordered field.

code

javascript · 8 lines
javascript
// Even insert spread; equality on _id still targets one shard
sh.shardCollection("iot.readings", { _id: "hashed" })

// Compound key with exactly one hashed field (MongoDB 4.4+)
sh.shardCollection("iot.readings", { deviceId: "hashed", ts: 1 })

// Ranged compound alternative: spreads writes AND keeps time windows targeted per device
sh.shardCollection("iot.readings", { deviceId: 1, ts: 1 })

go deeper

for a junior

Recall that the shard key spec chooses the strategy: a value of 1 partitions by value, the string "hashed" partitions by the hash of the value. Know that hashing is the usual answer to an ever-increasing key.

for a middle

Be able to explain that hashing preserves equality targeting but destroys range targeting, and give one workload for each strategy. Mention that a compound key can contain a single hashed field.

for a senior

Show that you would choose from the access patterns rather than by reflex — and be ready to argue that a compound ranged key often beats hashing for time-series data that is both written and read by time.

for a principal

Own the framing that this choice moves cost between the write path and the read path rather than removing it, and be able to say which of the two your system can afford to pay and why.

## The two partitioning strategies When you shard a collection you pick not just which field or fields form the key, but how MongoDB derives ranges from them. Ranged sharding divides the key space into contiguous intervals of the actual values: shard A owns keys from `MinKey` to 1000, shard B owns 1000 to 5000, and so on. Hashed sharding applies a hash function to the key value and divides the space of *hash* outputs instead. You choose it in the shard key specification itself: `{ userId: 1 }` is ranged, `{ userId: "hashed" }` is hashed. The underlying index differs too — a hashed shard key is backed by a hashed index. ## What hashing buys you The hash of a monotonically increasing sequence is not monotonic. That single property is the main reason hashed sharding exists. Sharding a high-volume event collection on `{ _id: 1 }` when `_id` is a default `ObjectId` funnels every insert into the top range and thus onto one shard, because `ObjectId` values encode creation time in their leading bytes and therefore rise over time. Sharding on `{ _id: "hashed" }` scatters consecutive inserts across the whole hash space, so writes land roughly uniformly on all shards from the first document. The same applies to any natural key with clumpy ordering — a country code prefix, a sequential invoice number, a device serial that increments per batch. A second benefit is that even distribution is achieved without you designing a compound key or pre-splitting anything: the hash function does the spreading. ## What hashing costs you Hashing destroys order, and everything that depended on order goes with it. **Range targeting.** With a ranged key on `createdAt`, a query for the last hour hits only the shards owning the top few ranges. With `{ createdAt: "hashed" }`, that hour's documents are scattered uniformly, so `mongos` must ask every shard and merge the results. For a dashboard whose whole workload is recent-window queries, hashing the timestamp trades a write hotspot for a read fan-out on every query. **Zone-based placement by value.** Pinning a value range to specific shards depends on the shard key preserving order, so it does not work on a hashed field. **Equality is still fine.** This is the part candidates most often get wrong in both directions. A hash function is deterministic, so `find({ userId: 42 })` on a `{ userId: "hashed" }` collection computes one hash and routes to the single shard owning it. Hashed sharding does *not* make every query a broadcast — only order-dependent ones. ## Compound hashed keys Since MongoDB 4.4 a compound shard key may include exactly one hashed field, in either the prefix or a later position. `{ deviceId: "hashed", ts: 1 }` spreads devices evenly across the cluster while giving each device's documents an ordered suffix that ranges can be split on, so no single device's data becomes indivisible. `{ tenantId: 1, sessionId: "hashed" }` keeps a tenant's data in an ordered, zone-pinnable prefix while ensuring the tenant's own documents scatter within it. This flexibility removes much of the old all-or-nothing feel of the choice. ## Caveats worth knowing Hashed indexes do not support arrays: you cannot hash a field whose values are arrays, so a multikey field is not a valid hashed shard key. Hashed indexes also do not support the unique option. And hashing collapses floating-point values that differ only in their fractional part, because values are converted to 64-bit integers before hashing — so a `price` field with cents is a bad hashed key, since many distinct prices map to the same hash and the resulting group becomes indivisible. ## How to decide Write the collection's real access patterns down, then ask two questions. First, does the natural key increase over time or cluster its values? If yes, ranged sharding on it alone will hotspot. Second, does the workload depend on value-ordered access — recent-window scans, ordered pagination across the key, zone placement? If yes, hashing that field pushes the cost onto every read. When both answers are "yes" — a time-series workload that both inserts monotonically and queries by time window — the usual answer is neither pure form but a compound ranged key whose prefix is a non-monotonic, high-cardinality dimension, such as `{ deviceId: 1, ts: 1 }`. Each device's writes spread by device, and a query for one device's last hour is still targeted. Reaching for a compound key rather than defaulting to hashed is what separates a considered answer from a memorised one.

  • With a hashed shard key on userId, can mongos still route find({ userId: 42 }) to a single shard?
    Yes. Hashing is deterministic, so `mongos` computes the hash of 42 and routes to the shard that owns that hash range. Hashed sharding only breaks routing for order-dependent predicates such as `$gt`/`$lt` on the hashed field, and for sorts or zone placement that rely on value order. Equality and `$in` on the key remain targeted.
  • Why is a monetary amount with cents a poor hashed shard key?
    Hashed indexes convert values to 64-bit integers before hashing, so numbers differing only in their fractional part collide onto the same hash. Many distinct prices then share one hash value, and documents sharing a hash value land in one indivisible group — reintroducing exactly the skew you used hashing to avoid.
  • You shard an events collection on { ts: "hashed" } and the last-hour dashboard slows down. What happened?
    The hour's documents are now uniformly scattered, so the router cannot narrow the query by shard key and must fan out to every shard and merge. The insert hotspot is gone but every read pays cluster-wide cost. A compound ranged key such as `{ deviceId: 1, ts: 1 }` usually serves both sides better.

saying these in an interview costs you the question

  • Claims a hashed shard key makes all queries scatter-gather
  • Uses hashed sharding on a field the workload queries by range
  • Thinks hashed sharding removes the need to think about frequency
  • Believes hashed and ranged can be mixed freely in one key
  • Assumes hashing prevents a dominant single value from skewing shards

context