skip to content

How would you pin European customers' documents to shards in Frankfurt using MongoDB zone sharding?

level: seniorimportance: should knowfreq 42%

answer

  1. a label on shards plus key ranges
  2. the field must be in the shard key
  3. two commands, one for shards, one for ranges
  4. the balancer enforces it, not the router
  5. placement is not residency

basics

~20 s

Tag the Frankfurt shards with a zone name using sh.addShardToZone, then map the shard-key ranges that hold European rows to that zone with sh.updateZoneKeyRange. The balancer then migrates those ranges onto the zoned shards and keeps them there.

solid answer

~50 s

Zone sharding is a placement constraint layered on the balancer. Three things must line up: 1. **The shard key must start with the field you want to partition on.** Zone ranges are expressed in shard-key values, so a residency field such as `region` has to be the leading field — for example `{region: 1, customerId: 1}`. 2. **Associate shards with a zone**: `sh.addShardToZone("shard-fra-0", "EU")` for each Frankfurt shard. 3. **Map key ranges to the zone**: `sh.updateZoneKeyRange("app.customers", {region: "EU", customerId: MinKey}, {region: "EU", customerId: MaxKey}, "EU")`. The lower bound is inclusive, the upper exclusive. The balancer, when running, then moves those ranges onto EU-zoned shards and refuses to place them elsewhere. Definitions live in `config.tags` on the config servers. Zones control *placement only*. Real data residency also requires that every member of those shards' replica sets — including secondaries and any backups — sits in the target region.

code

javascript · 14 lines
javascript
// shard key must lead with the field you zone on
sh.shardCollection("app.customers", { region: 1, customerId: 1 })

sh.addShardToZone("shard-fra-0", "EU")
sh.addShardToZone("shard-fra-1", "EU")

sh.updateZoneKeyRange(
  "app.customers",
  { region: "EU", customerId: MinKey },   // inclusive
  { region: "EU", customerId: MaxKey },   // exclusive
  "EU"
)

sh.status()

go deeper

for a junior

Know that MongoDB can pin parts of a sharded collection to specific shards by labelling those shards with a zone name and mapping shard-key ranges to it.

for a middle

Be ready to name both steps — associating shards with a zone and associating key ranges with it — and to explain why the field you zone on must be the leading part of the shard key.

for a senior

Show that you would plan the resulting migration wave into a balancer window, cover the whole key space so nothing leaks onto reserved shards, and verify placement in sh.status() afterwards.

for a principal

Own the distinction between placement and compliance: residency obligations cover every replica-set member, backup and downstream copy, and zone sharding is one control in that story rather than the whole of it.

## What a zone is A zone is a label applied to shards, plus one or more shard-key ranges associated with that label. The balancer honours it as a constraint: a range covered by a zone may only live on a shard in that zone. Nothing else changes — routing, indexes and query semantics are identical. Zones were called *tags* in older versions, which is why the metadata still lives in `config.tags` and shards carry a `tags` array in `config.shards`. ## The three moving parts **Shard key.** Zone ranges are ranges of shard-key values, so whatever you want to pin must be expressible as a prefix of the shard key. If you want per-region placement, shard on something like `{region: 1, customerId: 1}`. You cannot zone on a field that is not in the shard key, and you cannot retrofit this without resharding. **Shard-to-zone association.** `sh.addShardToZone("shard-fra-0", "EU")`, repeated for every shard in the region. A shard may belong to several zones, and a zone may span several shards. **Range-to-zone association.** `sh.updateZoneKeyRange(ns, min, max, zone)` where `min` and `max` are full shard-key documents. Bounds follow the usual convention: lower inclusive, upper exclusive. Using `MinKey` and `MaxKey` for the trailing fields covers every value of the leading field: ```javascript sh.addShardToZone("shard-fra-0", "EU") sh.addShardToZone("shard-fra-1", "EU") sh.updateZoneKeyRange( "app.customers", { region: "EU", customerId: MinKey }, { region: "EU", customerId: MaxKey }, "EU" ) ``` Removing is symmetric: `sh.removeRangeFromZone(ns, min, max)` and `sh.removeShardFromZone(shard, zone)`. ## What happens next, and when Nothing moves instantly. Zone assignment is a rule the **balancer** enforces on its next rounds, so if the balancer is stopped, or its active window is closed, the data stays where it is and `sh.status()` shows ranges sitting outside their zone. On a large collection, converging can take hours or days, bounded by the one-migration-per-shard concurrency limit. Ranges that fall outside **every** zone range are unconstrained: the balancer may place them on any shard, zoned or not. This trips people up on partial rollouts — you zone `EU` but forget `US`, and US data quietly lands on the Frankfurt shards. If a zone exists but no shard belongs to it, the covered ranges cannot be placed there and simply stay put. ## Zones are placement, not access control This is the point most worth making in an interview. A zone decides which shard stores a document. It does not: - **Guarantee residency by itself.** A shard is a replica set; if one of its secondaries is in another country, the data is in another country. Residency means constraining every member, plus backups, oplog copies and any analytics replication. - **Restrict who can read it.** Authorization is unchanged; any authenticated client with rights on the collection reads EU rows from wherever it connects. - **Localize latency automatically.** Clients still connect to a `mongos` and the query still travels to the owning shard. A user in Frankfurt hitting a router in Virginia gets a transatlantic round trip regardless of zoning. Local routers plus a shard key that keeps a request on one shard are what buy locality. - **Prevent cross-zone queries.** A query without the leading shard-key field still scatters to every shard, including ones in other regions. ## Operating zones `sh.status()` prints zones and their ranges per collection. Before adding a zone to an existing cluster, work out how much data the change will force across the network and put it inside a balancer window. When you retire a region, remove the ranges first and then the shard-zone associations, and watch the balancer drain it — removing associations while ranges are still mapped just leaves ranges with nowhere legal to go. Zones also give you a clean way to introduce heterogeneous hardware: label a few large-disk shards as an archive zone and map historical key ranges to them, keeping recent data on faster shards. The mechanics are exactly the same; only the meaning of the leading shard-key field changes.

  • A document's shard key falls outside every zone range you defined. Where does it live?
    Anywhere. Unzoned ranges are unconstrained, so the balancer places them on whichever shard evens out the collection — including shards that belong to a zone. If you need strict separation you must cover the whole key space with zones, which is why partial zone rollouts often leak data onto the shards you meant to reserve.
  • You add a zone to a collection that already holds a terabyte per region. What should you expect?
    A long rebalance, not an instant move. The balancer migrates the affected ranges over its next rounds, limited to one migration per shard at a time and to whatever active window you configured. Plan it as a scheduled data-movement project, watch config.changelog for progress, and expect the collection to look mis-zoned in sh.status() until it converges.
  • Does zone sharding reduce latency for users in that region?
    Only indirectly. Zoning decides which shard stores the data; the request path still runs client to mongos to shard. You get local latency when a router runs near the users and the query carries the shard key so it targets the local shard. Without both, a zoned deployment can still make every read cross an ocean.

saying these in an interview costs you the question

  • Says a zone can be defined on any indexed field
  • Believes zone assignment moves data immediately
  • Treats zone placement as proof of data residency
  • Forgets that unzoned key ranges can land on zoned shards
  • Thinks zones restrict which clients may read the data

context