How would you pin European customers' documents to shards in Frankfurt using MongoDB zone sharding?
answer
- a label on shards plus key ranges
- the field must be in the shard key
- two commands, one for shards, one for ranges
- the balancer enforces it, not the router
- placement is not residency
basics
~20 sTag the Frankfurt shards with a zone name using sh.addShardToZone, then map the shard-key ranges that hold European rows to that zone with sh.updateZoneKeyRange. The balancer then migrates those ranges onto the zoned shards and keeps them there.
solid answer
~50 sZone sharding is a placement constraint layered on the balancer. Three things must line up: 1. **The shard key must start with the field you want to partition on.** Zone ranges are expressed in shard-key values, so a residency field such as `region` has to be the leading field — for example `{region: 1, customerId: 1}`. 2. **Associate shards with a zone**: `sh.addShardToZone("shard-fra-0", "EU")` for each Frankfurt shard. 3. **Map key ranges to the zone**: `sh.updateZoneKeyRange("app.customers", {region: "EU", customerId: MinKey}, {region: "EU", customerId: MaxKey}, "EU")`. The lower bound is inclusive, the upper exclusive. The balancer, when running, then moves those ranges onto EU-zoned shards and refuses to place them elsewhere. Definitions live in `config.tags` on the config servers. Zones control *placement only*. Real data residency also requires that every member of those shards' replica sets — including secondaries and any backups — sits in the target region.
code
javascript · 14 lines// shard key must lead with the field you zone on
sh.shardCollection("app.customers", { region: 1, customerId: 1 })
sh.addShardToZone("shard-fra-0", "EU")
sh.addShardToZone("shard-fra-1", "EU")
sh.updateZoneKeyRange(
"app.customers",
{ region: "EU", customerId: MinKey }, // inclusive
{ region: "EU", customerId: MaxKey }, // exclusive
"EU"
)
sh.status()go deeper
Know that MongoDB can pin parts of a sharded collection to specific shards by labelling those shards with a zone name and mapping shard-key ranges to it.
Be ready to name both steps — associating shards with a zone and associating key ranges with it — and to explain why the field you zone on must be the leading part of the shard key.
Show that you would plan the resulting migration wave into a balancer window, cover the whole key space so nothing leaks onto reserved shards, and verify placement in sh.status() afterwards.
Own the distinction between placement and compliance: residency obligations cover every replica-set member, backup and downstream copy, and zone sharding is one control in that story rather than the whole of it.
## What a zone is A zone is a label applied to shards, plus one or more shard-key ranges associated with that label. The balancer honours it as a constraint: a range covered by a zone may only live on a shard in that zone. Nothing else changes — routing, indexes and query semantics are identical. Zones were called *tags* in older versions, which is why the metadata still lives in `config.tags` and shards carry a `tags` array in `config.shards`. ## The three moving parts **Shard key.** Zone ranges are ranges of shard-key values, so whatever you want to pin must be expressible as a prefix of the shard key. If you want per-region placement, shard on something like `{region: 1, customerId: 1}`. You cannot zone on a field that is not in the shard key, and you cannot retrofit this without resharding. **Shard-to-zone association.** `sh.addShardToZone("shard-fra-0", "EU")`, repeated for every shard in the region. A shard may belong to several zones, and a zone may span several shards. **Range-to-zone association.** `sh.updateZoneKeyRange(ns, min, max, zone)` where `min` and `max` are full shard-key documents. Bounds follow the usual convention: lower inclusive, upper exclusive. Using `MinKey` and `MaxKey` for the trailing fields covers every value of the leading field: ```javascript sh.addShardToZone("shard-fra-0", "EU") sh.addShardToZone("shard-fra-1", "EU") sh.updateZoneKeyRange( "app.customers", { region: "EU", customerId: MinKey }, { region: "EU", customerId: MaxKey }, "EU" ) ``` Removing is symmetric: `sh.removeRangeFromZone(ns, min, max)` and `sh.removeShardFromZone(shard, zone)`. ## What happens next, and when Nothing moves instantly. Zone assignment is a rule the **balancer** enforces on its next rounds, so if the balancer is stopped, or its active window is closed, the data stays where it is and `sh.status()` shows ranges sitting outside their zone. On a large collection, converging can take hours or days, bounded by the one-migration-per-shard concurrency limit. Ranges that fall outside **every** zone range are unconstrained: the balancer may place them on any shard, zoned or not. This trips people up on partial rollouts — you zone `EU` but forget `US`, and US data quietly lands on the Frankfurt shards. If a zone exists but no shard belongs to it, the covered ranges cannot be placed there and simply stay put. ## Zones are placement, not access control This is the point most worth making in an interview. A zone decides which shard stores a document. It does not: - **Guarantee residency by itself.** A shard is a replica set; if one of its secondaries is in another country, the data is in another country. Residency means constraining every member, plus backups, oplog copies and any analytics replication. - **Restrict who can read it.** Authorization is unchanged; any authenticated client with rights on the collection reads EU rows from wherever it connects. - **Localize latency automatically.** Clients still connect to a `mongos` and the query still travels to the owning shard. A user in Frankfurt hitting a router in Virginia gets a transatlantic round trip regardless of zoning. Local routers plus a shard key that keeps a request on one shard are what buy locality. - **Prevent cross-zone queries.** A query without the leading shard-key field still scatters to every shard, including ones in other regions. ## Operating zones `sh.status()` prints zones and their ranges per collection. Before adding a zone to an existing cluster, work out how much data the change will force across the network and put it inside a balancer window. When you retire a region, remove the ranges first and then the shard-zone associations, and watch the balancer drain it — removing associations while ranges are still mapped just leaves ranges with nowhere legal to go. Zones also give you a clean way to introduce heterogeneous hardware: label a few large-disk shards as an archive zone and map historical key ranges to them, keeping recent data on faster shards. The mechanics are exactly the same; only the meaning of the leading shard-key field changes.
- A document's shard key falls outside every zone range you defined. Where does it live?Anywhere. Unzoned ranges are unconstrained, so the balancer places them on whichever shard evens out the collection — including shards that belong to a zone. If you need strict separation you must cover the whole key space with zones, which is why partial zone rollouts often leak data onto the shards you meant to reserve.
- You add a zone to a collection that already holds a terabyte per region. What should you expect?A long rebalance, not an instant move. The balancer migrates the affected ranges over its next rounds, limited to one migration per shard at a time and to whatever active window you configured. Plan it as a scheduled data-movement project, watch config.changelog for progress, and expect the collection to look mis-zoned in sh.status() until it converges.
- Does zone sharding reduce latency for users in that region?Only indirectly. Zoning decides which shard stores the data; the request path still runs client to mongos to shard. You get local latency when a router runs near the users and the query carries the shard key so it targets the local shard. Without both, a zoned deployment can still make every read cross an ocean.
saying these in an interview costs you the question
- Says a zone can be defined on any indexed field
- Believes zone assignment moves data immediately
- Treats zone placement as proof of data residency
- Forgets that unzoned key ranges can land on zoned shards
- Thinks zones restrict which clients may read the data