When a keyspace is split so each key lives on one node, why must every caller resolve keys the same way?
answer
- one owner per key
- the rule has to be shared
- a fixed part and a changing part
- disagreement splits one key in two
basics
~20 sAssignment must be a deterministic function of the key, evaluated identically by every caller. If two callers disagree about where a key belongs, each reads and writes its own copy on a different node and neither sees the other's.
solid answer
~50 sSplitting a keyspace means each key has exactly one owning node, and the only thing that makes that true is a rule everyone evaluates the same way: a fixed part that turns a key into a `partition`, and a changing part that assigns partitions to nodes. Together those are the partition map. Determinism is required against a given version of the assignment, not forever — assignments move when nodes are added or fail, and what matters is that nobody acts on two versions at once. Stores differ in how hard they protect this: some nodes track which partitions they own and answer that a key lives elsewhere, others accept whatever arrives. Where they do not check, disagreement produces no error at all — just one key living as two independent entries. The shared rule is also what makes growth by adding nodes possible; without it, extra machines can only hold copies.
go deeper
Recall that splitting a keyspace gives each key exactly one owning node, and that this only works if every caller computes the same owner for a key.
Explain the two halves — a fixed key-to-partition function and a changing assignment of partitions to nodes — and why only the second moves when the tier grows.
Show what disagreement costs in production: two live entries for one key, a counter counted twice, a claim held twice, and no error anywhere to tell you.
Frame it as a fleet property: where callers resolve keys, correctness of the rule lives in every client version you deploy, which is a governance decision as much as a technical one.
## Why a partitioned tier needs a rule at all A volatile tier that no longer fits on one machine has two different answers available, and they are not variations of each other. One is to put a **copy of the whole keyspace** on more machines, which adds read capacity and survivability but no room. The other is to cut the keyspace into **partitions** and give each node a share, which is the only one of the two that lets the tier hold more than one machine's memory. The second answer needs something the single-machine design never needed: an answer to *which node holds this key*, available wherever a request begins. That answer is a rule, and it has to be **deterministic** — the same key resolves to the same partition every time — and it has to be **evaluated identically everywhere**. Those are two separate requirements, and the second is the one real deployments break. ## Ownership is a convention, not a physical fact Nothing stops most nodes in this class from storing any key handed to them. Ownership exists only because every party that resolves a key computes the same owner. Stores differ in how far they go to protect that: - some nodes track which partitions they own and refuse, or answer that the key lives elsewhere; - others accept whatever arrives, so a wrongly routed write is stored and a wrongly routed read reports the key as simply absent; - and in the shapes where the caller never resolves at all — an intervening proxy, or a directory consulted per lookup — the convention is enforced in one place instead of in every caller. That variation is worth stating plainly: whether disagreement between callers surfaces as an error or as silence is a property of the store you bought, not of the idea. ## The rule has two halves It helps to separate the part that is fixed from the part that moves. 1. **Key to partition.** A pure function of the key's bytes. It does not know the node count, does not change when a machine is added, and gives the same answer in every process. 2. **Partition to node.** The current assignment. This is the part that changes — when a node is added, when one fails, when assignments are moved while the tier keeps serving. Together they are **the partition map**. Stores package the two differently: some keep the assignment as an explicit table the rule reads, others fold node membership into the rule itself so the live node set is an input. Where the division lands changes how an update propagates, but the caller-visible requirement is identical in every arrangement — at any given moment, everyone agrees. So *deterministic* means **deterministic against a version of the assignment**, not permanent. Assignments are expected to move. What must never happen is two parties acting on different versions with nothing to notice. ## What disagreement actually does Suppose two services compute different owners for one key, on a store whose nodes do not check ownership: - each service writes to its own node and reads back its own value, so both are internally consistent and neither is right; - a counter is now two counters, and the total is whatever one caller happens to see; - a record marking work as already handled exists on one node, so the caller routing elsewhere does the work a second time; - a claim held under a key exists twice, so two holders each believe they are the only one; - nothing raises an error, because nothing is in a position to compare. That is why this is stated as an absolute rather than as a good practice. Split state produced by disagreement does not look like a defect in the tier; it looks like an application that has lost its mind. ## Why the rule is what buys you nodes Growing a tier comes down to two options: more memory on each machine, or more machines each holding a share of the keyspace. The second option exists **only because there is a rule**. Without one, extra machines can hold copies, or hold whatever each of them happened to be asked for — which is a set of unrelated stores, not a larger one. The deterministic key-to-node rule is the thing that turns a set of machines into one addressable keyspace. ## Where implementations genuinely differ | Question | One arrangement | Another | |---|---|---| | Who evaluates the rule | the caller, from its own copy of the map | an intervening proxy, or a directory asked per lookup | | Is membership inside the rule | no, the assignment is a separate table | yes, the live node set is an input to the rule | | What a node does with a foreign key | serves or stores it as if it owned it | answers that the key lives elsewhere | | How a change reaches the resolver | callers refresh, on a signal or on a schedule | one component is updated and callers notice nothing | Treat none of these as the model. The claim that holds across all of them is the one this subject is about: at any moment one key has one owner, and everything that resolves keys must agree on who that is.
- Does a deterministic rule mean the assignment of partitions to nodes can never change?No. Determinism is required against the current version of the assignment. Assignments move when nodes are added or fail; what must hold is that everyone resolving at a given moment agrees. The design problem is not preventing change but making sure a change reaches every place that resolves keys before those places act on the old one.
- Two client libraries in different languages reach one partitioned tier. What must you verify?That both implement the same key-to-partition rule exactly, including how the key's bytes are taken, and that both hold the same assignment. Where the store expects the caller to resolve, this becomes a fleet-wide correctness property that no single component can enforce. Behind an intervening proxy or a directory lookup it is not the caller's problem at all.
saying these in an interview costs you the question
- Thinks the store finds the owning node for you at write time
- Assumes any node can answer for any key once nodes are joined
- Treats disagreement between callers as an error someone reports
- Thinks each caller may use its own rule if it is self-consistent
- Believes adding nodes spreads an existing keyspace with no rule
- Confuses copies of the whole keyspace with pieces of it