Which component resolves a key to its node in a partitioned in-memory tier, and what does each option cost?
answer
- three places a map can live
- caller, proxy, or directory
- hops versus who must be told
- staleness moves, it never vanishes
basics
~20 sThree shapes exist: the caller holds a copy of the partition map, an intervening proxy holds it, or a directory is consulted per lookup. They differ in hops, in who must learn of a change, and in who can be stale.
solid answer
~40 sWith a **caller-side map**, every calling process holds the assignment and goes straight to the owning node — one hop, the lowest latency, and the map's correctness now lives in every client in the fleet. With an **intervening proxy**, callers address one place and the proxy resolves; they never see a redirect and never hold a map, but you pay a hop and now run a tier that must be scaled, monitored and failed over. With a **directory lookup**, an authoritative service answers per lookup, which costs a round trip unless you cache the answer — and a cached answer is a caller-side map with a deadline. None of the three removes staleness; they move it. Not every store offers all three, so the shape is often chosen for you.
go deeper
Remember that something has to turn a key into a node, and it is not always the caller: it may be the caller, a proxy in front of the nodes, or a service asked per lookup.
Be able to compare the three on hops, on who has to learn that an assignment moved, and on who is capable of noticing that the answer is out of date.
Argue from the propagation column rather than the latency column, and say explicitly that a cached directory answer is a caller-side map with a deadline.
Treat the shape as a fleet and procurement decision: it fixes how many places can hold a wrong map and constrains which stores in this class you can adopt without rewriting callers.
## Three places the answer can come from Once a keyspace is cut into **partitions**, something has to turn a key into the node that owns it. There are three shapes in use, and a discussion of this subject that assumes one of them has already gone wrong. 1. **Caller-side map.** Each calling process holds a copy of the partition map — the rule plus the current assignment — and resolves locally, then connects straight to the owning node. 2. **Intervening proxy.** Callers connect to one address. A proxy tier holds the map, resolves, and forwards. Callers hold nothing and learn nothing about the topology. 3. **Directory lookup.** A separate service is the authority on where a key's partition currently lives, and is asked for the answer. ## What each one costs | Shape | Who holds the map | Hops to the data | How it learns of a change | The bill | |---|---|---|---|---| | Caller-side map | every calling process | one | from an answer that the key lives elsewhere, a periodic refresh, or a push | every client version in the fleet is a place correctness can rot | | Intervening proxy | the proxy tier | two | the proxy is told, or watches the tier itself | a hop of latency, plus a component to scale, monitor and fail over | | Directory lookup | a service asked per lookup | two, unless the answer is kept | it is the source, so there is nothing to propagate | a lookup on the hot path, or a cache that reintroduces staleness | The hop counts are the easy part of the comparison and usually the least interesting. The column that decides real designs is the third one. ## How a change reaches the resolver Assignments move: a node is added, a node fails, partitions are moved between nodes while the tier keeps serving. Each shape finds out differently. - A **caller-side map** typically discovers a change reactively, by a node answering that the key lives elsewhere, and the caller is expected to refresh and remember. Some deployments also refresh on a schedule, or are pushed an update. Until the refresh lands, that caller is routing on an old assignment. - An **intervening proxy** concentrates the problem: one component refreshes, and every caller behind it is correct the moment the proxy is. The flip side is that the caller has no idea whether the proxy is current, and none of its own signals will tell it. - A **directory** consulted per lookup is never behind, because it is the source. The moment you keep its answers to avoid the extra round trip, you are holding a caller-side map whose refresh policy happens to be a deadline — with all the same consequences, bounded by that deadline rather than by how fast a change is pushed. ## What is the same in all three The underlying requirement does not change with the shape: at any moment, one key has one owner, and whoever resolves must agree with whoever else resolves. What the shape decides is **where a wrong answer can be produced and who is able to notice**. Staleness is not eliminated by any of the three; it is relocated — into every caller, into the proxy tier, or into a cache in front of the directory. ## Choosing between them - If the latency budget genuinely cannot afford a second network hop, the caller-side map is the only shape that avoids it, and you accept fleet-wide responsibility for the map. - If callers are written in several languages, or are maintained by teams that will not upgrade a library on your schedule, a proxy converts many correctness surfaces into one. - If the topology changes often and callers must never act on an old assignment, an authoritative lookup buys that at the price of a round trip, and you should be honest that caching it gives the price back. - If the store you are running offers only one of the three, the choice is already made, and the design question becomes how your application survives that shape's characteristic failure. That last point deserves emphasis, because it is where candidates most often assume one product's arrangement is the model. Some stores in this class expect the caller to resolve and will tell it when it is wrong. Others front the nodes with a proxy so the caller never resolves anything. Others again are just independent nodes with the whole convention living in the client library. An answer that names the three and says what each pays travels across all of them; an answer that describes redirects as *the* mechanism does not.
- Does putting a proxy in front remove the risk of a stale partition map?No, it relocates it. The proxy holds the map and can be stale on the caller's behalf, with the caller unable to tell. What you buy is one place to fix and one refresh path instead of one per client; what you pay is a hop, a tier to operate, and the loss of the caller's own signal that something moved.
- When does a cached directory answer stop being a directory lookup?As soon as the answer is kept for reuse. The caller is then holding a copy of the map with a deadline attached, and every stale-map consequence returns, bounded by that deadline rather than by how quickly a change is pushed. The honest description is a caller-side map whose refresh policy is an expiry.
- Why is the number of hops usually the least interesting column?Because the hop is a fixed, predictable cost you can measure once, while the propagation question decides how long a wrong answer can persist and whether anyone can see it. A design that saves a hop and silently routes on a week-old assignment has made the worse trade.
saying these in an interview costs you the question
- Assumes the caller is always told the key lives elsewhere
- Thinks a proxy removes staleness rather than relocating it
- Counts only latency and ignores who must learn of a change
- Treats a per-lookup directory as free because it is authoritative
- Believes every store in this class offers all three shapes