Where should the partition map live for tiers reached by services in several languages, and what should you refuse to depend on?
answer
- who pays when an assignment changes
- hop count is the easy column
- one surface or a fleet of them
- refuse to depend on an error
basics
~20 sDecide by who bears the cost of a change, not by hop count. Caller-side maps spread correctness across every client version deployed; a proxy concentrates it into one component you run; a directory is authoritative but priced per lookup.
solid answer
~50 sThe organisational question is not which shape is fastest but **where a wrong map can exist and who can see it**. Standardising on caller-side maps buys the lowest latency and makes the map a property of every client library version deployed anywhere — a correctness surface you cannot inventory. A proxy converts many surfaces into one, at the cost of a hop and a tier with its own capacity, monitoring and failure behaviour, and it hides from callers whether they were misrouted. A directory is the authority, and caching its answers quietly returns you to caller-side maps with a deadline. The dependency worth refusing outright is the assumption that a misroute surfaces as an error: stores in this class differ on whether nodes report foreign keys at all, so every design should survive a stale map producing an ordinary absence.
go deeper
Understand that the choice of where the map lives is made once for a whole organisation, and that it affects every service that talks to the tier.
Be able to argue the comparison on propagation rather than latency: how many places hold the map, and how each of them finds out that an assignment changed.
Show the operational follow-through: a versioned assignment, telemetry that reveals what a caller believed, and alerts on miss ratio and misroute replies rather than on the map itself.
Own the refusal list. Name the assumptions that vary across stores in this class, most importantly that a misroute announces itself, and state the migration cost the routing shape locks in.
## The decision is about who pays for a change Assignments move — nodes are added, nodes fail, partitions are relocated while the tier serves. Every routing shape works when nothing is changing. The standing decision is about the moment something does, and specifically about **how many places can hold a wrong answer, and which of them can tell you they are holding one**. That reframing matters because the obvious comparison — one network hop versus two — is measurable, bounded and usually not what hurts. A hop is a fixed tax you can budget. A fleet routing on a month-old assignment is an unbounded correctness problem you will discover from user reports. ## The three bills | Shape | What the organisation takes on | What it gives up | |---|---|---| | Caller-side map in every service | the lowest achievable latency and no component in the path | the map's correctness is now a property of every client version in the fleet, in every language, on every team's release schedule | | Intervening proxy | one place to hold, refresh and fix the map; callers become trivially portable | a hop, plus a tier to size, monitor and fail over, and callers that cannot tell whether they were misrouted | | Directory consulted per lookup | an authority that is never behind | a round trip on the hot path, or a cache that puts you back on caller-side maps with a deadline you chose | The first row is where most organisations end up by default, usually without deciding. It is a defensible choice, but only with the obligations that come with it: one sanctioned client per language, a way to see which version each service is running, and the ability to roll a map change independently of application releases. ## What you refuse to depend on The most useful principal-level output here is a short list of guarantees the organisation will not build on, chosen because they **vary across the stores in this class** and would therefore turn a store swap into a rewrite: - **Do not depend on a misroute surfacing as an error.** Some stores track partition ownership per node and answer that a key lives elsewhere. Others accept whatever arrives, so the same mistake returns an ordinary absence. Designs must survive the second case. - **Do not depend on absence meaning never-written**, for any state whose loss is expensive. That is the same rule stated from the application's side. - **Do not depend on the caller receiving any topology signal at all.** Behind a proxy it receives none, so no design should require one to stay correct. - **Do not depend on all three shapes being available.** Some stores offer exactly one, which turns a standing preference into a procurement constraint. What you *can* depend on across the whole class is narrow and worth naming: one key has one owner at any moment, the resolution rule is deterministic against the current assignment, and growth by adding nodes exists only because that rule exists. ## How the choice reaches back into procurement If every service is written to address a single endpoint that a proxy owns, then adopting a store that expects callers to resolve is not a configuration change — it is a change to every service. The reverse is equally true: a fleet that has standardised on a resolving client cannot trivially move behind a proxy that hides the topology, because its retry and refresh logic assumes signals it will stop receiving. The routing shape is one of the few technical choices in this area with a direct migration cost attached, so it belongs in the same conversation as which store you buy. ## Making the map an artifact Whichever shape wins, the practical follow-through is to stop treating the map as configuration that lives wherever it happens to live: 1. **One owner** for the assignment, and one path by which a change is published. 2. **A visible version** in the telemetry of whatever resolves keys, so answering *what did that caller think the assignment was* takes a query rather than an investigation. 3. **Alerting on the two symptoms**, not on the map itself: miss ratio against request volume, and the rate of replies saying a key lives elsewhere where the store produces them. 4. **A rollback path** independent of application deployment, because a bad assignment published to a fleet is otherwise repaired at the speed of your slowest release train. ## The blast radius question Finally, size the damage honestly. A wrong map does not degrade a fraction of traffic at random; it affects exactly the keys in the partitions that moved, across every caller that has not refreshed. If the tier holds state that exists nowhere else, that population is the one that experiences the incident — and whether a given workload may keep state with no other home is a separate standing decision that this one should not quietly pre-empt.
- What does the routing shape constrain about which stores you can adopt later?More than teams expect. If every service addresses a single endpoint a proxy owns, a store that expects callers to resolve keys becomes a change to every service, not a swap. A fleet standardised on resolving clients has the mirror problem: its refresh and retry logic assumes signals a proxy will never send.
- How do you cap the blast radius when a published assignment is wrong?Treat the map as a deployed artifact: one owner, a version visible in the telemetry of whatever resolves keys, and a rollback path independent of application releases. Alert on miss ratio against traffic and on the rate of replies saying a key lives elsewhere, so a bad map appears as a signal rather than as user reports.
- Which guarantees about key assignment hold across this whole class of stores?Only the narrow ones: one key has one owner at any moment, resolution is deterministic against the current assignment, and adding nodes as a growth path exists because that rule exists. Everything about how a wrong answer is reported, who holds the map and how a change propagates varies by product.
saying these in an interview costs you the question
- Chooses the shape purely on the extra network hop
- Assumes every store in this class offers all three shapes
- Treats the map as configuration rather than a versioned artifact
- Believes a misroute will always announce itself as an error
- Ignores that the shape decides the cost of swapping stores later