skip to content

Several teams share one volatile tier that will be split into partitions next quarter; what standing position do you set for multi-key designs?

level: principalimportance: should knowfreq 38%

answer

  1. one rule: no single address space
  2. inventory every multi-key call
  3. classify, do not default to co-location
  4. rework before the split, not on it
  5. refuse to depend on cross-partition units

basics

~20 s

Rule: no design may assume one address space. Every call naming more than one key is classified - co-located under a capped group, split, folded, or explicitly accepted as non-indivisible - and reworked before the split.

solid answer

~50 s

The position is one sentence: no design may assume that all keys share one address space. Everything else is how you make that checkable. Inventory every call that names more than one key, plus every group of steps submitted as a unit and every server-side unit, and classify each: co-locate under a capped group, split and compose in the caller, fold into one entry, or accept it is no longer indivisible and say what a reader does. Do the rework before the split, on the single node, where both the old and the new shape run — that turns split day into a routing change rather than a code change. Then name the guarantees the organisation refuses to depend on: cross-partition indivisibility of any kind, and any intervening layer silently composing multi-key calls, because both vary across the products you might buy.

go deeper

for a junior

Take away the sequencing point: the reworked shapes run correctly on a single node too, so this work is done before the split rather than during it.

for a middle

Be able to produce the inventory — calls naming several keys, groups submitted as a unit, server-side units — and to classify each into co-locate, split, fold, or accept.

for a senior

Argue the sequencing and the gating signal: rework while both shapes work, count what remains, and make that count the gate on split day rather than discovering the remainder in production.

for a principal

Set the standing rule and the guarantees you refuse to depend on, then govern co-location group size across teams, since one team's group sizes a node the whole organisation pays for and fails with.

## The position, in one sentence **No design may assume that every key it touches shares one address space.** That is the rule; the rest of this is what makes it enforceable in an organisation where several teams write to the same tier and none of them owns the split. The reason to state it as a standing position rather than a migration task is that the tier will be split again, moved to a hosted equivalent, or put behind a different routing shape at some point, and each of those changes what a multi-key call does. A codebase that stopped assuming one address space survives all three; a codebase that merely got through one split does not. ## Finding the designs that assume it The inventory is smaller than teams expect and harder to grep for than they hope. What you are looking for: - Any call that names more than one key. - Any group of steps submitted to the store as one unit. - Any unit the server executes on your behalf over several keys. - Any code that derives one key name from another and then reads both, which is the same assumption wearing a disguise. - Any place where a comment says "these are always written together" and nothing enforces it. The practical lever is the client boundary: if every team reaches the store through one thin internal interface, the calls are enumerable and the rework is mechanical. If every team uses the driver directly, the first task is the interface, not the split. That is also the point at which the organisation gets the ability to change routing shape later without touching application code. ## Classifying each one | Class | Rework | The cost you are accepting | |---|---|---| | Must run as one unit, small group | Co-locate under a capped group | The group can only be placed whole, and one node now holds all of it | | Values independent | One call per key, composed in the caller | More calls, and the group is no longer seen at one instant | | Values are really one thing | Fold into one entry | Entry size and that group's traffic on one node | | Indivisibility not actually needed | Accept it, and state what a mixed view means | The guarantee, deliberately, in writing | The fourth row is the one to make respectable. Teams will reach for co-location by default because it is the change that requires no thought, and an organisation that lets them ends up with a tier where placement is decided by whoever shipped first. ## Sequencing: rework before the split Every one of these reworks is correct on a single node as well. Four single-key calls work fine on one machine; one entry holding a group works fine on one machine; a co-location marker in a key name is inert on one machine. That is the whole sequencing argument: **do the reworks while both shapes still run**, so the split becomes a routing change with no application release attached to it. A team that leaves the rework until split day has converted an incremental change into a big-bang cutover, and the failures arrive on whichever code path is least exercised. It also gives you a real progress signal. The number of remaining multi-key calls is countable, it goes down, and it gates the split. ## The guarantees the organisation refuses to depend on Write these down, because they are the ones that differ between products and the tier is among the components most likely to be replaced by a hosted equivalent: 1. **Cross-partition indivisibility**, in any form. Nothing this tier offers spans partitions as one unit. 2. **An intervening layer composing multi-key calls.** Where a proxy fans a call out and merges the results, the call works and was never indivisible — and it stops working entirely if the routing shape changes to one that refuses instead. Depending on it means depending on an accident of deployment. 3. **A uniform answer about what happens on a violation.** One arrangement refuses, another composes, another never offered the operation. Code that handles only the refusal is code written against one arrangement. ## Governing co-location A co-location marker is cheap to introduce and expensive to remove: undoing one means renaming keys in a running system. So cap the group size, require that a new marker names the operation that justifies it, and give every group an owner who is accountable for its size. Because the tier is shared, one team's over-broad group sizes a node that everybody is paying for and failing with, which makes this a cross-team standard rather than a team-local choice. ## What skipping this costs The cross-partition refusal arrives in production, on the least-tested path, during the change window when everything else is also new. Or worse, it does not arrive at all: an intervening layer composes the call, the group is quietly read torn, and the first symptom is a customer-visible inconsistency nobody can reproduce. Either way, the migration you postponed is now a key rename under load — and teams, having been burned, start treating the tier as if it were one machine again.

  • What do you say to a team that proposes co-locating everything so nothing has to be reworked?
    That it moves the problem from the code to the placement, where it is harder to undo. Groups can only be placed whole, so co-locating broadly reproduces the single node they are trying to keep, complete with its ceiling — and unwinding it later means renaming keys in a running system rather than changing a call.
  • How do you know the rework is finished, given the split has not happened yet?
    Count what is left. With access to the store behind one internal interface, every call naming more than one key is enumerable, and the count gating the split goes to zero or to a reviewed list of capped co-located groups. Without that interface, you cannot know, which is why the interface comes first.
  • Does any of this change if the tier is put behind a proxy rather than split directly?
    The rule does not change; the symptoms do. Behind a proxy some multi-key calls keep working by being fanned out and composed, which hides the problem instead of surfacing it. That is more dangerous than a refusal, because a torn group is discovered by a customer rather than by an error.

saying these in an interview costs you the question

  • Waits for the split and treats the refusals as the inventory
  • Defaults every case to co-location because it needs no code change
  • Assumes the routing layer will hide the difference from applications
  • Believes the reworks cannot be shipped until the tier is actually split
  • Treats co-location group size as each team's own business on a shared tier