skip to content

Keyspace Design

How a value is addressed when nothing enforces a schema: what a key encodes, how many keys a design makes, how large one entry grows, and how you find what you did not name.

on this pageshow

questions

21

Two teams share one flat keyspace and one writes entries under the key user:1042 — what should that key carry instead?

level: juniorimportance: must knowfreq 70%

answer

  1. nothing here enforces a name
  2. the key is the model
  3. whose entry is this?
  4. owner, entity, identifier, format version

basics

~20 s

On a shared flat keyspace the key string is the entire addressing model, so it must say whose it is: an owning prefix, then the entity, then the identifier, and a format version where the value's shape may change.

solid answer

~50 s

`user:1042` names an entity and an identifier and nothing else, so it is a name any other team on the tier may also choose — and the store has no schema to object. Write the key from a template of segments joined by one separator fixed across the organisation: an **owning prefix** naming the team or service, the **entity**, the **identifier**, and a **format version** where the value's serialized shape may change under a rolling deployment. The prefix is the segment doing the work: it makes an entry attributable to whoever wrote it, and it means a second team arriving on the tier does not silently land on top of yours. The separator is a convention rather than something the store enforces, and stores differ in which bytes a key may hold, so fix it against the store you actually run.

go deeper

for a junior

Recall the four segments and why they exist: owner, entity, identifier, and a format version where the value's shape can change. Be able to say out loud that nothing in the store stops a second team writing the same name.

for a middle

Explain the failure mechanics, not the convention: one name addresses one entry, a write replaces the value, and both callers still succeed. Say what the separator does and does not guarantee, and what happens when an identifier contains it.

for a senior

Show that you have enforced a convention rather than written one — keys built through a single function, a format version segment planned before the first shape change, and a clear statement that a prefix buys attribution rather than isolation.

for a principal

Frame it as an organisational contract: what every key on a shared tier must encode so an entry is attributable years later, and how you get a convention adopted when nothing in the store will ever reject a key that ignores it.

## The store has no schema, so the key is the addressing model An in-memory store of this class has no table, no column list and nothing that declares what may exist. The address string is the model: the only promise is that a value written under a key comes back under that same key. A relational engine refuses a second table with a name already taken; here nothing refuses anything. So `user:1042` is not a declaration that these entries belong to your service. It is a claim on a name, made by whoever wrote first and taken over by whoever writes next. That is why interviewers open here. An unprefixed key is cheap to type and expensive on a tier two teams share, and the cost arrives months later, in someone else's service. ## What a collision actually looks like Two services writing the same key do not get an error. One name addresses one entry, and a write replaces the value that was there. Both writes succeed, both reads succeed, and nothing fails at the moment the two designs meet: - if the two services store **different shapes**, the reader fails to parse a value it did not write — a confusing failure a long way from its cause; - if they store the **same shape with different meanings** — one counting attempts, the other counting successes — nothing fails at all, and the number is simply wrong; - if one service attaches a lifetime to the entry and the other does not, each is surprised by the other's behaviour without ever seeing the other's code. ## The segments a key on a shared tier should carry | Segment | The question it answers | What goes wrong without it | |---|---|---| | Owning prefix | Whose entry is this? | Two teams claim one name; nobody can attribute an entry found later | | Entity | What kind of thing is it? | `1042` alone is meaningless to the next reader | | Identifier | Which one? | Every instance of the entity fights for one name | | Field | Which part of the thing? | Only relevant where the design addresses parts separately rather than holding the whole thing under one key | | Format version | Which shape is the value in? | Two deployed versions of an application read each other's writes and cannot parse them | A worked template: `acme:billing:invoice:5561:v2`. Read left to right it says the organisation or team, the service or domain, the entity, the identifier, and the shape of the value. Someone who finds that key in a listing two years from now can answer who to ask. The field segment is conditional in a way worth stating: whether one entry per field is even an option depends on the store. Where the server treats values as **opaque bytes it only hands back**, every change is a full read-modify-write and a key per field is the only way to address a part. Where the server understands the value as **structure it can read or change in part**, the same design can be one entry addressed by field instead. The naming choice follows that, not the other way round. ## The separator is a choice, not a rule A colon is conventional and nothing more; the store does not parse it, does not index it and does not reserve it. Stores differ in which bytes a key may contain and how long a key may be, so fix the separator against the store you run rather than against habit. The one hard requirement is that the separator must not appear inside a segment value — an identifier that is an email address or a free-text name will eventually contain whatever character you picked, and then a reader splitting the key gets the wrong segments. Either choose a character that cannot occur, or encode the segment on the way in. The same discipline applies to prefixes that are prefixes of each other: `user` is a prefix of `username`, so any convention that matches on a leading string should match on the separator too. ## What a prefix does not buy you - It **reserves nothing**. Another team can still write your name; the prefix makes that unlikely and always attributable, not impossible. - It isolates **names only**. Where the store offers a named container above the key as an alternative to a prefix, that container isolates names only as well: the memory, the ceiling and the operator stay shared. - It says nothing about **where an entry lives** when the keyspace is split across nodes; what a segment does to placement is a separate subject. - Longer keys are not free, but how much those extra bytes cost is a sizing question rather than a naming one. ## Put the convention in code, not in a document A convention that lives only in a wiki page is followed until someone is in a hurry. Build keys through one small function per service — owner, entity, identifier in, key string out — so that the convention has exactly one place to be wrong, and so that changing it later is a change to one function rather than an archaeology exercise across every call site.

  • Where in the key does a format version belong, and what does it buy during a rolling deployment?
    Put it at the end, alongside the identifier, so the owner and entity segments stay stable. During a rolling deployment two application versions run at once; if the value's serialized shape changed, a version segment means each writes and reads its own shape rather than handing the other a value it cannot parse. Old-shape entries then age out or are removed deliberately once no reader wants them.
  • Your convention fixes a separator, but an identifier already contains that character — what breaks, and what do you do?
    Splitting the key gives the wrong segments: an identifier containing the separator reads as two segments, so tooling and any prefix match attribute the entry to the wrong owner or entity, and two different identifiers can produce the same key string. Either pick a separator that cannot occur in any identifier you accept, or encode the segment on the way in so the raw character never reaches the key.
  • Does adding an owning prefix stop another team writing your entries?
    No. Nothing in the store reserves a name; a prefix is a habit, not a lock. What it buys is that an accidental clash becomes unlikely, and that every entry found later is attributable to an owner who can be asked. Keeping two teams genuinely apart by credentials, quotas or separate deployments is a different lever entirely.

A shared filing room where anyone may create a folder with any label. The room does not stop two departments labelling a folder invoices, and nothing announces the clash — the second department's folder simply stands where the first one did. The only defence is the label itself: department, document kind, number.

saying these in an interview costs you the question

  • Thinks the store rejects a key another team has already written.
  • Treats the separator as something the store parses or enforces.
  • Believes an owning prefix isolates memory or capacity, not just names.
  • Assumes every store offers a named container above the key.
  • Puts a mutable attribute such as a status into the identifier segment.
  • Says an unprefixed key is fine because their service is the only writer.
open as a page

A store answers only by key, but a service must find an account by email address. What has to exist, and who maintains it?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A second entry: an application-maintained index keyed by the email address, holding the account's key. The application writes, updates and repairs it, and nothing in the store's base data model builds such a lookup or notices when it is wrong.

open as a page

Why is asking a serving in-memory store for every key matching a prefix in one operation not a harmless query?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A whole-keyspace listing costs the distinct-key count, not the number of matches, and the reply is assembled whole before anything is sent. Where the store executes one operation at a time, every waiting caller pays it.

open as a page

A key scheme in an in-memory store mints one entry per user per day per field. How do you size it before shipping?

level: middleimportance: must knowfreq 62%

basics

~20 s

Multiply the factors out to a distinct-key count, then price each entry as key bytes plus value bytes and multiply again. Compare the result with the tier's budget, and identify which factor is unbounded before arguing about the bytes.

open as a page

On a store that hands values back as opaque bytes, what does one entry grown to hundreds of megabytes cost to read, write and remove?

level: middleimportance: must knowfreq 60%

basics

~20 s

Every operation on that entry is proportional to its whole size: a read ships all of it, any change rewrites all of it, and removing it is real work that someone has to pay for.

open as a page

Writing an entry and its application-maintained index takes two operations. What can a reader see between them, and what if the second fails?

level: middleimportance: must knowfreq 63%

basics

~20 s

Between the two operations the index and the entry disagree, so a reader sees either side of the window: an index pointing at an unchanged entry, or an entry no index knows of. If the second never lands, the disagreement is permanent.

open as a page

What does an incremental cursor walk over a live keyspace actually guarantee about the entries it returns, and what does it not?

level: middleimportance: must knowfreq 62%

basics

~20 s

A cursor walk returns every entry present throughout it at least once, may return one twice, and may or may not return entries that arrived or left mid-walk. It is not a snapshot, so no exact count rests on it.

open as a page

A service creates a new key in an in-memory store on every request and deletes none. What does the store do about that growth?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Nothing. Each write is individually valid and there is no schema against which a key count could be judged, so a key-minting design shows up only in the total entry count and the memory used, never in a refused write.

open as a page

An in-memory store holds a one-byte flag under a ninety-character key, forty million times over. What is that ratio costing?

level: middleimportance: should knowfreq 48%

basics

~20 s

The addressing has become the data: about 3.6 gigabytes of key text against 40 megabytes of payload. Keys are stored bytes, they travel on every call, and at high counts a key can far outweigh what it addresses.

open as a page

One store offers a named container above the key, another only a flat keyspace — what does each isolate for three teams?

level: middleimportance: should knowfreq 52%

basics

~20 s

A named container above the key isolates names only — containers share one memory budget, one process and one operator. A flat keyspace reaches the same separation with an owning prefix, which works on any store.

open as a page

An entry the server understands as a collection has grown for eight months — do you split it across keys or cap it on write?

level: middleimportance: should knowfreq 54%

basics

~20 s

Cap it when the tail is no longer wanted and the bound can be enforced on the write path. Split it across keys when every member still matters and only the single-entry packaging is the problem.

open as a page

In a store that answers only by key, which index shape suits a lookup by exact value, and which suits 'the twenty most recent'?

level: middleimportance: should knowfreq 52%

basics

~20 s

An exact-value lookup needs one index entry per value, holding an unordered member collection of matching keys. A ranked lookup needs a score-ordered collection. Both are application-maintained entries, and which the store can hold at all varies.

open as a page

A team wants to cut key bytes in an in-memory store by hashing each long key to a short fixed-width string. What does that trade?

level: seniorimportance: should knowfreq 44%

basics

~20 s

It trades legibility and a small chance of silent collision for a saving of exactly the bytes removed times the distinct-key count. Compute that product first: a collision here is one caller reading another's value, unreported.

open as a page

A live tier's key convention must gain an owning prefix while callers keep reading and writing — how do you run that change, and how do you know the old shape is dead?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Deploy readers that try the new key and fall back to the old, then flip writers, then let entries with a lifetime age out or backfill the rest. The old shape is dead when the instrumented fallback stays at zero.

open as a page

A store executes operations on several worker threads. How much of the oversized-entry problem does that actually remove?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Worker threads narrow the damage rather than remove it. Other callers keep being served, but the large operation still holds one worker for its whole duration, competes for the same memory and allocator, and stays slow for the caller that asked.

open as a page

An application-maintained index has drifted from the entries it points at over months. How do you detect the divergence and repair it while callers keep reading?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Compare the index against the authoritative source, not the tier, in both directions: index entries pointing at entries that are gone, and records with no index entry. Repair at read time, and rebuild into fresh keys readers move to once complete.

open as a page

A nightly reconciliation job walks the keyspace of a live store whose entries are split across several nodes; what must that job tolerate to be correct?

level: seniorimportance: should knowfreq 52%

basics

~20 s

One walk per node, assembled at different times: the result may repeat entries, may miss entries written while it ran, and describes no single instant. Every action must be idempotent, and a difference needs a second pass before anything destructive.

open as a page

A team proposes a third application-maintained index on a shared tier. What does owning one commit them to, and when should the lookup live elsewhere?

level: principalimportance: should knowfreq 40%

basics

~20 s

It commits every present and future writer of that data to maintaining the index, plus a permanent repair pass, extra memory, and a rebuild after anything that empties the tier. A lookup that must be right belongs to the durable engine.

open as a page

Your in-memory store offers no incremental traversal at all, yet you need an inventory of what it holds; how do you get one?

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

Not from the store: reconstruct the key set from outside - the system of record, the key template applied to identifiers you can iterate, or a list the write path records. A store without traversal answers only for keys you can already name.

open as a page

Across a shared tier serving many teams, what must a key convention encode so that any entry found on call is attributable?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Enough to answer, from the key string alone, who owns the entry, what it is, which value shape it holds and whether removing it is safe. The store will never reject a key that ignores the convention.

open as a page

You own a shared in-memory store used by many teams — what rule would you set for what one entry may hold?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

A stated bound in bytes and in members, derived from the latency and memory you are willing to spend on the worst operation rather than from whatever ceiling the store enforces, applied on the write path and re-examined whenever a new growth shape appears.

open as a page