skip to content

A store answers only by key, but a service must find an account by email address. What has to exist, and who maintains it?

level: juniorimportance: must knowfreq 74%

answer

  1. the store answers one question only
  2. any other lookup is a second entry
  3. address to key, then key to entry
  4. you write it, you repair it

basics

~20 s

A second entry: an application-maintained index keyed by the email address, holding the account's key. The application writes, updates and repairs it, and nothing in the store's base data model builds such a lookup or notices when it is wrong.

solid answer

~40 s

The store maps a key to a value and answers nothing else, so a lookup by email address is not a query the tier can run — it is a second entry the application creates and keeps in step. The usual shape is a forward entry, `account:1042`, holding the account, and an index entry, `index:account-by-email:[email protected]`, holding the string `account:1042`. A read becomes two round trips: resolve the address to a key, then read the entry. From that moment every write path that touches the address must write the index too, every delete path must remove it, and the divergence between them is yours to detect and repair. Some deployments run an optional indexing add-on that maintains such lookups server-side; without one, correctness is entirely the application's.

go deeper

for a junior

Remember the founding fact: the store answers by key, so any other lookup is a second entry your own code wrote. Name the two round trips — address to key, key to entry.

for a middle

Explain what the index entry holds for a one-to-one lookup against a one-to-many one, and say which of those shapes exists only where the server understands the value rather than handing back opaque bytes.

for a senior

Show the obligations you just created: every write path, every delete path, and a repair job that never ends. Say how a service that does not know the index exists will break it.

for a principal

Frame it as a commitment rather than a feature. Ask whether this lookup belongs in an unschematised keyspace at all, or whether the engine that already maintains its own indexes should answer it.

## The only question the store answers An in-memory store of this class maps a **key** to a value. Every read names the key it wants, the store locates that one entry and hands the value back. There is no schema, no declared column list and no type the store understands well enough to index, so there is nothing for it to build a lookup out of on your behalf. An **entry** is one key plus the value it addresses, and the key string is the entire data model. That single fact is where every other access path starts. A caller holding an email address holds something the store has never been told about. It cannot answer "which entry has this address" for the same reason it cannot answer "how many accounts are suspended": the value is not a queryable surface, and in a large part of this class the server treats the value as **opaque bytes it only hands back**, so it could not look inside even if it wanted to. ## The lookup is a second entry you write The standard design is two entries per fact: - the **forward entry** — `account:1042` holding the account itself; - the **application-maintained index** — a second entry the application writes so it can look something up by a value, for example `index:account-by-email:[email protected]` holding the string `account:1042`. A read by address becomes two round trips: read the index entry to get the key, then read the entry. A read by key stays one. That extra hop is the cheapest part of the deal. ## What the index entry holds depends on the lookup | Lookup | Index entry | What it holds | |---|---|---| | One address, one account | one entry per address | the account's key as a plain value | | One status, many accounts | one entry per status value | an unordered member collection of account keys | | "Most recent twenty" | one entry for the access path | a score-ordered collection of account keys | Which of these you can actually build varies by store. Where the server understands the value as a collection and can change part of it, adding or removing one member is a single small operation. Where the server treats values as opaque bytes, the many-member forms have to be encoded into one value that you read whole, change and write back — which is slower, bounded by how large one entry may grow, and races with every other writer. On such a store the common design is one small index entry per indexed value and no many-member index at all. ## What you have just taken on 1. **Every write path.** Creating an account writes two entries. Changing the address writes the new index entry and removes the old one. A service that writes accounts without knowing the index exists silently breaks it. 2. **Every delete path.** Deletes are the paths that get forgotten, and a forgotten delete leaves an index entry pointing at a key that no longer resolves. 3. **A window.** The two writes are separate operations, so between them the index and the entry disagree, and a concurrent reader can land inside that window. 4. **Repair.** Because nothing enforces agreement, divergence accumulates, and detecting and fixing it is a permanent job rather than a launch task. 5. **Memory.** The index is real entries with real per-entry cost, and it grows with the number of distinct indexed values, not with the number of accounts. ## Where something else maintains it Nothing in the base data model of this class maintains such a lookup. Two things change that. Some deployments add an **optional indexing add-on** that indexes values server-side; where one is in use, the store maintains the index and correctness moves from your code to that component, along with its own operational surface. And the durable engine the data usually also lives in maintains its own indexes as part of committing a row — which is exactly why the honest answer to some of these lookups is that they should be asked there rather than here. ## What an interviewer is listening for The weak answer is "add an index on email", spoken as if the tier had indexes. The strong one names the second entry, the two round trips, the write and delete paths that now have an obligation, and the fact that the store will never tell you the index is wrong. A candidate who volunteers "and who repairs it when it drifts" before being asked has run one of these in production.

  • Why does the index entry usually hold the key rather than a copy of the value?
    Because a copy is a second thing to keep in step. Holding the key means the index can only ever be wrong about which entry to visit, and the entry itself is still the one source of the value. A copy adds a second way to be stale and doubles what a repair pass has to compare — it buys one saved round trip for a permanent correctness cost.
  • A read by address finds no index entry. Does that prove no such account exists?
    No. It proves the index has no entry for that address, which is also what you see when the index write was lost, when a lifetime attached to the index entry ran out before the account's, or when a writer that did not know about the index created the account. Treating absence in an application-maintained index as absence in the data is the mistake that makes duplicate accounts.
  • What does the design cost when the same account must be findable by three different values?
    Three index entries, three writes on create, three removals on delete, and three ways to drift. Each additional access path multiplies the write amplification and the repair surface, which is why the count of access paths — not the count of accounts — is the number to challenge in a design review.

A closed-stack library shelves books by call number and nothing else. To find everything by one author you use a card catalogue — cards a librarian typed and files by hand. The shelves do not update the cards; when a book is reshelved or discarded and nobody pulls the card, the catalogue confidently sends the next reader to an empty slot, and the fix is to re-derive the cards from the acquisitions ledger rather than from the shelves you already doubt.

saying these in an interview costs you the question

  • Talks about adding an index on the field, as if the tier had indexes
  • Assumes the store keeps the index in step with the entry
  • Forgets the delete and address-change paths entirely
  • Treats a missing index entry as proof the account does not exist
  • Copies the whole value into the index to save a round trip