skip to content

A team proposes a third application-maintained index on a shared tier. What does owning one commit them to, and when should the lookup live elsewhere?

level: principalimportance: should knowfreq 40%

answer

  1. an obligation, not a feature
  2. every writer that will ever exist
  3. ask what a wrong answer costs
  4. uniqueness checks do not belong here
  5. name the owner of the repair

basics

~20 s

It commits every present and future writer of that data to maintaining the index, plus a permanent repair pass, extra memory, and a rebuild after anything that empties the tier. A lookup that must be right belongs to the durable engine.

solid answer

~50 s

An application-maintained index is not a feature you add, it is an obligation you take on forever, and it is taken on by everyone who writes the data — not only the team proposing it. Approving one commits the organisation to every write path maintaining it, in every service, present and future; a repair pass with a named owner and a divergence signal; the memory it occupies; and a rebuild after anything that empties the tier. The third index is usually where write amplification and repair surface stop being worth it. Decline it when the lookup is a correctness question — a uniqueness check, an entitlement decision — and let the durable engine that maintains its own indexes answer it. Where a deployment runs an optional indexing add-on, correctness moves to that component along with its operational surface.

go deeper

for a junior

Recall that an index here is code someone has to keep writing — approving one means someone maintains it for as long as the data exists.

for a middle

Count the obligations concretely: operations per logical write with three indexes, the memory the index entries occupy, and the delete paths that must all be updated.

for a senior

Argue from what a wrong answer costs. Show why a uniqueness or entitlement check belongs to the engine that can refuse the write, and what the recovery procedure must say.

for a principal

Set a rule rather than a verdict: declared indexes, one owning team, a divergence signal, correctness-critical lookups out of the tier, and access-path count as a review trigger.

## What is actually being requested An **application-maintained index** is a second entry the application writes so it can look something up by a value. It looks like a small addition — one more key template, a couple of extra writes — and it is approved that way. What is really being requested is a permanent, organisation-wide obligation, and the third one is usually where that becomes visible. ## The commitments, in the order they get forgotten 1. **Every writer, forever.** The index is only as correct as the least careful write path. That includes services that do not exist yet, a data-fix script someone runs at two in the morning, and a migration job. None of them is stopped by the store, which has no idea the index exists. 2. **Every delete path.** Deletes are the paths with the least test coverage and the most branches, and a forgotten delete leaves a pointer to something gone. 3. **A repair pass with an owner.** Not a launch task — a scheduled job, driven from the authoritative source, reporting a divergence count that somebody looks at. An index without a named owner for its repair is an index that is quietly wrong. 4. **Memory.** Index entries are real entries with real per-entry cost, and they grow with the number of distinct indexed values, which is not the number of records. On a shared tier that memory is taken from every other tenant. 5. **A rebuild after the tier is empty.** The tier is not the only copy of anything, so any event that clears it — a restart on a store keeping nothing across one, a restore, a resize — leaves the index to be rebuilt before the access path works again. That rebuild has to be in the recovery procedure, or the first sign of it is a user report. 6. **Write amplification.** With three indexes, one logical change is one entry write plus up to six index operations, because a changed value means a removal from the old index entry and an add to the new one. That is the number to put in front of the reviewer. ## Where the lookup should live instead | Signal | Better home | |---|---| | A wrong answer is expensive — uniqueness, entitlement, billing | the durable engine, which maintains its index as part of committing the write | | The access path combines several values or needs ranking with filtering | a system whose job is answering questions about values | | The lookup is read rarely but must always be right | the source of record, with the tier holding nothing about it | | The lookup is read constantly, tolerates being briefly wrong, and has one value | an application-maintained index is a reasonable answer | The distinguishing question is not "is it fast enough there" — it usually is — but **what a wrong answer costs**. A uniqueness check through an application-maintained index is the classic mistake: the index is missing an entry for a few milliseconds or a few months, the check passes, and a duplicate account is created that no constraint prevented. The durable engine can refuse that write outright, and nothing in an unschematised keyspace can. ## The third option worth naming Some deployments run an **optional indexing add-on** that indexes values server-side. Where one is in use the trade changes shape rather than disappearing: the store maintains the lookup, so the divergence window and the repair pass move out of your code, and in exchange you take on that component's memory behaviour, its own operating surface and a dependency your key convention now assumes. It is a real answer to this question; it is not a free one, and the details of any particular one belong to the product that ships it. ## What a principal-level answer sounds like It is not "no". It is a rule the organisation can apply without you in the room: - every application-maintained index is **declared** next to the entry it shadows, so the next reader of the keyspace can see it exists; - it has **one owning team**, which owns the repair pass and the divergence signal; - correctness-critical lookups do not live in the tier, and that is written down rather than argued case by case; - the count of access paths on one entity is a **review trigger**, because each one multiplies writes and repair surface; - the recovery procedure names which indexes must be rebuilt, and from what. That is the difference between owning a keyspace and hosting one: the rule survives the people who wrote it, and the engineer on call in two years can still tell what the keyspace is supposed to contain.

  • Why is a uniqueness check through an application-maintained index unsafe even with a grouped write?
    Because the check and the write are still two decisions separated by a window, and because the index can be incomplete for reasons unrelated to this request — a writer that never maintained it, a removal under memory pressure, a lost write. A grouped write narrows the race; it does not make the index authoritative. Only the engine that can refuse the write can enforce uniqueness.
  • The team argues the index costs almost nothing because the entries are tiny. What do you put against that?
    The write amplification and the repair surface, not the bytes. One change becomes several operations across every service that writes, and the index adds a permanent job, an owner, a signal and a line in the recovery procedure. The memory is usually the smallest of the costs, and quoting it is how the real ones get skipped.
  • What has to happen to these indexes after the tier is restored or restarted empty?
    They have to be rebuilt from the authoritative source before the access paths are trusted, and until then a lookup that finds nothing must not be read as 'no such record'. If that step is not in the recovery procedure, the tier comes back looking healthy while every lookup by value silently reports absence.

saying these in an interview costs you the question

  • Argues the cost is the bytes rather than the write amplification
  • Approves an index with no named owner for its repair
  • Enforces uniqueness through a lookup the store does not maintain
  • Assumes only the requesting team will ever write the data
  • Leaves index rebuilds out of the recovery procedure
  • Treats an indexing add-on as removing the trade rather than changing it