An application-maintained index has drifted from the entries it points at over months. How do you detect the divergence and repair it while callers keep reading?
answer
- drift is the normal state
- never verify the tier against itself
- both directions, or half a check
- rebuild beside, then move readers
- dual-write during the build
basics
~20 sCompare the index against the authoritative source, not the tier, in both directions: index entries pointing at entries that are gone, and records with no index entry. Repair at read time, and rebuild into fresh keys readers move to once complete.
solid answer
~50 sDrift is expected, not exceptional, so the design question is what you compare against. Comparing the index to the tier is circular — the tier is the thing you suspect — so drive the pass from the durable system of record the accounts also live in: for each record, compute what the index should say and check it, which finds records the index never learned about. Then walk the index's own entries to find pointers to entries that are gone or no longer carry the looked-up value. Repair has three gears: verify-and-fix at read time, which costs one comparison; a scheduled reconciliation reporting a divergence count as a health signal; and a full rebuild into a separate set of index keys, with readers moved only once it completes, so nobody reads a half-built index. Rate-limit the rebuild — it is write traffic on a serving tier.
go deeper
Recall that nothing repairs the index for you, so a job somewhere has to compare it with the truth and fix what it finds.
Explain the two directions of the check and what each one uniquely finds — an unindexed record against a dangling index entry.
Show the live rebuild: build beside, dual-write during the build, move readers at the end, throttle throughout, and report the divergence count as a signal.
Make the repair pass a condition of approving the index at all, with a named owner and a divergence budget, and count its memory and write cost in the tier's capacity plan.
## Why drift is the normal state An **application-maintained index** is a second entry the application writes so it can look something up by a value. Nothing in the store enforces agreement between it and the entries it points at, so every one of these leaks a little divergence, and they accumulate: - a write path in a service that does not know the index exists; - a delete path that removed the entry and forgot the index; - a process killed between the two writes; - a write refused because the tier was at its memory ceiling; - a **lifetime attached to one of the two entries** and not the other, so one disappears first; - a bug fixed six months ago whose damage is still sitting in the keyspace. So the honest posture is that the index is wrong by some unknown amount right now, and the engineering question is how you find out and how you fix it without an outage. ## Do not compare the tier against itself The first instinct — rebuild the index from what the tier holds — is circular. The tier is the thing under suspicion, and it is also volatile: entries may have been removed under memory pressure or lost across a restart, so an index rebuilt from it inherits every gap. It also costs the tier a great deal of read work to produce a source of truth it does not actually have. The repair pass is driven from the **authoritative source** the data also lives in — the durable engine of record. That gives both directions of the check: | Direction | Driven from | Finds | |---|---|---| | Forward | each record in the source | records the index never learned about, or that it points at the wrong key | | Reverse | each entry of the index | pointers to entries that are gone, or that no longer carry the looked-up value | Only the forward direction can find an **unindexed record**, because an entry nobody indexed is one nothing in the tier will ever lead you to. Only the reverse direction can find a **dangling index entry** for a record the source deleted long ago. A pass that runs one direction and calls the index verified has checked half of it. ## Three gears of repair 1. **Read-time verification.** After resolving a value to a key and reading the entry, confirm the entry still carries the value that was looked up. If it does not, treat the lookup as a miss and remove the offending index entry. One comparison per read, and it keeps the paths people actually use from compounding — but it never sees an index entry nobody reads, and never discovers a record that was never indexed. 2. **Scheduled reconciliation.** A periodic job running both directions, writing the fixes and — the part teams skip — reporting the divergence count. That number is the health signal: a stable small number is the background rate, and a step change points at a write path that stopped maintaining the index. 3. **Full rebuild.** Reserved for the cases where divergence is large or its cause is unknown. ## Rebuilding while callers read A rebuild must not be done in place, because a half-rebuilt index is worse than a drifted one: a reader that arrives mid-rebuild sees an index that is missing most of its content and concludes that nothing exists. The shape that works is: 1. Build the replacement under a **separate set of index keys**, driven from the authoritative source, while the existing index keeps serving reads. 2. Have writers maintain **both** index key sets for the duration, so the new one does not fall behind the moment it is built. 3. When the build completes and the divergence count on the new set is acceptable, move readers to it. 4. Stop the dual write and remove the old index entries. The dual-write step is what makes it safe and is what gets forgotten: a rebuild that snapshots the source while writes continue produces an index that is already stale when it is switched to. ## The costs to say out loud A rebuild is **write traffic against a tier that is serving**, and on a store that executes one operation at a time, an unthrottled rebuild competes directly with every caller; where the store executes operations on several worker threads it narrows rather than stops the tier, but the network and the memory ceiling still notice. Rate-limit it, and remember that the replacement index occupies memory at the same time as the old one, so the peak footprint is roughly double — which on a tier near its ceiling is exactly how a repair job becomes an incident.
- Why can read-time verification not replace the scheduled pass?Because it only inspects what someone looked up. An index entry nobody reads stays wrong indefinitely, and a record that was never indexed is invisible to it — the lookup that would reveal it is exactly the one that returns nothing. Read-time verification bounds the damage on hot paths; it never measures the divergence.
- What would you emit from the reconciliation job besides the fixes?The divergence count, split by direction, as an ongoing signal rather than a log line. A stable background number tells you the design's normal rate; a jump in unindexed records points at a write path that stopped maintaining the index, usually a new service or a new code path, and that is the finding worth an alert.
- Why is a rebuild in place worse than leaving the index drifted?Because during the rebuild the index is mostly empty, and an empty index entry is indistinguishable from 'no such record'. Callers doing uniqueness checks create duplicates, and callers doing lookups report data loss. A drifted index is wrong about a few things; a half-built one is wrong about nearly everything, briefly.
saying these in an interview costs you the question
- Rebuilds the index from the tier, which is the thing under suspicion
- Runs one direction only and calls the index verified
- Rebuilds in place while readers are using the index
- Runs the rebuild unthrottled against a serving tier
- Treats repair as a launch task rather than a permanent job
- Forgets that both index copies occupy memory during a rebuild