When a product-search index is fully rebuilt into a shadow index, how do you keep the writes that arrive during the rebuild?
answer
- position first, snapshot second
- overlap is expected
- second consumer feeds the shadow
- switch only after catch-up and checks
basics
~20 sRecord the change stream's position before the bulk load starts, load a snapshot into the shadow index, then replay every change from that position with version-guarded writes until it catches up. Only then switch the alias; keep the old index for rollback.
solid answer
~50 sA full reindex builds a new **shadow index** beside the live one, while readers keep querying the live one through an **alias**. The rebuild takes hours and the catalog keeps changing, so I record the change stream's position **before** starting the snapshot read, bulk-load the snapshot into the shadow, and then run a second consumer that replays changes from the recorded position into the shadow. The snapshot and the replay overlap, so every write is **version-guarded** and duplicates are skipped. When the shadow's lag is within the normal SLO and verification passes (document counts, sampled query comparisons), I switch the alias atomically, make the shadow consumer the live one, and keep the old index for a rollback window. The change log must be retained for longer than the rebuild takes, or the replay has a gap.
code
pseudocode · 10 linesstart_pos = change_log.current_position()
shadow = create_index(new_schema)
for batch in read_snapshot(source_db): // snapshot read begins after start_pos
shadow.bulk_write_if_newer(batch)
shadow_consumer = consume(change_log, from = start_pos,
apply = shadow.write_if_newer)
wait_until(shadow_consumer.lag() < lag_slo)
if not verify(shadow, live):
abort() // alias untouched, readers unaffected
alias.point_to(shadow)go deeper
Recall that a rebuild goes into a new index while readers use a stable alias, and that the alias moves only when the new index is ready.
Explain the ordering: record the position, snapshot, replay, catch up, then switch, and why version guards make the overlap harmless.
Show the operational checks: log retention against rebuild time, throttled load, verification before the switch, and a rollback window with the old consumer still running.
Decide when a reindex is worth doing at all, and how to make it a routine, rehearsed operation rather than a rare, risky event for the catalog team.
## Why a full reindex happens Most index changes are incremental, but some force every document to be rewritten: - a field changes type, for example a text field that must become a numeric filter; - text analysis changes, so existing documents were prepared the old way; - the index is split into a different number of shards; - a new field must be populated for every product; - the index drifted or was damaged and must be rebuilt from the source of truth. For an e-commerce catalog the rebuild may take hours, and shoppers must keep searching throughout. ## The shadow index and the alias A **shadow index** is a new, empty index built with the new schema next to the **live** one. Readers never query an index by its concrete name; they query an **alias**, a stable name that points at exactly one index. Switching the alias from the live index to the shadow is a single metadata change, so queries move over at once and can be moved back just as quickly. The hard part is not the switch, but ensuring the shadow is complete and current at the moment of switching. ## The catch-up sequence 1. **Record the change stream's current position** (the offset in the log that feeds the indexer). 2. **Create the shadow index** with the new schema and write-friendly settings. 3. **Bulk-load a snapshot** of all products read from the database after step 1. 4. **Start a second consumer** that replays changes from the recorded position into the shadow, while the existing consumer keeps updating the live index. 5. **Wait for catch-up**: the shadow consumer's lag falls within the normal SLO. 6. **Verify**: compare document counts, sample queries against both indexes and compare results, and check error logs for rejected documents. 7. **Switch the alias** to the shadow, and make the shadow's consumer the live one. 8. **Keep the old index** for a rollback window, then delete it. ```pseudocode start_pos = change_log.current_position() shadow = create_index(new_schema) for batch in read_snapshot(source_db): // snapshot read begins after start_pos shadow.bulk_write_if_newer(batch) shadow_consumer = consume(change_log, from = start_pos, apply = shadow.write_if_newer) wait_until(shadow_consumer.lag() < lag_slo) if not verify(shadow, live): abort() // alias untouched, readers unaffected alias.point_to(shadow) ``` ## Why the order of steps 1 and 3 matters If the snapshot started **before** the position was recorded, changes committed between the two would be in neither the snapshot nor the replay, and the shadow would silently miss them. Recording the position first guarantees that every change is in the replay, the snapshot, or both. "Both" is the normal case, which is why every write must be **version-guarded**. A change can be in the snapshot and replayed later; the replayed copy carries a version equal to the stored one and is skipped. A replayed change can also be older than the snapshot row, and the guard rejects it. Deletes need versioned tombstones for the same reason. ## Two ways to handle in-flight writes | Approach | How it works | Trade-off | |---|---|---| | **Replay from a recorded position** | A second consumer reads the change log from the recorded position into the shadow | Live path untouched; needs log retention longer than the rebuild | | **Dual-feed during the rebuild** | The indexer writes every new change to both live and shadow while the bulk load runs | No long retention needed; the snapshot can overwrite newer dual-fed data unless writes are version-guarded, and the live indexer's code changes | Both approaches rely on the same versioning; replay is usually simpler to reason about because the live path does not change. ## Risks to plan for - **Log retention**: if the rebuild takes 10 hours and the change log keeps 7 days, there is room; if retention is shorter than the rebuild, replay starts from data that is already gone. - **Load**: the snapshot read and the bulk load compete with production traffic, so throttle both. - **Capacity**: two full indexes exist side by side, so plan roughly double the storage for the duration. - **Rollback**: keep the old index and the old consumer's position until the new one has served real traffic without problems.
- What do you verify before switching the alias to the shadow index?Document counts against the source of truth and the live index, allowing for the known lag. A sample of real queries run against both indexes, with the result differences reviewed, because a schema or text-analysis change is expected to change some results but not to empty them. Also the count of documents the shadow rejected, and the shadow consumer's lag, which should be inside the normal SLO.
- What happens if the change log's retention is shorter than the rebuild takes?By the time the shadow consumer starts, the events after the recorded position may already be deleted, so the replay has a gap and the shadow silently misses changes. Either extend retention for the duration, shorten the rebuild by throttling less or parallelising the load, or use dual-feed so new changes reach the shadow as they happen.
- How do you roll back if problems appear after the switch?Point the alias back at the old index, which was kept for this reason. The old index must still be current, so keep its consumer running during the rollback window, rather than stopping it at the switch. Once the new index has served real traffic cleanly for the agreed window, stop the old consumer and delete the old index.
saying these in an interview costs you the question
- Taking the snapshot first and recording the log position afterwards is fine.
- The replay never contains changes that the snapshot already holds.
- Writes can be paused for the whole rebuild without anyone noticing.
- Switching the alias as soon as the bulk load finishes is safe.
- The old index can be deleted the moment the alias moves.