skip to content

In a replicated wide-column store whose replicas repair each other, why must a delete outlive replica repair, and what happens if its tombstone is purged too early?

level: seniorimportance: must knowfreq 46%

answer

  1. a replica missed the delete
  2. repair copies what is newest
  3. no marker, data looks new
  4. grace period vs repair cadence
  5. one owner avoids it

basics

~10 s

If a replica missed a delete and the others have purged the tombstone, repair sees the old data only there and copies it back. Tombstones must outlive the longest gap between successful repairs.

solid answer

~50 s

In stores whose replicas reconcile by comparing data, a delete is safe only while its **tombstone** exists. Suppose a replica is down when a row is deleted. The others store the tombstone; the down replica still holds the row. When it returns, repair compares replicas: the tombstone is newer than the row, so the delete wins and spreads. But if the others have **already purged** the tombstone, repair sees a row on one replica and **nothing** on the others, treats the row as live data they are missing, and copies it back: the deleted data is **resurrected**. Hence tombstones are kept for a **grace period** that must exceed the longest gap between successful repairs (and any replica outage), and repair must reliably run within it; a replica down longer must be rebuilt, not simply rejoined. Stores that serve each range from one server have no peer reconciliation within a cluster, so they can drop markers at a major compaction.

go deeper

for a junior

Know that deleted data can come back in a replicated store if the delete marker disappears before every replica has seen it.

for a middle

Walk through the resurrection sequence and explain why repair treats absence as missing data.

for a senior

Set the grace period against the real repair cadence and outage lengths, monitor repair completion, and rebuild long-absent replicas.

for a principal

Be ready to own the policy linking repair schedules, grace periods and deletion guarantees, including what the business is promised about deleted data.

## The setting This problem belongs to stores where **several replicas each hold a copy** of the data and **repair** them against each other — comparing what each replica holds and copying whatever is newest to the replicas missing it. The mechanics of that repair (anti-entropy trees, read repair, stored hints for down nodes) are a subject of their own. Here the question is how deletes survive it. ## Why tombstones must be kept at all A delete is recorded as a **tombstone** with a timestamp. During repair: - replica A has the **row** (it missed the delete); - replicas B and C have the **tombstone**, newer than the row. Repair sees a newer tombstone and spreads it to A. The delete wins. This works **only because the tombstone still exists** on B and C. ## How resurrection happens Now suppose B and C have **purged** the tombstone during compaction before A was repaired: 1. A still holds the row. 2. B and C hold nothing for that key — neither row nor tombstone. 3. Repair sees data on A and **absence** on B and C. Absence is indistinguishable from "never received it". 4. Repair copies the row from A to B and C. 5. The deleted row is **back** on every replica. The same happens if a replica stays offline longer than the tombstone lifetime and then rejoins. ## The rule: a delete must outlive repair Stores in this model keep tombstones for a configurable **grace period** before compaction may purge them. Safety requires: | condition | why | |---|---| | grace period longer than the maximum interval between successful full repairs | every replica sees the tombstone before it can vanish | | repair actually completes on schedule | a failing repair job silently breaks the guarantee | | a replica down longer than the grace period is rebuilt, not rejoined | its stale data would otherwise be treated as live | | hints or queued writes older than the grace period must be discarded, not replayed | replaying them could reintroduce old data | ## The trade-off in the grace period - **Longer** grace: safer against slow repairs and long outages, but tombstones pile up, costing disk and read time. - **Shorter** grace: less tombstone overhead, but repair must run more often and more reliably. Tables that never delete explicitly and only age data out through one uniform time-to-live sometimes use a shorter grace period — a decision to make deliberately. ## Why the other model does not face this within a cluster In stores where **each key range is served by one server** and redundancy comes from a replicated storage layer beneath it, there is no peer reconciliation of rows inside the cluster: one server applies every delete. Once a major compaction has rewritten a range's files, no older data for it remains, so delete markers can be dropped then. Cross-cluster copies, where used, receive the delete as a replicated mutation. ## Interview angle Tell the story of the missed delete step by step, state the rule — tombstones must outlive the repair cycle — and name the operational duties it creates: repair on schedule, rebuild long-absent replicas.

  • A replica was offline for longer than the grace period. What should operators do before it serves traffic again?
    Treat its data as untrustworthy for deletes. Wipe it and rebuild it from the other replicas, rather than rejoining it and letting repair spread rows the others deleted and have since purged.
  • Why can a failing repair job be more dangerous than a slow one?
    If repair silently stops completing, tombstones still expire on schedule. Any replica that missed a delete during that time can resurrect data once the markers are gone, with no error to warn anyone.

saying these in an interview costs you the question

  • Setting the tombstone grace period shorter than the repair interval
  • Rejoining a replica that was offline longer than the grace period without rebuilding it
  • Believing repair can tell a purged delete apart from missing data
  • Assuming every wide-column store needs a tombstone grace period for this reason