skip to content

How do you test that no data is corrupted when a service is killed abruptly mid-write?

level: seniorimportance: should knowfreq 44%

answer

  1. Restarting proves almost nothing
  2. Pick the moment, do not roll dice
  3. Only promises must be kept
  4. Lost, duplicated, orphaned
  5. Crash the recovery too

basics

~20 s

Kill deliberately at the boundaries between irreversible steps, restart, then judge by invariants over the stored data: acknowledged operations minus applied ones must be empty, nothing applied twice, no orphaned rows, and cross-store totals reconcile.

solid answer

~40 s

Choose the kill points rather than killing at random: before the commit, after the commit but before the event is published, after publishing but before the acknowledgement, and during the startup recovery itself — then repeat each with randomised jitter. Kill hard, so you exercise the crash path and not the shutdown path. The oracle is a set of invariants evaluated after restart, never the logs or a health check. Have the client record every operation the system acknowledged, then assert three sets are empty: acknowledged-but-not-applied, applied-more-than-once, and rows failing a structural or cross-store invariant. Operations that were never acknowledged may legitimately be present or absent, so only the acknowledged set is a durability obligation. Add domain invariants a generic tool cannot guess, and allow a bounded settle window after restart before asserting.

code

pseudocode · 15 lines
pseudocode
kill_hard(service, at = "40ms after commit")
restart(service)
wait_until_settled(max_seconds = 30)

acknowledged = client.acknowledged_operations()   # 1847 track adds
applied = store.applied_operations()

lost = acknowledged - applied
duplicated = applied.appearing_more_than_once()
orphans = store.query("track rows whose playlist row is absent")

assert lost == empty
assert duplicated == empty
assert orphans == empty
assert store.track_count(9142) == feed.track_added_events(9142)

go deeper

for a junior

Know that an abrupt kill can leave an operation half-done, and that checking the service restarted is not checking the data. Be ready to name the three damage shapes: a partial write, an orphaned record, and an operation applied twice after redelivery.

for a middle

Explain why a write to one store plus a message to another is not atomic, and what each kill point can leave behind. Be ready to describe an invariant you would evaluate after restart instead of relying on logs or a health check.

for a senior

Show the full method: chosen kill points with jitter, a hard kill rather than a graceful stop, an acknowledged-versus-applied reconciliation, domain invariants, a bounded settle window, and a run that crashes the recovery itself. Expect to be asked how you avoid false failures from asynchrony.

for a principal

Own how far this is worth taking across a portfolio: which services carry durability promises strong enough to justify continuous crash testing, what the reconciliation costs to build and keep honest, and how you stop teams from quietly redefining a promise so the check keeps passing.

## The question is not whether it restarts Killing a process is easy; deciding whether the data survived is the hard part, and "the service came back up" answers a different question. An abrupt kill — no shutdown hook, no draining, no flush — leaves a system somewhere between two consistent states, and a crash test is only as good as the invariants you can evaluate afterwards. ## The three damage shapes **Partial writes.** An operation that spans more than one durable step is interrupted between them. Inside a single transactional store this is usually safe; across two stores it never is. The classic case is a write to the primary store followed by a publish to a message broker: kill between them and the record exists with no event, kill after the publish but before the acknowledgement and the event exists and may be delivered again. **Orphaned records.** A row whose parent or paired row was never written, or was written and then rolled back while the child was not. These accumulate silently and only surface later as a query returning something impossible. **Replayed messages.** At-least-once delivery means that anything unacknowledged when the process died will be delivered again. That is correct behaviour from the broker; the defect is a handler that applies it twice. ## Designing the kill Random kills find shallow bugs. Choose the kill points deliberately, at the boundaries between irreversible steps: before the commit, after the commit but before the event is published, after the publish but before the acknowledgement, during the startup recovery itself. That last one matters more than people expect — recovery code is the least exercised code in the system, and a recovery that is not itself idempotent turns one crash into a permanent corruption. Then repeat each point with randomised jitter over many runs, because the interesting window is often microseconds wide. Kill hard. A graceful stop tests the shutdown path, which is a different and much friendlier path than the one an out-of-memory kill or a yanked node takes. ## The oracle The verdict must come from **invariants over the resulting state**, not from the logs and not from the health check. Build the workload so that the expected end state is knowable: a client that records every operation it issued and, separately, every operation the system acknowledged. After the restart, compute three sets: - **Lost** = acknowledged minus applied. Must be empty. Anything here is a durability bug: the system promised and did not deliver. - **Duplicated** = operations applied more than once. Must be empty. Anything here is an idempotency bug. - **Orphaned** = rows failing a structural invariant, plus cross-store totals that do not reconcile. Note the asymmetry that makes this fair: an operation issued but **not** acknowledged may legitimately be absent *or* present — the system never promised either way — so only the acknowledged set is a durability obligation. Getting that distinction right is what separates a crash test from a flaky one. Add domain invariants that no generic tool can guess: a balance equals the sum of its entries; a collection's item count matches the count of item-added events; no state machine sits in a transient state with no owner. ## A worked example A playlist service accepts batched track additions. A test issued adds continuously, killed the process with an uncatchable signal 40 milliseconds after each commit, and restarted it. Of 1,847 acknowledged adds, three were absent from the playlist yet present in the activity feed — a dual-write divergence, the record and its event written to two stores with no atomic link. A separate run killed the process before the broker acknowledgement and produced 14 duplicated tracks on redelivery, because the handler keyed on nothing and simply appended. Both defects were invisible to the service's own health check, which reported healthy within four seconds of restart on every run. The fixes are the standard ones — make the handler idempotent on the operation's own identity, and make the record-plus-event pair atomic rather than two independent writes — but the point for a test engineer is that neither defect would have been found by a test whose assertion was "the service restarts". ## Practical cautions - Cross-store reconciliation needs a quiet moment or a consistent snapshot; comparing two live stores under load produces false failures and teaches the team to ignore the check. - Distinguish "lost" from "not yet applied": allow a bounded settle window after restart, then assert, rather than asserting instantly and blaming the system for asynchrony. - Run the reconciliation before and after the fault. If it fails before, the test found a pre-existing bug, and reporting it as a crash defect wastes an investigation. - Automate the cleanup of what the run leaves behind, or the next run starts from someone else's corruption and every result becomes arguable.

  • How do you choose where in the flow to kill the process?
    At the boundaries between irreversible steps — before the commit, between the commit and the publish of its event, between the publish and the acknowledgement, and inside the startup recovery — because those are the gaps where two durable facts can disagree. Then repeat each point with randomised jitter across many runs, since the vulnerable window is often microseconds wide and a single well-timed run proves little.
  • What tells you that a redelivered message was applied twice rather than once?
    Not the delivery count, which the sender controls, but the end state: an operation identity appearing more than once in the applied set, plus domain totals that overshoot — an item count higher than the count of distinct add operations, a balance larger than the sum of its acknowledged entries. Redelivery itself is expected behaviour; the defect is only visible as duplicated effect.
  • Why does a kill during the startup recovery deserve its own case?
    Recovery code is the least exercised code in the system and it runs precisely when the state is already inconsistent. If replaying or repairing is not itself idempotent and restartable, a second crash mid-recovery converts a survivable interruption into permanent corruption. Test it by killing at intervals through the recovery and asserting the same invariants once it finally completes.

Pulling the power mid-sentence: the question is not whether the machine boots again, but whether the sentence is either finished or absent — never half-written.

saying these in an interview costs you the question

  • Says the data is fine because the service restarted cleanly
  • Kills the process at one convenient point only
  • Assumes one store's transaction protects a write to a second store
  • Treats redelivery as impossible rather than expected
  • Compares row counts but never cross-store totals
  • Never crashes the system during its own recovery

context