skip to content

A service writes a new order to a base orders table, then makes a second, separate call to write a pointer into an index table keyed by customer email so orders can be looked up by email later. The process crashes after the first write succeeds but before the second write is attempted. What is the resulting failure mode, and what are two concrete ways to guard against it?

level: seniorimportance: must knowfreq 70%

answer

  1. orphaned base row = missing index entry, silent failure
  2. dangling pointer = index entry with no base row, loud failure (404)
  3. transactional write removes intermediate state where scope allows
  4. outbox pattern: write a pending-update record in the same transaction as the base write
  5. reconciliation job as backstop

basics

~20 s

The order exists but is invisible to anyone searching by email, because its index entry was never created. Fix it by either doing both writes as one atomic operation when possible, or by having a background process that finds and repairs these gaps automatically.

solid answer

~50 s

This produces an orphaned base row / missing index entry: the order is real and durable in the base table, but any query path that goes through the email index table will never find it, because that index entry simply doesn't exist. It's silent — there's no error, just an invisible order. Two concrete guards: (1) Use a transactional write where the store's scope allows it (e.g., a scoped multi-item transaction covering both tables), so the two writes succeed or fail together and this partial state can't occur. (2) Where that's not possible, use an outbox pattern — write the base row and a 'pending index update' record together in one transaction against the base table alone, then have an idempotent background worker drain the outbox and create the index entry, retrying until it succeeds, with a reconciliation job as a backstop to catch anything the worker itself failed to process.

go deeper

for a junior

Should be able to describe in plain language that the order becomes 'invisible' to email search, without needing the outbox/transaction terminology.

for a middle

Should be able to name at least one concrete mitigation (transactional write or a retry-with-cleanup approach) and explain why the crash creates a partial state.

for a senior

Should know the outbox pattern by name or mechanism, explain why it's safe by construction, and understand the scope limits of transactional writes that make the outbox necessary in the first place.

for a principal

Should design the full three-layer defense (transaction where scope allows, outbox + idempotent worker otherwise, reconciliation as backstop) and reason about the operational cost/monitoring each layer requires.

## Why it is dangerous The scenario is the canonical partial-failure case for the index table pattern, and it's worth being exact about why it's dangerous: it fails silently. The base write succeeded, so from the perspective of the code path that created the order, everything worked — there's no exception, no error log, no retry triggered. The order is sitting in the base table, complete and correct, fully queryable by its own primary key. The only thing wrong is that a second, logically related write never happened, and nothing about the system's normal operation surfaces that gap until someone specifically tries to find that order through the email index and gets a false negative — a missing search result that looks exactly like 'this customer has no orders,' which is indistinguishable from the truth unless you know to be suspicious of it. ## Why it happens Why this happens is structural, not accidental: two writes against two independent tables, issued as two separate network operations from application code, have no inherent atomicity unless something explicitly provides it. Any of these leaves the sequence half-done: - a crash; - a timeout; - a deploy that kills the process mid-request; - a throttling response from the second table that the caller doesn't retry correctly. Because the base write happened first (it usually should, since the base table is the source of truth and losing an order entirely is worse than losing its discoverability by email), the specific shape of the failure is always the same: base row present, index entry missing. The mirror-image failure — an index entry pointing at a base row that was never created, or that was later deleted — is the other half of this same problem class (a dangling pointer), and shows up as a downstream 404 or null when the application follows the pointer, which is at least loud rather than silent. ## The transactional approach The two mitigations named above work at different points in the pipeline. The transactional approach removes the failure mode structurally: if the store's transaction API can scope both the base write and the index write into one atomic unit (as discussed for the general consistency-strategy question — DynamoDB's TransactWriteItems, or an Azure Table Storage entity-group transaction when both entities happen to share a partition key), then there is no intermediate state for a crash to land in; the operation either fully happens or fully doesn't, and a retry after a crash is safe because nothing partial was left behind. Its limitation is scope: it only works when both writes fit within whatever boundary the store's transaction mechanism allows, which frequently excludes cross-table, cross-partition scenarios like customer-email index tables keyed differently from the base table's own key. ## The outbox pattern The **outbox pattern** handles the cases the transaction can't reach. The trick is to convert the second write into something that can be transactionally attached to the first write, by writing a small 'index update needed' record into the same base table (or a companion outbox table in the same transactional scope as the base write) in the very same operation that creates the order. That outbox record is now guaranteed to exist if and only if the order exists, because they were written together. A separate, decoupled background worker then reads unprocessed outbox records, performs the actual index-table write, and marks the record processed — and because this worker's job is only to complete work that's durably recorded, it can retry indefinitely without risk of losing track of a pending update, and it can be built idempotently (writing the same index entry twice is harmless) so retries after its own partial failures are safe too. ## Reconciliation as the backstop Neither mechanism alone is a complete guarantee in practice — the outbox worker itself can have bugs, get stuck, or fall arbitrarily behind under load — so production systems typically add a third layer: a periodic reconciliation job that scans the base table (or consumes its change stream) and independently verifies that every base row has a corresponding index entry, repairing any gap it finds. This is the backstop that catches the failure the outbox mechanism was supposed to prevent but didn't, in whatever rare edge case slipped through — a worker crash mid-processing, a bug in the idempotency key, a dropped message in a queue-based variant. A concrete real-world instantiation of this three-layer approach is common in DynamoDB-backed services: 1. a conditional/transactional write creates the base item plus a DynamoDB Streams-visible marker; 2. a Streams-triggered Lambda performs the async index-table write; 3. a scheduled Lambda or Step Functions job runs a full-table reconciliation on a slower cadence (hourly or daily) to catch what the streaming path missed.

  • Why is the base write typically done first, rather than the index write?
    Because the base table is the source of truth for the entity itself — losing the order entirely would be a worse outcome than the order existing but being temporarily undiscoverable by email. Ordering the writes so the more critical one happens first, and building recovery around the less critical one, is a deliberate risk-prioritization choice.
  • How does the outbox pattern avoid the same partial-failure problem it's trying to solve?
    It works because the outbox record and the base row are written in the very same transaction against the same table, so they're atomic with each other by construction — there's no separate network call that can fail independently. The actual index-table write is deferred to a background worker that can safely retry against the durable outbox record without any risk of losing track of the pending work.
  • If you find a dangling index entry (pointing at a base row that no longer exists) during reconciliation, is it safe to just delete it?
    Generally yes, since by definition it's pointing at nothing valid and any query following it would fail anyway — but the reconciliation job should log or alert on these findings rather than silently deleting them, since a spike in dangling entries can indicate a bug in the deletion path (e.g., the base row's delete isn't cleaning up its own index entries) that needs fixing at the source.

It's like mailing a package (the base write) but forgetting to update the tracking website (the index write) because you got interrupted right after dropping it at the post office. The package is real and on its way, but anyone checking the tracking site sees nothing and assumes it was never sent.

saying these in an interview costs you the question

  • Assumes a crash between two independent writes can't really happen in practice / dismisses the scenario as unlikely enough to ignore
  • Proposes only 'add error handling' without describing a concrete mechanism (transaction, outbox, reconciliation)
  • Doesn't distinguish the orphaned-row case (silent) from the dangling-pointer case (loud, 404-like)
  • Assumes retrying the whole request from the client is sufficient without discussing idempotency of the retried writes
  • Treats a reconciliation job as optional/unnecessary once an outbox or transaction exists

context