skip to content

You are setting the durability configuration for a new platform whose data ranges from payment ledgers to high-volume device telemetry. How do you decide what each dataset's commit should wait for, and what evidence backs the decision?

level: principalimportance: nice to knowfreq 32%

answer

  1. Start from cost of a lost ack, not from knobs
  2. Two axes: replayable? externally acted on?
  3. Critical / standard / regenerable → three commit levels
  4. Outbox always critical
  5. Tail lag is the real loss window; untested failover = hypothesis

basics

~20 s

Classify data by two questions: can the writer replay it, and did anyone outside act irreversibly on the acknowledgement? That yields an acceptable loss window per dataset, which maps to a commit level — relaxed local, node-durable, or replica/quorum-acknowledged — and the latency budget each costs. Then verify by testing real failover.

solid answer

~60 s

Do not pick one global setting. Classify each dataset on two axes: **replayability** (can the producer regenerate this?) and **external consequence** (did a user, a payment network, or another service act on the acknowledgement?). Those give an acceptable loss window, which maps directly onto a commit level: - **No acknowledged loss permitted** — ledgers, payments, identity, consent, outbox rows: quorum-synchronous commit across separate failure domains. Budget the extra milliseconds per commit. - **Seconds tolerable** — most application data: node-durable commit plus asynchronous replication, with replication lag monitored and alerted as the real loss window. - **Regenerable, high volume** — telemetry, clickstream, derived tables, caches: relaxed local commit is fair game, since the producer can replay. The evidence I want before committing to it: measured commit latency at each level on the real hardware and topology, the current replication lag distribution (p50/p99, not an average), the write volume per class, and a tested failover measuring actual loss and time-to-recovery rather than the documented figures. Then write it down as policy with owners, because the expensive failure mode is the classification quietly drifting as new tables appear.

go deeper

for a junior

Recognize that not all data needs the same durability, and that payment data is treated differently from metrics.

for a middle

Give the classification axes and name the three commit levels with what each survives.

for a senior

Add measurement: latency per level on the real topology, lag distribution as the live signal, tested failover, and the outbox exception.

for a principal

Own it as policy — classes with owners, durability as a schema-review field, alerting tied to each class's stated loss window, and an explicit argument for why uniform maximum durability is the wrong default.

## Start from consequences, not from knobs The engineering error is choosing a setting and then rationalizing it. Start from what a lost acknowledged transaction *costs*: - **Money or legal exposure** — a payment recorded to a customer but absent from the ledger, a consent withdrawal that vanished. Cost is unbounded and reputational. - **A broken promise to a user** — an order confirmation for an order that does not exist. Recoverable with support effort, corrosive at scale. - **A gap in a stream nobody will notice** — five seconds of missing metrics. Cost is approximately zero. Each class implies a different acceptable loss window, and the window is the input to the configuration, not the output. ## The two questions that classify almost everything 1. **Can the writer replay it?** If the producer retains the data and can resend (device buffers, a message log with retention, a source file, a rebuildable projection), lost transactions are recoverable operationally and the durability requirement drops sharply. 2. **Did something outside the database act irreversibly on the acknowledgement?** A card charged, an email sent, a message published, a third party told. If yes, losing the record creates a discrepancy no restart can fix, and the acknowledgement must not have been given until the write was safe at the required scope. Anything answering 'no replay' and 'yes external act' goes to the strongest setting. Anything 'replayable' and 'no external act' can take the cheapest. The middle is a judgement call, and defaulting the middle to node-durable plus monitored asynchronous replication is defensible almost everywhere. ## Mapping classes to configuration | Class | Example | Commit waits for | Rough latency added | |---|---|---|---| | Critical | ledger, payments, identity, outbox | quorum of replicas in separate failure domains | replica round trip: sub-ms same rack, single-digit ms cross-zone | | Standard | user profiles, orders, content | local durable flush; async replication | device flush only | | Regenerable | telemetry, clickstream, caches, derived tables | memory, flushed in the background | none on the critical path | Two structural points. First, **the outbox belongs in the critical class** even if the business rows it accompanies do not — losing an outbox row silently drops a side effect another system is waiting for. Second, cross-**region** synchronous commit is a different order of expense (tens to over a hundred milliseconds per commit) and should be reserved for a small, explicitly named set of data, if used at all; most organizations meet regional-disaster requirements with asynchronous cross-region replication plus a stated non-zero loss window. ## The evidence to gather - **Measured commit latency per level on the actual hardware and topology.** Cloud network storage and power-loss-protected NVMe differ by orders of magnitude; guessing is worthless. - **Replication lag distribution**, p50 and p99, under peak write load — not the average, because the loss you take is the lag at the *worst* moment, and lag is bursty. - **Write volume and transaction shape per class**, since flush-bound cost is per commit; sometimes batching removes the entire problem at no durability cost, which is always the better trade. - **A tested failover.** Kill the primary under load and measure the transactions actually lost and the time to serve writes again. A documented loss target that has never been observed is a hypothesis. ## Governance, because configuration drifts The long-term failure is not choosing wrong once; it is that six months later new tables have appeared and nobody classified them. So: write the classes down with an owner each, make the class an explicit decision in the schema-change checklist, alert on replication lag against the class's stated window, and re-run the failover test on a schedule. Also make the application idempotent on retry regardless of class — no durability setting removes the indeterminate-commit-response problem. ## What a strong answer sounds like Refuse the single global setting. Classify by replayability and external consequence, map to an acceptable loss window, map that to a commit level, price each in measured latency, protect the outbox specially, and close with verification: monitored lag as the live signal and a periodically tested failover. Acknowledging that most data sits comfortably in the middle tier — rather than gold-plating everything to quorum-synchronous — is the mark of judgement, because uniform maximum durability buys latency cost across the whole platform to protect data that did not need it.

  • Why not simply configure the strongest durability everywhere and stop reasoning about it?
    Because the cost lands on every commit across the whole platform: quorum acknowledgement adds a replica round trip to writes that did not need it, cutting per-connection throughput and raising tail latency for user-facing paths. It also couples write availability to replica health more widely than necessary. Uniform maximum durability is a real choice, but it should be made knowingly with the latency budget measured, not adopted to avoid classification work.
  • How do you keep the classification from rotting as the platform grows?
    Make it a required field in the schema-change review — every new table declares its durability class and owner — and alert on replication lag against each class's stated window so drift shows up as a signal rather than a surprise. Re-run the failover exercise on a schedule and publish the measured result. Without those, the policy becomes a document nobody reads within two quarters.
  • Where do batching and transaction shape fit into this decision?
    They often remove the pressure that made relaxed durability tempting in the first place. Because durable commit cost is per commit rather than per row, grouping a chatty ingestion path into batches can deliver most of the throughput gain while keeping full durability. Always evaluate that first; weakening the guarantee should be the remedy only after cheaper structural changes are exhausted.

saying these in an interview costs you the question

  • Applying one durability setting to the entire platform without classifying data
  • Leaving the outbox or side-effect-triggering tables in the relaxed tier
  • Quoting average replication lag as the expected data loss instead of the tail
  • Claiming a data-loss target that has never been verified by an actual failover test
  • Reaching for relaxed durability before trying batching, which costs no guarantee at all
  • Assuming a durability setting removes the need for idempotent retries

context