skip to content

What does a deduplication record buy a payment service, and what does a customer see if a restart emptied the store holding it?

level: juniorimportance: should knowfreq 58%

answer

  1. state that exists nowhere else
  2. the second arrival must look different
  3. empty store, so a fresh request
  4. the side effect runs a second time
  5. degrades to at-least-once, does not break

basics

~20 s

A deduplication record marks a request identifier as already handled, so a repeated arrival of the same work is skipped instead of charged again. If a restart empties the store, the record is gone and the customer is charged twice.

solid answer

~40 s

A deduplication record is the entry that says *this request identifier has already been handled*, so the second arrival of the same work is recognised and skipped. It is state that exists nowhere else: nothing can recompute it, so once it is gone the second arrival looks exactly like a first one. A restart is one way it goes. Stores in this class differ here - some keep nothing at all across a restart, others reload a copy from disk that usually lags the most recent writes - but either way the records covering requests still in flight are the ones most likely to be missing. The service then performs the side effect a second time and the customer is charged twice. Nothing was lost; once-only handling simply degrades to at-least-once.

go deeper

for a junior

Be able to say in one line what the entry holds - this request identifier was already handled - and that a store of this class can lose it without being broken.

for a middle

Explain why nothing can rebuild the record: it is not a copy of anything, so an empty store makes a repeat arrival indistinguishable from a first one.

for a senior

Name the cause of the disappearance you are discussing and price it. A restart, a removal under memory pressure and a promoted copy are three different incidents.

for a principal

Decide whether an occasional duplicate side effect is a priced degradation or an unacceptable outcome, because that one call decides whether this record may live on a volatile tier at all.

## What this entry actually is A deduplication record is a small entry whose **presence is the whole message**: it says that a unit of work identified by some stable identifier - a request identifier, a message identifier, an order reference minted by the sender - has already been handled. The value may be nothing more than a marker, or a short state such as `in-progress` or `finished`. The key is what carries the meaning. It belongs to a family of things a shared volatile tier is asked to hold that are **not a copy of anything else**. A cached product page can be fetched again from the database that owns it, so a miss costs a fetch. A deduplication record cannot be fetched again from anywhere, because the fact it records - *we already did this* - lives nowhere else. That single property is what makes losing it interesting. ## Why a second arrival exists at all Senders retry because they did not learn the outcome, not because they know the first attempt failed. A timed-out connection, a proxy that closed the socket while the work was still running, a consumer that crashed after doing the work but before recording its progress, an operator replaying a batch that looked stuck - in every one of those the sender's honest state is *unknown*, and the safe thing for a sender in an unknown state is to send it again. So the receiving service sees the same work twice and has to decide which arrival it is looking at. The record is how it decides: - **record present** - this is a repeat, so skip the side effect; - **record absent** - this is new, so do the work and write the record. Everything this workload promises rests on the second line being reached only when the record is genuinely absent because the work genuinely never happened. ## The ways the record disappears Four causes are worth separating, because they are different incidents with different fixes. | Cause | What is true of the record | What the customer gets | |---|---|---| | Its deadline passed | absent, by your own design | a retry after the deadline is charged again | | Removed while the tier was under memory pressure | absent, and not because of your design | duplicates in bursts, correlated with other people's traffic | | The process restarted keeping nothing | a whole population of records absent at once | every request still in flight is charged again on retry | | A copy was promoted that never received the write | absent on the node now serving | duplicates only around a failover | Stores in this class differ on the last three in ways worth saying out loud. Some keep nothing whatsoever across a restart; others reload a copy from disk, which typically lags the newest writes by some interval, so the records most likely to be missing are exactly the ones for requests that were in flight. Some deployments acknowledge a write as soon as the receiving node has it and let copies catch up afterwards; others will not acknowledge until a copy holds it, and some let a caller ask for that per call. Whether the fourth row can happen to you is therefore a property of your store and your configuration, not of the class. ## What the customer sees The consequence always has the same shape: **the side effect happens a second time**. A card is charged twice, a payout is sent twice, a shipment is created twice, a welcome message arrives twice, a callback fires twice at a partner who then does their own work twice. Notice what does *not* happen. Nothing is lost. No request goes unprocessed. The honest way to name the outcome is that once-only handling **degrades to at-least-once**: the work still happens, it may just happen more than once. That phrasing earns its keep in an interview, because it tells the listener what you would go looking for after an incident - duplicate effects, not missing ones - and because it sets up the only question that matters next: is a duplicate here a wasted computation, or a second charge? ## What the record is not Two boundaries stop this drifting: - It is **not the guarantee**. An entry on a tier that can remove it for reasons that have nothing to do with your request cannot be the only thing standing between a customer and a second charge. Where the duplicate is intolerable, the guard belongs in the durable store that records the effect, which can declare the identifier unique; the volatile record then does the cheap, useful job of keeping most duplicates away from that store. - It is **not a cache entry**. A cache miss costs a fetch from the source that still holds the value. This miss costs a second side effect, and no amount of warming will refill it.

  • Why is this record not simply a cache entry with a different name?
    A cache entry is a copy of something a slower source still holds, so a miss costs a fetch and nothing else. A deduplication record is the only evidence the work was done; a miss cannot be filled from anywhere, and instead of recomputing a value the system quietly performs the side effect again.
  • If the record disappears, has any data been lost?
    No. Nothing stored elsewhere is gone. What is lost is the knowledge that the work already happened, so the failure shows up as a repeat rather than a gap: the payment is taken a second time, the shipment is created twice, the message arrives again. Look for duplicates after such an incident, not for missing work.

saying these in an interview costs you the question

  • Says the store is broken when the record is missing, rather than volatile by design.
  • Assumes every store of this class preserves what was written across a restart.
  • Calls the record a cached value, though nothing anywhere can refill it.
  • Believes losing the record loses the work, rather than repeating it.
  • Treats a volatile record as proof the charge can only ever happen once.