A team adds a processed-message dedup table to every consumer in their event-driven system by default, even for consumers whose downstream operation is already a naturally idempotent upsert keyed by the entity's primary key. What is the cost of that blanket policy, and when should a team skip the dedup table entirely?
answer
- dedup table = ongoing write + storage + cleanup cost
- upsert-by-ID is naturally idempotent, skip the table
- guard non-idempotent side effects only (deltas, emails, external calls)
- whole-handler test, not just the main DB write
- search-index reindex vs marketing email example
basics
~20 sA separate 'have I seen this before' table costs extra storage, a database write on every message, and upkeep to stop it growing forever. If the actual operation is already safe to repeat on its own, like setting a status to a fixed value, that extra table isn't needed at all.
solid answer
~60 sA processed-message dedup table isn't free: it's an extra write, often in the same transaction, on every single message, an ever-growing table that needs a retention/cleanup policy so old entries don't degrade lookup performance or blow storage, and one more piece of infrastructure to test, monitor, and reason about. When the downstream operation is already naturally idempotent, an upsert keyed by the entity's own primary key, a SET status = X write, a PUT-style replace, redelivery is harmless without any dedup store at all, because reapplying the same operation converges to the same state regardless of how many times it runs. In that case a blanket policy of adding a dedup table to every consumer is pure overhead: extra latency and storage bought for a guarantee the operation already provides for free. Teams should reserve the dedup table for operations that are inherently non-idempotent by nature, deltas, appends, side effects on external systems like sending an email or charging a card, and skip it for consumers whose write is already a pure, deterministic upsert.
go deeper
Should sense that adding extra bookkeeping has some cost, even if they can't fully articulate storage/latency trade-offs yet.
Should identify at least one naturally idempotent operation type, upsert-by-ID or absolute write, as a case where a dedup table isn't needed.
Should articulate the concrete costs, extra write, storage growth, cleanup burden, and correctly classify a mixed handler, idempotent write plus non-idempotent side effect, as still needing a guard.
Should frame this as a per-consumer design decision informed by a cost/benefit analysis, discuss retrofitting/validation strategy, and recognize that adding guarding machinery introduces its own new failure modes.
## What the dedup table actually charges you A processed-message dedup table, or equivalent unique-constraint scheme, is a **real, ongoing cost, not a one-time setup cost**. - **Every message pays an extra write.** Every message that flows through a consumer using it incurs at least one extra database write, the INSERT that claims the message ID, on top of whatever the actual business write already costs, which directly adds latency to the hot path and roughly doubles the write volume against that table's underlying storage. - **The table itself grows without bound** as long as the consumer runs, since every processed message adds a row that, absent a cleanup job, is never removed; left unmanaged, this degrades index lookup performance over time and eventually becomes a non-trivial storage line item, especially for high-throughput consumers processing millions of events a day. - **It also adds an entire subsystem** — the dedup table, its constraint, its cleanup/TTL job, its failure modes such as the check-then-act race and the partial-claim-then-crash scenario — that has to be built, tested, and monitored, none of which is needed if the downstream operation is already safe to repeat. ## When that cost is worth paying That cost is worth paying when the operation being guarded genuinely isn't repeat-safe: - applying a delta, - appending a row to an audit log or ledger, - calling an external non-idempotent API where many payment gateways, unless they accept their own idempotency key, will create a second charge on a second call, - or sending a notification where a duplicate is directly visible to the end user. These operations have no notion of already at the target state; running them again is observably different from running them once, so something external to the operation has to intervene to stop the second run. ## What is already idempotent without any machinery But a large share of consumer operations in practice are already naturally idempotent without any extra machinery: - an upsert keyed by the entity's own primary key, - a `SET status = 'SHIPPED'` absolute write, - replacing a document wholesale in a document store by its ID, - or writing to a cache with a TTL keyed by the entity ID. For all of these, processing the same message once or five times produces the identical final row; there is no window where reprocessing does anything different from the first processing, so a dedup table adds cost without adding any correctness the operation didn't already have. A team applying a blanket every consumer gets a dedup table policy pays the storage/latency/complexity tax on these consumers for zero marginal benefit. ## The one question to ask per consumer The decision of when to skip the dedup table comes down to answering one question per consumer: **does reapplying this exact operation, with this exact payload, converge to the same final state no matter how many times it runs?** | Answer | The operations it covers | What to do | |---|---|---| | If yes | pure upsert-by-ID, absolute-value writes, deterministic replaces | skip the dedup table and let natural idempotency carry the guarantee | | If no | deltas, appends, external non-idempotent calls, anything with observable side effects beyond database state such as emails, SMS, or webhooks fired to a third party | the dedup table, or an equivalent idempotency-token scheme scoped to that specific side effect, earns its cost | ## The test applies to the whole handler **A subtlety worth naming:** naturally idempotent has to hold for the whole operation, not just the primary database write. A consumer that does a naturally idempotent upsert but also, in the same handler, fires a webhook to a partner system on every invocation is not idempotent overall, even though its database write alone would pass the test. The webhook call is the non-idempotent part that still needs guarding, even if the upsert next to it doesn't. ## Two consumers, one event stream **A concrete real-world example:** a search-index consumer that reindexes a product document on every product-updated event typically doesn't need a dedup table — indexing the same document twice with the same content just re-writes the same index entry, a classic naturally idempotent upsert. Contrast that with a marketing-automation consumer reacting to the same event by sending a your-product-was-updated email; that side effect has no natural idempotence, so it genuinely needs a processed-message check, or a downstream idempotency key passed to the email provider, even though it's reacting to the exact same event stream as the reindexing consumer next to it.
- How would you retrofit an existing blanket-dedup-table consumer to skip it safely once you've confirmed the operation is naturally idempotent?Remove the dedup check and its associated write, but keep monitoring/alerting on duplicate-processing counts for a while after the change, since naturally idempotent claims are easy to get subtly wrong, for example a side effect hiding inside the handler. It's safer to validate the claim in a staging environment by deliberately redelivering messages and diffing the resulting state before removing the guard in production.
- If a consumer's main write is a naturally idempotent upsert but it also increments a metrics counter for observability on every message, does it need a dedup table?Only if the metrics counter's accuracy actually matters for a decision; if it's just an approximate operational dashboard, a slightly inflated count from occasional redelivery is usually an acceptable trade-off rather than justifying a full dedup table for the whole consumer. If the counter feeds billing or an SLA calculation, it needs guarding just like any other non-idempotent side effect.
- Does using a dedup table ever make an already-idempotent operation less correct?Not less correct, but it can introduce new failure modes that didn't exist before, the check-then-act race, the partial-claim-then-crash scenario, or a TTL cleanup job that deletes a still-relevant entry too early, so adding the table isn't a strictly safe default choice; it trades one class of risk, rare already-solved duplicates, for a different class of risk, bugs in the new guarding mechanism itself.
It's like installing a security guard checking ID badges at the entrance to every single room in a building, even the ones that only ever contain a whiteboard you can erase and rewrite as many times as you like without consequence. The guard is worth it at the vault door; at the whiteboard room, it's pure overhead.
saying these in an interview costs you the question
- Treats a dedup table as free to add with no downside
- Can't identify any operation that's naturally idempotent without a dedup table
- Misses that a handler with a naturally idempotent DB write can still have a non-idempotent side effect elsewhere in the same handler
- Has no plan for the dedup table's unbounded growth (no TTL/cleanup)
- Applies 'skip the dedup table' as a blanket rule without checking the specific operation first