When is an after-commit effect that can be lost acceptable, rather than writing the notification into the same unit of work?
answer
- compare two costs, not one
- reconstructible from committed state
- who would notice the absence
- durable intent implies possible repeats
- best-effort must be stated, not assumed
basics
~20 sWhen the effect is reconstructible from committed state, cheap to redo, and its temporary absence costs only staleness. Effects that move money, are irreversible, or hold the only copy of something must be recorded inside the transaction instead.
solid answer
~50 sBoth options cost something, so the decision is a comparison. A plain after-commit send is one line, adds no schema and no delivery machinery, but can vanish without a trace. Writing the intent as a row in the same unit adds a table, a delivery step, ordering questions and repeat-tolerance on every receiver — in exchange for the change and the intent never disagreeing. Accept the lossy version when the effect can be rebuilt from state the database already holds, when something periodic will converge the two sides anyway, and when a delay or a miss is a stale screen rather than a wrong outcome. Refuse it when a lost effect creates a divergence nobody will notice and nothing can repair. And be honest that the durable option upgrades the guarantee to at-least-once, which is a requirement you are placing on every receiver.
go deeper
The takeaway to hold is that work placed after a commit can be lost. Before writing it there, ask whether anything would notice if it never ran.
Be able to sort effects into reconstructible and not. A cache refresh that the next read rebuilds is a different risk class from a payment or a customer-facing confirmation.
Argue the choice with detection in mind: how the absence would be found, how long the two sides may disagree, and what repairs them. Reject in-callback retry as a durability story.
Make it policy rather than a per-case instinct, and own the second-order cost: durable intent pushes repeat-tolerance onto every receiver, including teams you do not control, so that requirement has to be negotiated and written down.
## The decision, framed honestly Everything after a commit is on the far side of the transaction's protection. There are two ways to live with that, and both cost: | | Plain after-commit effect | Intent written inside the unit of work | |---|---|---| | What it guarantees | Runs only if the change is real | Change and intent share one fate | | What it costs | Can be lost silently | A table, a delivery step, operations around it | | Failure shape | Missing effect, correct data | Delayed or repeated effect | | Burden on receivers | None extra | Must tolerate a repeat | | Right when | The effect is reconstructible | The effect is irreversible or unique | Neither is the "senior" answer by default. Reaching for the durable option everywhere buys reliability the system does not need and pays for it in machinery every team member must then understand; reaching for the plain send everywhere eventually loses something that mattered. ## When loss is genuinely tolerable - **The effect is reconstructible from committed state.** Invalidating a cached value, refreshing a materialised view, re-indexing a record — if the next read or the next scheduled pass rebuilds it, a lost effect costs staleness for a bounded time. - **Something already converges the two sides.** A periodic reconciliation, a nightly rebuild, or a receiver that polls anyway turns a lost message into a delay. - **The effect is advisory.** Warming a cache, nudging a dashboard, emitting a metric. Nobody makes an irreversible decision on it. - **The volume makes each individual event worthless.** If losing one of ten thousand telemetry events changes no decision, durability machinery is spend with no return. - **The window of exposure is short and observable.** You can say, with a number, how long the two sides can disagree and how you would see it. ## When the intent must be inside the transaction - **The effect moves money or makes a promise to a person.** A charge, a payout, a confirmation the customer will act on. - **The effect is irreversible.** Once out, it cannot be recalled, so it must not happen for a change that rolled back — and must not be silently skipped for one that did not. - **It carries the only copy of something.** If the payload cannot be rebuilt from the committed data, losing the send destroys information. - **Nothing will notice the absence.** This is the criterion teams weigh least and should weigh most. A missing effect that no reconciliation, no dashboard and no customer will surface is a permanent, invisible divergence. - **A regulator, an audit, or a downstream contract requires the record.** Then the record is part of the write, not a consequence of it. ## The questions to ask in the design review 1. **If exactly this effect vanished right now, how would we find out — and how long would that take?** If the honest answer is "a customer would tell us in a month", the effect is not a candidate for a plain send. 2. **Can the effect be rebuilt from what committed?** If yes, a repair job is cheaper than a delivery pipeline. 3. **What is the cost of the opposite failure?** Durable intent implies delivery that may repeat. If a repeat is worse than a miss for this receiver, adding durability without repeat-tolerance makes things worse, not better. 4. **How many receivers, and do we control them?** Imposing repeat-tolerance on teams you do not own is a negotiation, not a decision. 5. **Is this the only effect of its kind, or the first of many?** One lossy send is a judgement call; a pattern of them is an architecture, and it should be a deliberate one. ## The trap in the middle The tempting compromise — retry the send inside the after-commit callback a few times — buys much less than it looks like. It survives a brief unavailability of the receiver and nothing else: it does not survive the process dying, it holds a thread and possibly a request while it waits, and it hides the failure rate behind eventual success. Retrying is a reasonable **optimisation** on top of a recorded intent; it is not a substitute for one. The second trap is claiming a guarantee the design does not have. If the effect is best-effort, the word *best-effort* belongs in the code, the runbook and the conversation with whoever depends on it. Most real damage from this area comes not from choosing the lossy option but from choosing it silently, so that a team downstream builds on a promise nobody made. ## The compression Ask what the absence of the effect costs and who would notice. If the answer is staleness that repairs itself, send after the commit and say so. If the answer is a divergence nobody sees and nothing fixes, the intent belongs in the transaction — and the price of that is receivers that can absorb a repeat.
- Why is retrying the send inside the after-commit callback a weak substitute for recording the intent?It only covers a receiver that is briefly unavailable. It does not survive the process dying, it occupies a thread while it waits, and eventual success hides how often the first attempt fails. Retry belongs on top of a recorded intent, where a later attempt can pick the work up again, not in place of one.
- What new obligation does writing the intent inside the transaction place on receivers?Repeat tolerance. Once the intent is durable, something will keep trying until it believes the effect happened, and the boundary between 'sent' and 'confirmed' is not perfectly observable — so the same effect can arrive more than once. Guaranteeing you never lose one means giving up the claim that you never send one twice.
- How do you decide the acceptable window in which the two sides may disagree?From what the business does during it. Name the decision made on the stale side and its cost per unit of time, then set the window below the point where that cost matters and monitor the actual lag against it. A window nobody can state as a number is a window nobody is enforcing.
saying these in an interview costs you the question
- Adds durable intent everywhere without asking what loss would cost
- Calls a best-effort send reliable in the design document
- Thinks in-callback retries make the effect durable
- Ignores that durable delivery forces receivers to tolerate repeats
- Cannot say how a lost effect would ever be noticed