A backfill re-sent last month's customer emails — how do you make pipeline side effects replay-safe?
answer
- not everything a task does is a row
- the mailbox has no partition to overwrite
- separate the computing from the telling
- gate delivery on the run being current
basics
~20 sSeparate computing data from delivering it. Put external effects — email, webhooks, partner file drops — in a step that only fires for current runs, gated on a backfill flag or on the window being recent, so a replay recomputes tables without re-notifying anyone.
solid answer
~50 sExternal effects have no partition to overwrite, so the replace-the-slice trick that makes table writes idempotent does nothing for them. Fix it structurally, in this order. **Split the graph**: compute and publish the data in one pipeline, deliver notifications from a separate one triggered only by current runs, so backfilling compute never touches the mailbox. **Gate what you cannot split**: pass a backfill indicator into the run, or compare the window's end to the current time, and skip delivery when the run is a replay. **Make the effect keyed** where the receiver supports it — write to an outbound table keyed by recipient and window and let a delivery step claim unsent rows, so a replay updates a row rather than sending twice. And **default to off**: notification suppression during backfills should be a platform behaviour, not something each author remembers. Remember the invisible side effect too — downstream pipelines triggered on completion will fan out once per replayed window.
code
python · 13 lines# deliver only for runs that represent a current window
DELIVERY_MAX_AGE = timedelta(hours=6)
def publish_summary(interval_start, interval_end, is_replay, now):
rows = build_summary(interval_start, interval_end) # idempotent write
upsert_summary(rows, interval_start)
stale = (now - interval_end) > DELIVERY_MAX_AGE
if is_replay or stale:
log.info("delivery suppressed: replay=%s stale=%s window=%s",
is_replay, stale, interval_start)
return
send_summary_email(rows)go deeper
Know that emails, webhooks and file drops cannot be undone, so a task that sends one is not made safe by an idempotent table write. Recognise the shape when you see it in a task.
Explain the two practical gates — an explicit replay indicator and a window-age check — and how writing an outbound row keyed by recipient and window converts a send into a write you can safely repeat.
Show you have handled the incident: split compute from delivery, walk the downstream graph for triggered pipelines, suppress freshness alerts for replays, and dry-run a backfill with an inert client to see what it would have sent.
Make suppression the platform default rather than per-team discipline: one wrapper for every external client, replay context passed into every run, and delivery from a replayed run requiring explicit opt-in.
## Why side effects are the exception Everything else in this topic works because the task owns a slice of storage it can replace. An email has no slice. Once it is delivered it cannot be un-delivered, and the recipient has no way to know it is a duplicate of something they read three weeks ago. The same is true of a webhook to a partner, a file dropped on an SFTP server, a payment call, a page to an on-call engineer, and an increment to an external counter. So the design question is not "how do I make the send idempotent" — usually you cannot — but "how do I make sure the send does not happen during a replay at all". ## Pattern 1: split compute from delivery The most robust fix is structural. One pipeline computes and publishes data for a window; a second pipeline reads the published result and delivers it. Backfilling the first re-computes tables and touches nothing external. The delivery pipeline runs on its own cadence for current windows only, and if you ever *do* want to re-deliver, you trigger it explicitly. This also fixes a subtler problem: it makes "we recomputed the numbers" and "we told people the numbers" two separately auditable events, which is what you want when a stakeholder asks whether the figures in Tuesday's email were the corrected ones. ## Pattern 2: gate the effect on the run being current When splitting is impractical, make the delivery step conditional. Two gates are common and they answer different questions: - **An explicit replay indicator** passed into the run — the operator states this is a backfill and the task branches. Precise, but relies on the flag actually being set. - **A recency test** — the window's end is compared to the current time, and delivery is skipped if the run is older than some threshold. It needs no operator discipline and it also catches the case where a run was stuck in a queue for two days and its notification is now misleading anyway. Both are better than nothing, and belt-and-braces is reasonable for anything with money or customers on the far end. ## Pattern 3: key the effect so the receiver can ignore repeats Some effects can be made safe at the boundary. Instead of calling the mail provider from the transformation task, write a row into an outbound table keyed by (recipient, window, message type) with a status column. A separate delivery step claims rows that have not been sent and sends them. A replay of the transformation re-writes the same row — an upsert on that key — rather than producing a second send request. Many partner APIs also accept a caller-supplied request key so they can drop retries themselves; use it when it exists. This converts an unrepeatable action into a table write, which you already know how to make idempotent, plus one narrow step whose only job is delivery. ## Pattern 4: a backfill mode with an inert provider For teams with many notification-emitting tasks, wire the external client through configuration: in backfill mode the mail/webhook/SFTP client is a no-op that logs what it would have sent. Authors get the behaviour for free, and the log becomes a useful diff — "this replay would have sent 61 emails" is exactly the review artefact you want before a large backfill. ## The side effects people forget **Triggering downstream work.** If downstream pipelines start when this one completes, a 90-window backfill kicks each of them 90 times. Everything they do — including *their* notifications — fans out. Check the whole downstream graph, not just your own task, before replaying. **Publishing to a shared, externally visible target.** Rebuilding a mart in place means consumers read a partially rebuilt table for the duration. Build aside and swap. **Alerting and data-quality checks.** Replayed runs can fire freshness or volume alerts that page a human at 3am for a window from last March. Suppress alerting for replay runs the same way you suppress delivery. **Metrics and counters.** Anything the task increments — a business counter, a usage meter that a customer is billed on — is an additive external effect wearing ordinary clothing. ## Making it stick A convention that depends on every pipeline author remembering to check a flag will fail on the day someone is under pressure. Push it down: give tasks a documented "is this a replay" input, route all external clients through one wrapper that respects it, default notification tasks to suppressed during backfills, and require an explicit opt-in to deliver from a replayed run. Then the failure mode reverses — the worst case becomes a notification that did not go out, which is recoverable, rather than 60 emails to every customer, which is not.
- Which is better, an explicit backfill flag or a recency check on the window?They cover different failures, so use both for anything customer-facing. An explicit flag is precise but depends on the operator setting it. A recency check needs no discipline and also catches a legitimately scheduled run that sat queued for two days, whose notification would be misleading regardless of why it was late.
- A team wants to keep the notification inside the transformation task. What is the minimum safe design?Have the task write an outbound row keyed by recipient, window and message type — upserted, so a replay overwrites rather than adds — and let a separate delivery step claim unsent rows. That turns the unrepeatable action into a table write you already know how to make idempotent, with one narrow step responsible for actually sending.
- What side effect of a backfill do teams most often miss entirely?Downstream fan-out. If other pipelines trigger on this one's completion, replaying 90 windows starts each downstream pipeline 90 times, and their side effects fire too — including their alerts and their notifications. Walk the downstream graph before a backfill and pause or gate consumers, not just the pipeline you are replaying.
saying these in an interview costs you the question
- Says the email is fine because the table write is idempotent
- Plans to raise the mail provider's rate limit before backfilling
- Relies on every pipeline author remembering to check a backfill flag
- Ignores downstream pipelines that trigger once per replayed window
- Lets freshness and volume alerts fire for replayed historical windows