skip to content

Your webhook worker's retries cause duplicate deliveries upstream. How do you decide what it may replay?

level: principalimportance: nice to knowfreq 25%

answer

  1. the duplicate lands in someone else's system
  2. not one policy for every endpoint
  3. replay needs the bytes kept somewhere
  4. memory equals concurrency times payload
  5. switch it off per route without a deploy

basics

~20 s

Treat replay as a decision about the receiver's data, not your client's convenience. Decide per endpoint with the team absorbing the duplicates, make the payload buffering that enables replay explicit and bounded, and let a route's replay be switched off without a release.

solid answer

~50 s

The retry count is the easy part; the real decision is which requests this client is allowed to put back on the wire, and that decision is not wholly yours — a replayed POST becomes a second row in someone else's system, so the team receiving the deliveries has standing to overrule you. I make it per endpoint rather than globally: replay is on by default only where the receiver has told us a repeat is harmless, and failures that occur before any bytes were written are treated separately from failures after, because only the latter can duplicate work. The enabling mechanism costs memory — replay requires holding the payload, so the bill is in-flight deliveries times payload size, and that number caps worker concurrency. Where that bill is too large, the honest answer is to re-drive delivery from the durable queue instead of holding bytes in RAM. Finally, I want a duplicate rate we both measure and a per-route switch that turns replay off without a deploy.

go deeper

for a junior

Understand that a retry can cause the same work to happen twice on the receiving side, and that replaying a request at all requires the payload to have been kept somewhere.

for a middle

Be able to explain the mechanics that make replay possible and what they cost: the payload must be held in memory for the retry window, which scales with how many deliveries are in flight.

for a senior

Show that you distinguish failures before the request was written from failures after, and that you size worker concurrency against the memory the buffered payloads actually consume.

for a principal

Own the policy: which endpoints may be replayed is agreed with the teams absorbing duplicates, the memory bill is a stated capacity decision, and the switch is per route and changeable without a release.

## What is actually being decided When a webhook delivery worker retries, three separate questions get bundled together and only one of them is technical: 1. **Can this request be replayed at all?** A `*http.Request` body is read once; replay requires the payload to still exist somewhere. That is mechanics. 2. **Should this request be replayed?** A retry after the bytes reached the receiver may create a second charge, a second email, a second row. That is the receiver's problem, not yours. 3. **What does enabling replay cost us?** Holding payloads in memory for the duration of a retry window is a capacity decision with a number attached. The reason this is a leadership call rather than an engineering preference is question two. Your client's convenience is being paid for out of another team's correctness budget, and the person absorbing duplicate deliveries can legitimately tell you to stop. ## Establish what a duplicate costs the receiver, per endpoint A global "retry POSTs three times" policy is a promise you made on someone else's behalf. Replace it with a per-endpoint decision, recorded next to the endpoint: - **Harmless repeat.** The receiver told you a second identical delivery is absorbed. Replay freely. - **Costly repeat.** A duplicate has a real-world effect — money, a message to a human, an irreversible state change. Replay only failures that provably happened before the request was written, and otherwise fail the delivery and surface it. - **Unknown.** Treat as costly. "We have never asked" is not "it is fine". The useful discipline is that this list has an owner on each side, and it changes by agreement rather than by someone tuning a constant in a config file. ## Separate the two failure positions A connection that was never established cannot have duplicated anything; a response that never arrived after the body was written may have been fully processed. The two deserve different defaults, and a client that collapses them into one "retryable error" category is throwing away the only information that makes the decision safe. In Go terms, this is why the client's error handling matters as much as the retry loop: the loop can only be as careful as the classification feeding it. ## Price the buffering, then bound it Replay is only possible if the payload survives the first attempt. In practice that means holding a `[]byte` per in-flight delivery for the whole retry window, so the memory bill is: concurrent deliveries x payload size x retry window occupancy That is a number you can compute before shipping, and it is the number that decides your worker's concurrency. A hundred concurrent deliveries of 4 KB events is nothing; a hundred concurrent deliveries of 8 MB payloads with a two-minute retry window is a memory profile that will page someone. When the bill is too high, the alternative is not "retry without a body" — that is a broken request — it is to stop holding the bytes at all: acknowledge nothing, leave the job on the queue, and let the next drain re-read the payload from durable storage. That moves replay from a hidden property of a client to a visible property of the pipeline, where it is already observable and already bounded. ## Make the policy visible and reversible Two properties matter more than the exact policy: **Visible.** A retry that is invisible at the call site is a behaviour change nobody reviewed. If retries live in a shared transport, the fact that a given call may be repeated should still be discoverable from the code that makes it, and the default for write methods should be off. **Reversible without a release.** When the receiving team calls to say duplicates are hurting them, the fix must be a configuration change on a route, not a deploy. Design for that on day one; you will need it during an incident, which is exactly when a deploy is most expensive. ## Instrument what the other team sees You cannot negotiate about duplicates without a number. Emit attempts per delivery and the count of deliveries that succeeded on attempt two or later — that second figure is the upper bound on duplicates you may have caused, and it is the figure to put in front of the receiving team. Without it the conversation is anecdote against anecdote, and the usual outcome is that retries get switched off entirely for endpoints where they were genuinely useful. ## Where I land, and what would change my mind Default: replay on for endpoints whose owner has agreed it is harmless, off for everything else, bytes buffered only within a concurrency limit derived from the memory bill, and the whole policy switchable per route at runtime. What changes my mind is scale — once payloads are large enough that buffering distorts the service's memory profile, re-driving from the queue is better on every axis except latency, and I would give up the latency.

  • Why is a single global retry policy for all outbound POSTs a poor default?
    Because the cost of a duplicate is a property of the receiving endpoint, not of your client. One endpoint absorbs a repeat silently, another sends a second email to a customer. A global setting applies your convenience uniformly to receivers with wildly different tolerances, and the ones who suffer have no say in the value you picked.
  • When would you stop buffering payloads for replay and re-drive from the queue instead?
    When the memory bill — concurrent deliveries times payload size across the retry window — starts shaping the service's capacity rather than being a rounding error. Re-driving reads the payload again from durable storage, so memory stops scaling with the retry window, the attempt is visible in the queue's own metrics, and the bound is the queue's redelivery policy instead of a constant in your code.
  • What would you put in front of the receiving team when they report duplicate deliveries?
    The count of deliveries that succeeded on a second or later attempt over the period, which is the upper bound on duplicates your client could have caused, alongside which endpoints they came from. That turns the discussion from blame into a per-endpoint decision, and it is also the metric that tells you whether switching replay off actually helped.

saying these in an interview costs you the question

  • Treats the retry count as the whole decision
  • Applies one replay policy to every outbound endpoint
  • Never asked the receiving team what a duplicate costs them
  • Ignores the memory held by buffered payloads awaiting replay
  • Requires a deploy to disable retries during an incident