Your service missed GitHub webhook deliveries during an outage. How do you recover and prevent recurrence?
answer
- delivery is best-effort, not a log
- GitHub kept a record of every attempt
- replay is an API call, not magic
- the same GUID arrives twice
- a slow sweep catches what nobody noticed
basics
~20 sTreat GitHub webhook delivery as best-effort. Recover by listing the stored deliveries for the hook and redelivering the failed ones through the deliveries API, then make the receiver acknowledge fast, queue the work, deduplicate by delivery GUID, and reconcile periodically against the API.
solid answer
~50 sGitHub records every delivery attempt for a webhook — the event, the timestamp, the response status, and the payload — and exposes them: `GET /repos/{owner}/{repo}/hooks/{hook_id}/deliveries` lists them, and `POST /repos/{owner}/{repo}/hooks/{hook_id}/deliveries/{delivery_id}/attempts` replays one. A GitHub App has the equivalent under `/app/hook/deliveries`. So the immediate recovery is: identify the outage window, list deliveries, filter to non-2xx statuses, and replay them once the receiver is healthy. The prevention work is architectural. Return 2xx within GitHub's roughly ten-second window by writing the payload to durable storage or a queue and processing asynchronously, so a slow downstream never turns into a failed delivery. Make every handler idempotent and deduplicate on `X-GitHub-Delivery`, because replays and retries produce duplicates by design. Do not rely on ordering. And keep a periodic reconciliation job that compares your state against the API, because that is the only mechanism that catches events nobody noticed were missing.
code
bash · 11 linesHOOK=292430182
curl -s -H "Authorization: Bearer $TOKEN" \
-H "Accept: application/vnd.github+json" \
"https://api.github.com/repos/acme/widgets/hooks/$HOOK/deliveries?per_page=100" \
| jq -r '.[] | select(.status_code != 200) | "\(.id)\t\(.event)\t\(.delivered_at)\t\(.status)"'
# replay one, pacing the loop so the receiver is not flooded
curl -s -X POST -H "Authorization: Bearer $TOKEN" \
-H "Accept: application/vnd.github+json" \
"https://api.github.com/repos/acme/widgets/hooks/$HOOK/deliveries/$DELIVERY_ID/attempts"go deeper
Know that a webhook delivery can fail and that GitHub keeps a record of each attempt with its response status, which someone can look at and resend.
Explain the mechanics: the deliveries listing and the redelivery endpoint, the roughly ten-second response window, and why deduplicating on the X-GitHub-Delivery GUID is necessary.
Show the incident shape end to end — bound the window, list and replay failures at a safe pace, then harden with fast acknowledgement, queued processing, idempotent handlers and scheduled reconciliation.
Own the reliability contract: what staleness the business tolerates, whether the receiver is shared or per-team, how delivery health is monitored independently of application metrics, and how correctness is proven after an outage.
## Why the outage lost events at all A webhook delivery is a single HTTP request from GitHub to your endpoint. If your endpoint is down, slow, or returns a non-2xx status, that delivery is recorded as failed. There is no queue on your side holding it and no guarantee that anything will re-drive it for you, so a receiver that is offline for an hour has an hour-shaped hole in its event stream. The design assumption to state plainly in an interview is: **webhook delivery is best-effort and at-least-once, never exactly-once and never guaranteed.** ## Step 1 — bound the damage Establish the window: when did the receiver start failing and when did it recover? Your own logs give the start; the delivery log confirms it from GitHub's side. Then enumerate what was affected. For a repository webhook: `GET /repos/{owner}/{repo}/hooks/{hook_id}/deliveries?per_page=100` Each entry carries `id`, `guid`, `event`, `action`, `delivered_at`, `status`, `status_code`, `duration`, and whether it was itself a `redelivery`. Filter to the window and to entries whose `status_code` is not 2xx. For a GitHub App the same data is under `GET /app/hook/deliveries`, authenticated with the app JWT — which matters, because an app receives events from every installation and one failed window can span many customers. Note that the delivery log is a rolling window, not an archive. If you discover the gap weeks later, the deliveries may no longer be listed and redelivery is no longer an option — which is precisely why reconciliation exists. ## Step 2 — replay Once the receiver is healthy and can tolerate the load: `POST /repos/{owner}/{repo}/hooks/{hook_id}/deliveries/{delivery_id}/attempts` Each call re-sends the original payload. Two cautions. First, pace the replay — firing thousands of redeliveries in a burst can trip secondary rate limits on the API calls that drive it, and can overwhelm the receiver you just brought back. Second, the redelivered request carries the **same** `X-GitHub-Delivery` GUID, so if your handler is not idempotent you will double-process anything that partially succeeded before the failure. ## Step 3 — verify After replaying, do not assume the state is correct. Run the reconciliation job (below) over the affected repositories and compare against the API. That converts "we replayed 412 deliveries" into "our state matches GitHub's". ## Preventing recurrence **Acknowledge before you work.** The receiver's only job in the request path is: verify the signature, persist the raw payload with its GUID and event name, return 2xx. Everything else happens in a worker. GitHub expects a response in about ten seconds; any design where a downstream call sits inside the request handler will fail deliveries the moment that downstream is slow. **Make the ack path the most reliable thing you own.** It should have no dependency that can be down while the receiver is up — ideally an append to a durable queue or a single insert. If even that is unavailable, returning 5xx at least leaves an accurate failed entry in the delivery log for you to replay. **Deduplicate on the delivery GUID.** Store processed GUIDs with a TTL longer than any replay you would realistically perform. Combined with idempotent handlers, this makes replay safe and makes at-least-once delivery a non-issue. **Do not depend on ordering.** Two events fired within the same second can arrive in either order, and a replay arrives long after its neighbours. Where order matters, do not reconstruct state from the event sequence — use the event as a trigger to re-read current state from the API, which is order-independent by construction. **Reconcile on a schedule.** A slow sweep — hourly or nightly, depending on how much staleness you can tolerate — that lists the relevant resources through the API and repairs anything that disagrees with your store. Reconciliation is what makes the whole system correct rather than merely fast; webhooks provide latency, reconciliation provides truth. **Monitor the delivery log itself.** Poll the deliveries endpoint and alert on a rising count of non-2xx statuses. Your own error rate can look fine while GitHub is timing out on you, so this is genuinely independent signal. **Watch for auto-disable.** A webhook whose endpoint fails persistently can end up disabled, at which point new events stop arriving entirely and your error rate looks healthy again — the silence is the symptom. Check that the hook is still active as part of the recovery. ## How to present this in an interview Lead with the assumption (best-effort, at-least-once), then the mechanism (delivery log plus redelivery endpoint), then the architecture (fast ack, queue, idempotent handlers, GUID dedupe, no ordering assumptions), and close with the safety net (reconciliation and delivery-log monitoring). The signal an interviewer is looking for is that you do not treat a webhook stream as a reliable log — you treat it as a fast hint over an authoritative API.
- Why is redelivery not enough on its own as a recovery strategy?Because the delivery log is a rolling window and only covers webhooks you know failed. A gap discovered weeks later has nothing left to replay, and a receiver that returned 200 while failing internally leaves no failed entries at all. Reconciliation against the API is the only mechanism that repairs state you did not know was wrong.
- What makes replaying thousands of deliveries risky?Two things. The replayed request carries the original X-GitHub-Delivery GUID, so non-idempotent handlers double-process anything that partially completed. And a burst of replays both hammers the receiver you just restored and drives enough API calls to trip secondary rate limits. Pace the replay and verify idempotency before starting it.
- How do you detect that your receiver is failing when your own metrics look clean?Poll the deliveries endpoint for the hook and alert on non-2xx status codes and rising durations. A receiver can time out at the edge, or be unreachable entirely, without your application ever recording an error. Also check that the webhook is still active, since a persistently failing endpoint can end up disabled and the resulting silence looks like health.
- How should a handler cope with two events arriving out of order?Do not reconstruct state from the sequence. Use the event as a signal that something about a resource changed, then read the current state from the API and write that. This makes handling naturally order-independent and idempotent, so a stale replayed event converges on the same result rather than reverting newer data.
saying these in an interview costs you the question
- Assumes GitHub retries failed deliveries until success
- Processes the event inline and then returns 2xx
- Treats the webhook stream as an authoritative event log
- Ignores duplicate deliveries because they should not happen
- Replays every delivery at once after an outage