skip to content

When an SMS provider used by a notification service starts returning errors and timeouts, how should retries and failover to a second provider work?

level: seniorimportance: should knowfreq 50%

answer

  1. not every error deserves a retry
  2. 429 is a request to slow down
  3. backoff, jitter, cap, deadline
  4. a secondary nobody used is untested
  5. accepted is not delivered

basics

~20 s

Classify failures: never retry permanent rejections, slow down on HTTP 429, and retry transient errors with capped exponential backoff and jitter. A per-provider circuit breaker shifts traffic to a warm secondary, and delivery-status callbacks, not API success, decide whether a provider is healthy.

solid answer

~50 s

First I **classify the failure**. A permanent rejection, such as an invalid number or a blocked recipient, is marked failed and never retried. HTTP 429 means slow down: honour `Retry-After` and lower that provider's send rate. Timeouts, connection errors and 5xx responses are **transient**: retry with exponential backoff, jitter and a cap, within the message's deadline. A **circuit breaker** per provider tracks error rate and latency; when it opens, workers route to a **secondary provider** that is already carrying a small share of traffic, so its sender registration and capacity are known to work. Timeouts are ambiguous, so failing over after one can double-send; that is acceptable for codes and usually not for marketing. Finally, an accepted request is not a delivered message: I consume **delivery-status callbacks** and fail over when the delivered rate drops, even if the API still returns success.

go deeper

for a junior

Recall that some provider errors are worth retrying and some are not, and that a backup provider can take over during an outage.

for a middle

Explain the failure classes, exponential backoff with jitter, why retries need a cap and a deadline, and what HTTP 429 with Retry-After asks the client to do.

for a senior

Show circuit-breaker states, a warm secondary with a steady traffic split, the timeout duplicate trade-off per message class, and delivery-rate alerts driven by callbacks.

for a principal

Weigh single versus multi-provider strategy: integration and compliance cost, per-region routing, contract minimums, and how much outage risk each message class justifies paying for.

## Classify before you retry A notification worker calls an external email or SMS provider over HTTP. Treating every failure the same way either wastes money retrying hopeless sends or gives up on messages that would have succeeded a second later. The first step is classification: | Failure | Meaning | Action | |---|---|---| | Invalid number, unknown recipient, content rejected | permanent | mark `FAILED`, no retry, maybe flag the contact | | Recipient has blocked the sender | permanent | mark `FAILED`, add to suppression | | HTTP 429 with `Retry-After` | throttled | pause this provider until then, reduce rate | | HTTP 5xx, connection reset, timeout | transient | retry with backoff, same idempotency key | | Authentication failure | configuration | stop sending, page the on-call | Classification tables are provider-specific in detail, because providers report errors differently, but every integration should map its errors onto these classes. ## Retrying transient failures Retries follow a few rules: - **Exponential backoff with jitter.** Wait roughly 1, 2, 4, 8 seconds, each randomised, so thousands of workers do not retry in lockstep. - **A cap and a deadline.** Stop after a fixed number of attempts or when the message's deadline passes; a login code that is already expired is dropped, not retried. - **A retry budget.** Limit retries to a fraction of normal traffic, so a provider incident does not turn into a self-inflicted retry storm that multiplies load exactly when the provider is weakest. - **The same idempotency key** on every attempt, so a provider that supports client keys can discard repeats. Messages that exhaust their retries move to a failed state or a dead-letter queue with the last error attached, where they can be inspected or replayed after the incident. ## Circuit breaker and failover A **circuit breaker** per provider watches recent outcomes: 1. **Closed.** Traffic flows; error rate and latency are measured over a rolling window. 2. **Open.** When errors cross a threshold, workers stop calling this provider and route new sends to the secondary. 3. **Half-open.** After a cool-down, a small trickle of traffic probes the primary; success closes the breaker, failure reopens it. Failover only works if the secondary is ready. A provider that has never carried real traffic often fails in surprising ways: sender numbers not registered for a country, templates not approved, account limits far lower than expected. Many teams therefore keep a **steady traffic split**, for example a small percentage on the secondary at all times, so it is known-good and its capacity is measured. Routing can also be per country or per message class, because provider quality varies by destination. ## The duplicate risk during failover A timed-out request may already have been sent by the primary. Failing it over sends it again through the secondary, and the two providers do not share idempotency keys. The choice depends on the message class: - **Login codes:** a duplicate is mildly confusing, a missing code blocks sign-in, so failover after a timeout is usually right. - **Order updates:** a duplicate is tolerable, so failover is usually fine. - **Marketing:** a duplicate costs money and goodwill, so it is often better to fail over only messages that were never attempted and leave ambiguous ones. ## Delivery-status callbacks An HTTP success response means the provider **accepted** the message, not that the handset or mailbox received it. Providers report later through **delivery-status callbacks** (webhooks) with states such as delivered, undelivered or bounced. Handling them well means: - correlating by the provider message id stored in the send ledger; - verifying the callback's signature before trusting it; - treating callbacks as possibly late, duplicated or out of order, and only moving a message's state forward; - feeding bounce and invalid-number results into suppression so future sends skip them. Callbacks are also the best health signal. A provider can keep returning success while a downstream carrier silently drops messages; the only symptom is a falling **delivered rate**. Alerting and breaker logic should include that rate, per provider and per destination region, not just API errors. A strong answer covers classification, bounded retries, a breaker with a warm secondary, the duplicate trade-off per message class, and callbacks as the real measure of success.

  • Why keep a small share of traffic on the secondary provider even when the primary is healthy?
    A secondary that carries no traffic is unverified: sender registrations, template approvals and account limits may be wrong, and nobody finds out until the primary fails. A steady small split proves the integration works, measures its real delivery rate and capacity, and keeps its configuration from drifting, so failover moves traffic to a path already known to work.
  • How should a notification service handle delivery-status callbacks that arrive out of order?
    Model message status as a state machine that only moves forward, for example accepted, then sent, then delivered or failed, and ignore a callback that would move the state backwards. Deduplicate callbacks by provider message id and status, verify their signatures, and store the event time from the callback rather than the arrival time.

saying these in an interview costs you the question

  • Every provider error should be retried until it succeeds.
  • An HTTP 200 from the provider proves the user received the SMS.
  • Retry immediately on HTTP 429 because it is a brief glitch.
  • A secondary provider can sit unused until the day it is needed.
  • Failing over after a timeout can never cause a duplicate.