skip to content

In a marketplace notification service, a 20-million-message marketing campaign delays login one-time codes by 40 minutes; how would you redesign sending so codes stay fast?

level: seniorimportance: must knowfreq 62%

answer

  1. what did codes and campaign share
  2. queue, workers, and one more
  3. provider throughput is a shared budget
  4. a code has a shelf life
  5. measure age, not length

basics

~20 s

Classify messages at ingest and give critical codes their own lane: separate queues, dedicated workers and a reserved slice of provider throughput. Pace campaigns instead of dumping them, expire codes that outlive their validity, and alert on oldest-message age per lane.

solid answer

~50 s

The codes waited because they shared a queue, a worker pool and a provider send budget with the campaign. I would **classify at ingest** into lanes such as critical (codes, security alerts), transactional (order updates) and bulk (marketing), each with its **own queue and worker pool**. Separate queues are not enough if every lane draws on the same provider rate limit, so I would **reserve throughput per lane**, for example with a token bucket per lane, and ideally use separate sender identities so campaign throttling or complaints cannot touch codes. Campaigns should be **paced** into the bulk lane at a controlled rate rather than enqueued all at once. Codes get a **deadline**: a code older than its validity is dropped and recorded, not sent. I would alert on the **age of the oldest message** in the critical lane, and shed bulk traffic first under pressure.

go deeper

for a junior

Recall that urgent messages such as login codes need their own path so large campaigns cannot delay them.

for a middle

Explain the three shared resources, queue, workers and provider budget, and how separate lanes with reserved capacity isolate them.

for a senior

Demonstrate the full fix: per-lane token buckets, separate sender identities, paced campaigns, code deadlines, oldest-age alerts and a predefined shedding order.

for a principal

Own the lane taxonomy and budget split as policy: who may classify a type as critical, how reserved capacity is paid for, and how campaign goals trade against code latency.

## Why the codes were late A login **one-time code** is useful for a few minutes and the user is staring at the screen waiting for it. A marketing campaign is large and can arrive whenever. When both pass through one queue, the arithmetic is unforgiving. Assume, for illustration, the SMS and email providers together accept 5,000 messages per second for this account. Draining 20,000,000 campaign messages takes 20,000,000 / 5,000 = 4,000 seconds, about 67 minutes. A code enqueued part-way through waits behind whatever is ahead of it, so a 40-minute delay is exactly what the design produces. There are three shared resources, and all three need separating: - the **queue** (codes wait behind campaign messages); - the **worker pool** (all workers are busy with campaign sends); - the **provider send budget** (rate limits and throughput are consumed by the campaign). ## Lanes by message class Classify each notification type once, in configuration, into a **lane**: | Lane | Examples | Latency goal | Under pressure | |---|---|---|---| | Critical | login codes, payment codes, security alerts | seconds | never shed | | Transactional | order shipped, refund issued | minutes | delay briefly | | Bulk | campaigns, recommendations | hours | pause or shed first | Each lane gets its **own queue** and its **own worker pool** with reserved concurrency, so bulk workers can be saturated without taking a worker away from codes. The general mechanics of priority queues, such as avoiding starvation, are a separate topic; the point here is isolation between lanes that have very different latency needs. ## Separating the provider budget Separate queues still fail if every lane calls the same provider account under one rate limit: the bulk lane can use the whole budget and the critical lane's workers get throttled. Remedies: 1. **Per-lane token buckets.** Split the known provider limit, for instance reserving 10% for critical and transactional traffic. With the illustrative 5,000 per second, 500 per second is reserved and the campaign drains at 4,500 per second, taking about 20,000,000 / 4,500 ≈ 4,444 seconds, roughly 74 minutes, which is acceptable for marketing. 2. **Borrowing.** Let the bulk lane borrow unused reserved tokens, but the critical lane can always reclaim its share immediately. 3. **Separate sender identities.** Where the provider model allows it, send codes from a different account, sender number pool or email sending domain. Campaigns generate complaints and throttling; isolating identities keeps that reputation damage away from codes. 4. **Respect back-pressure per lane.** When a provider returns HTTP 429, slow the lane that is sending most, not all lanes equally. ## Pacing campaigns A campaign should not be written into the queue as 20 million messages at once. A **campaign scheduler** releases batches at a target rate, checks lane health before each batch, and pauses automatically when the critical lane's latency rises. Pacing also spreads load on templating, preference lookups and cap counters, which are shared too. ## Deadlines on codes Every critical message should carry an **expiry**. A worker that dequeues a code older than its validity window drops it and records `EXPIRED`: sending a dead code confuses the user and wastes a paid SMS. The user will usually have requested a new one already, which is a new message with its own key. ## Checking the redesign against the incident Replay the original incident against the new design to confirm each fix does its job: - the campaign is released in paced batches, so the bulk queue never holds all 20 million messages at once; - critical workers are idle-ready because bulk load cannot occupy them; - the reserved 500 messages per second comfortably exceeds an illustrative critical peak of 200 per second; - any code that still waits past its validity is dropped and counted, which turns a silent delay into a visible metric. ## Observability and shedding - Alert on the **age of the oldest waiting message** per lane, not only on queue length; a short queue can still be stuck. - Track end-to-end latency from event creation to provider acceptance for codes, as a service-level objective. - Define the shedding order in advance: pause bulk, then delay transactional, never touch critical. - Load-test with a campaign running, because the failure only appears when lanes compete. A strong answer names all three shared resources, fixes each one, and adds pacing and deadlines so the system degrades in the right order instead of equally.

  • Why is queue length a weaker alert than oldest-message age for the critical lane?
    Critical traffic is small, so its queue can be short while every message in it is stuck because workers or the provider budget are exhausted. Oldest-message age measures what the user feels, how long the next code has waited, and it rises immediately when consumption stalls, regardless of how many messages are waiting.
  • What should the campaign scheduler do when the critical lane's latency starts rising?
    Pause or slow the release of new campaign batches automatically, and let the bulk lane lend back any borrowed provider budget. Messages already in the bulk queue can wait. Resuming should be gradual once critical latency is back under its target, so the campaign does not immediately recreate the pressure.

A hospital does not make emergency patients queue behind routine check-ups: it has a separate entrance, dedicated staff and reserved beds. Routine appointments are rescheduled first when the building fills up.

saying these in an interview costs you the question

  • Separate queues alone fix it even when all lanes share one provider limit.
  • Add more workers to the shared pool and the codes will catch up.
  • Expired codes should still be sent so the user sees something arrived.
  • Queue length is the best signal that codes are delayed.
  • The campaign team can mark its messages urgent to finish faster.