In a marketplace notification service, what stages does one event pass through before a push, email or SMS reaches a provider?
answer
- one event, several channels
- who is the user, what did they allow
- check before you build the message
- per-channel template and locale
- queue in front of the provider call
basics
~20 sAn event is enriched with recipient data, filtered by preferences, opt-outs, quiet hours and frequency caps, routed to allowed channels, rendered from a localised template, deduplicated by an idempotency key, then queued per channel for workers that call providers and record status.
solid answer
~40 sI would describe it as a pipeline of small stages behind one entry point. An `OrderShipped`-style event arrives on a queue; the service **enriches** it with the recipient's contact points, locale and time zone; it runs **policy checks** (opt-outs, per-category preferences, quiet hours, frequency caps); it **routes** to the channels that survive; it **renders** a localised template per channel; it stamps an **idempotency key** and writes a send record; then it enqueues one message per channel. Channel workers call the push, email or SMS provider with retries and record the provider's message id, and **delivery-status callbacks** update that record later. Centralising this means every feature team gets consent, caps and deduplication for free instead of reimplementing them.
go deeper
Recall the stages in order: enrich, policy checks, route, render, deduplicate, enqueue, send and track. Be ready to say why consent is checked before the message is built.
Explain what data each stage needs, why the provider call sits behind a per-channel queue, and what the send record holds and who reads it.
Show where the pipeline fails in production: stale preferences while queued, provider outages isolated per channel, and campaigns crowding urgent messages without separate lanes.
Frame the central service as a platform contract: which rules are mandatory for every sender, how new notification types are onboarded, and who owns the routing rules.
## Why one central pipeline A marketplace sends many kinds of messages: order updates, one-time login codes, price-drop alerts, marketing campaigns. If every feature called an email or SMS provider directly, each team would have to reimplement consent checks, rate limits, retries and duplicate protection, and one team forgetting an opt-out check becomes a legal and trust problem. A **notification service** is the single component that turns a business event ('order 812 shipped') into zero or more delivered messages on the right **channels** (mobile push, email, SMS), under one set of rules. ## The stages in order 1. **Ingest.** Producers publish a domain event or call a send API with a notification type, recipient id and payload. The service acknowledges quickly and does the rest asynchronously. 2. **Enrich.** Look up the recipient: registered push tokens, verified email address, verified phone number, preferred language and time zone. 3. **Policy checks.** Apply global opt-outs and per-category **preferences** ('no marketing by SMS'), **quiet hours** in the user's local time, and **frequency caps**. A message that fails here is suppressed, deferred or folded into a digest, and the reason is recorded. 4. **Route.** Pick channels from the notification type's routing rule and the checks above: a login code may be SMS-first, an order update push-first with email as a fallback. 5. **Render.** Fill a per-channel, per-locale **template** with the payload. Push needs a short title and body, email needs a subject plus HTML and text parts, SMS needs a short plain-text body. 6. **Deduplicate.** Derive an **idempotency key** from the event and recipient and write a send record, so a redelivered event does not produce a second message. 7. **Enqueue per channel.** Put each rendered message on a channel queue, ideally split by priority so a one-time code never waits behind a campaign. 8. **Send and track.** Channel workers call the provider under its rate limit, retry transient failures, store the provider's message id and later apply **delivery-status callbacks** (delivered, bounced, failed). ## What each stage owns | Stage | Main input | Output or side effect | |---|---|---| | Enrich | recipient id | contact points, locale, time zone | | Policy checks | preferences, cap counters | allow, defer, digest or suppress | | Route | notification type rule | list of channels | | Render | template id, locale, payload | channel-specific message body | | Deduplicate | event id, recipient, channel | send record keyed for idempotency | | Send | rendered message | provider message id, status updates | ## Why the queues sit where they do The expensive and unreliable part is the provider call: providers are remote, rate-limited and occasionally slow or down. Putting a **queue in front of the channel workers** means: - a burst of events is absorbed instead of timing out the producer; - each channel scales and fails independently, so an SMS outage does not stop email; - retries happen in the worker without the producer knowing or caring; - separate queues per priority keep urgent messages moving during a campaign. Policy checks run **before** rendering and sending for a simple reason: a suppressed message should cost nothing, and the decision should use the user's current preferences at send time. Some systems re-check opt-outs just before the provider call as well, because a message can sit in a queue for minutes while the user unsubscribes. ## What the send record is for The send record written at the deduplicate stage is the service's memory of each message. It typically holds the idempotency key, the recipient, the channel, the template, the state (`PENDING`, `SENT`, `DELIVERED`, `FAILED`, `SUPPRESSED`) and the provider's message id. It answers support questions ('did the customer get the code?'), feeds frequency caps and digests, lets callbacks find the message they refer to, and blocks duplicate sends. ## Common mistakes - Letting feature code call providers directly 'just this once', which bypasses consent and caps. - Checking preferences only in the client app, even though email and SMS never pass through the app. - Rendering before policy checks, which wastes work and can leak a suppressed message into logs or previews. - Treating the provider's success response as delivery; it only means the provider accepted the message. - Using one shared queue for every notification type, so a campaign delays login codes. A junior engineer is not expected to design every stage in depth, but should be able to list them in a sensible order and explain why consent checks come early and the provider call sits behind a queue.
- Why might a notification service re-check opt-outs right before calling the provider, even after checking at routing time?Messages can wait in a channel queue for seconds or minutes, and a user may unsubscribe in that gap. A cheap re-check against the current preference record just before the provider call closes that window. It matters most for marketing, where sending after an opt-out breaks consent rules, and least for one-time codes, which are sent within seconds and are exempt from marketing preferences anyway.
- Where should localisation happen in the pipeline, and what does it need from enrichment?Localisation happens at the render stage. It needs the recipient's preferred language, and often their time zone and currency, so enrichment must load those before rendering. Templates are stored per notification type, channel and locale, with a fallback locale when a translation is missing, so a missing translation degrades to a default language instead of failing the send.
It works like a post room: incoming requests are checked against a do-not-mail list, addressed and packed in the right format, and handed to the right courier. No department is allowed to post letters itself.
saying these in an interview costs you the question
- Each feature team should call the email or SMS provider directly for speed.
- Opt-outs can be enforced by hiding messages in the mobile app.
- A success response from the provider means the user received the message.
- Preference checks can run after the message has been sent to the provider.
- One shared queue for all notification types is fine at any scale.