In a marketplace notification service, how do per-user frequency caps and digests stop a burst of events from flooding one user?
answer
- count per user, channel, category
- window: fixed or sliding
- parallel workers race the counter
- over the limit is not always drop
- some messages must never wait
basics
~20 sA frequency cap counts recent sends per user, channel and category and blocks sends over the limit; blocked low-value messages are dropped, deferred or buffered into a digest that is flushed later as one summary. Transactional messages such as one-time codes are exempt.
solid answer
~50 sA **frequency cap** is a sliding-window counter keyed by user, channel and category, for example 'at most three social pushes per day' as an illustrative rule. The policy stage **checks and increments atomically**, usually in an in-memory key-value store, so parallel workers cannot all squeeze under the limit. When a message is over the cap it is dropped, deferred to the next window, or added to a **digest bucket** for that user. A scheduled job flushes each bucket at the end of its window, or when it gets large, and renders one summary such as 'five new offers on items you watch'. Caps apply to marketing and social messages; **one-time codes, security alerts and order updates are exempt**, because suppressing them breaks the product. Quiet hours work the same way: non-urgent messages are deferred until the user's local morning.
go deeper
Recall that a cap limits messages per user in a time window and that capped low-value messages can be folded into a digest instead of sent one by one.
Explain the counter key, fixed versus sliding windows, why check-and-increment must be atomic, and how a digest bucket is filled and flushed.
Show operational care: exemptions defined in configuration, jittered release after quiet hours, monitoring open digest buckets, and choosing drop, defer, digest or collapse per type.
Treat caps as a shared budget across product teams: who sets the limits, how teams compete for a user's attention, and how you measure opt-out rates against engagement.
## The flood problem Some events arrive in bursts: twenty price drops on a wishlist during a sale, a seller answering ten questions in a row, a popular listing collecting likes. Turning each event into its own push trains users to disable notifications or uninstall the app, and each SMS costs money. The notification service therefore needs rules that limit **how many** messages a user gets and **when**, without losing the messages that matter. ## Frequency caps A **frequency cap** is a limit such as 'no more than N messages of category C on channel X per user per time window'. The numbers are product decisions; the mechanism is the same: - **Key.** Usually `user + channel + category`, sometimes also a global per-user cap across all categories. - **Window.** A fixed window (calendar day in the user's time zone) is simple; a **sliding window** (last 24 hours) avoids a burst straddling midnight. A sliding window can be stored as a sorted set of send timestamps or approximated with two fixed-window counters. - **Atomic check-and-increment.** Many workers process messages in parallel. If each reads 'two sent, limit three' and then sends, all of them pass. The check and the increment must be one atomic operation in the counter store. - **When to count.** Counting at decision time is simpler and slightly conservative; counting only after a successful send is more accurate but needs care to avoid the same race. ## What happens to a capped message | Outcome | Suits | Cost | |---|---|---| | Drop | low-value marketing | user never sees it | | Defer | time-insensitive updates | arrives late, may be stale | | Digest | many similar small events | one summary instead of many | | Collapse | repeated updates to one object | only the newest state is shown | **Collapsing** deserves a note: when an order moves from 'packed' to 'shipped' within minutes, the second update can replace the first rather than add to the count, using a collapse key such as the order id. ## Digests A **digest** turns many small events into one message. 1. Instead of sending, the policy stage appends the event to a **digest bucket** keyed by user and category, and records when the bucket was opened. 2. A scheduler flushes buckets whose window has ended, for example hourly or once a day at a local time the user chose, or earlier when a bucket passes a size threshold. 3. The flush renders a summary template ('you have 7 new offers'), sends it through the normal pipeline with its own idempotency key, and clears the bucket. 4. Items that became irrelevant in the meantime, such as an offer that expired, are dropped at flush time. The flush job must be safe to run twice: a bucket id plus window end makes a good idempotency key for the summary message. ## Quiet hours and exemptions **Quiet hours** are evaluated in the recipient's local time zone, which enrichment provides. Non-urgent messages that fall inside them are deferred to the end of the quiet period, and often merged into a digest. Some classes must **bypass** caps and quiet hours: - one-time login or payment codes, which the user is waiting for right now; - security alerts such as a new-device sign-in; - transactional updates the user explicitly relies on, such as delivery arriving today. The exemption belongs to the **notification type's configuration**, not to the calling team's code, so nobody can mark a campaign as 'urgent' to dodge the cap without review. ## Pitfalls - A cap that counts all channels together can let a flood of pushes block an important email. - Deferred messages released together at 08:00 local time create a synchronised spike per time zone; spread releases over a few minutes. - A digest that is never flushed because the scheduler missed its window loses messages quietly; monitor the age of the oldest open bucket. - Counting suppressed messages against the cap makes the cap stricter than intended. In an interview, the key points are: the cap is an atomic windowed counter on a well-chosen key, the over-cap path is a deliberate choice (drop, defer, digest, collapse), and critical transactional messages are exempt by configuration.
- How would you avoid a spike when thousands of deferred messages are released at the end of quiet hours?Add jitter to the release time, for example spreading releases over several minutes after the quiet period ends, and let the channel queue's rate limit smooth what remains. Because quiet hours end at a local time, the spike repeats per time zone, so the spread should be applied to every zone's release rather than only the largest.
- What is a collapse key and when is it better than a digest?A collapse key marks messages that describe the same object, so a newer message replaces an older unsent one instead of adding to it. It suits status updates on one order, where only the latest state matters. A digest suits many different items, such as offers on several listings, where the user wants a summary of all of them.
saying these in an interview costs you the question
- Frequency caps should apply to one-time login codes too.
- Read the counter, then increment it later; races are too rare to matter.
- Capped messages should be retried on another channel that has spare quota.
- A digest can be flushed without an idempotency key because it runs on a schedule.
- Quiet hours can be evaluated in the server's time zone.