skip to content

For a chat platform whose largest groups have 100,000 members, how would you decide how much presence information to fan out?

level: principalimportance: nice to knowfreq 30%

answer

  1. changes times watchers
  2. most dots are never seen
  3. subscribe to what's on screen
  4. a count, not a roster
  5. thresholds as configuration

basics

~20 s

Presence fanout grows as status changes times watchers, which is unaffordable in huge groups. Send live presence only for what is on screen and for small circles like contacts, batch and debounce changes, and replace per-member presence with an approximate online count in large groups.

solid answer

~50 s

Start with the arithmetic. As an illustration, suppose 20,000 of 100,000 members are online and each flips status twice an hour. Pushing each flip to every online member means 40,000 x 20,000 = **800 million** deliveries an hour, about 222,000 per second, for one group, to show dots nobody scrolls to. So I would treat presence as a **fidelity budget**. First, **subscribe on view**: a client asks for live presence only for the people currently on screen, and drops the subscription when they scroll away. Second, give **full fidelity to small circles**: 1:1 chats, contacts and small groups. Third, show large groups an **approximate online count** that refreshes every tens of seconds. Fourth, **batch and debounce**: send deltas every few seconds and suppress flapping. Fifth, serve **last seen on demand** instead of pushing it. The decision is a product-cost trade: I would set group-size thresholds from measured delivery cost and from how often users actually look at presence.

go deeper

for a junior

Remember that showing who is online can cost more than the messages themselves, because each status change may go to many people.

for a middle

Explain the formula of changes times watchers, and how subscribing only to users on screen shrinks the watcher count.

for a senior

Run the numbers for a large group, and propose batching, debouncing, approximate counts and a separate presence tier that can degrade under load.

for a principal

Own the fidelity budget: tie group-size thresholds to measured cost and real usage, and agree with product on which surfaces lose per-person presence.

## Why presence is a fanout problem Detecting presence is cheap: one heartbeat-refreshed entry per online device. **Distributing** it is the expensive part. Each status change has to reach everyone currently *watching* that user, and in a big group that can be everyone. The cost grows as: **status changes per hour x watchers per change** Messages have the same shape, but people *want* every message. Most presence updates in a huge group are never looked at. ## The arithmetic Illustrative assumptions for one 100,000-member group: - 20,000 members online at a time; - each online member changes status about twice an hour (app backgrounded, reconnect, and so on); - naive design: every change is pushed to every online member. Changes per hour: 20,000 x 2 = **40,000**. Deliveries per hour: 40,000 x 20,000 = **800,000,000**, about **222,000 per second** for a single group. Add many such groups, plus users who belong to several, and presence can easily outweigh message traffic. It also burns client battery and bandwidth. ## The levers | Lever | What it does | Cost it cuts | What users lose | |---|---|---|---| | **Subscribe on view** | live presence only for users on screen | watchers per change | nothing visible | | **Small-circle full fidelity** | live dots in 1:1, contacts and small groups | limits live fanout to small audiences | none where it matters | | **Approximate count** | "about 4.2k online" in big groups | per-member fanout | per-person dots | | **Batch deltas** | send changes every few seconds | messages per change | a few seconds of freshness | | **Debounce flapping** | hold offline for a grace period | spurious changes | slightly late offline | | **Last seen on demand** | fetch when a profile opens | background pushes | nothing visible | **Subscribe on view** usually saves the most. A client showing twenty member avatars subscribes to twenty presence channels, not 100,000. When the user scrolls, it drops old subscriptions and adds new ones. The server keeps a *watcher set* for each user, and it stays small because screens are small. ## Making the decision There is no single right answer. It is a budget. A reasonable process: 1. **Measure** how often each surface actually shows presence (1:1 header, member list, group header) and how often users open it. 2. **Price** each surface's fanout using the formula above and real online and flip rates. 3. **Set thresholds**: for example, full live presence up to a few hundred members, subscribe-on-view in larger groups, and an approximate count only above some size. Treat the thresholds as tunable configuration, not constants. 4. **Degrade under load**: when the presence tier is under pressure, lengthen batch intervals or pause large-group presence first, before any message traffic suffers. 5. **Watch the cost**: alert on presence deliveries per second and on watcher-set sizes, and compare them with message deliveries. ## Trade-offs to state explicitly - **Freshness against cost**: batching and debouncing trade seconds of staleness for large savings, which is acceptable for a dot and not for a message. - **Consistency of experience**: users may notice that big groups show no per-person dots. Product and design have to accept that, and it is a conversation a lead owns. - **Privacy interacts with fanout**: users who hide their presence must be filtered out *before* fanout, which also shrinks it. - **Isolation**: keep presence on its own tier and budget, so a presence storm, for example after a mass reconnect, cannot slow message delivery. ## What stays out of scope How the client draws the dots, and how the gateway fleet survives a reconnect storm, are separate designs. This decision is about **how much status information the server agrees to distribute, and to whom**.

  • How does subscribe-on-view work on the server side?
    The client sends the list of user IDs currently visible, and the server adds the client to each of those users' watcher sets. When the view changes, the client sends additions and removals. A status change fans out only to that user's watcher set, which stays small because a screen shows only a few dozen people. Subscriptions expire with the connection.
  • Why should presence run on a separate tier from message delivery?
    Presence traffic is bursty and loses little value when delayed, while messages must not be delayed. On a shared tier, a presence storm, such as many clients flipping state after a mass reconnect, can slow message delivery. Separating them lets presence be batched, throttled or paused under load without affecting messages.

saying these in an interview costs you the question

  • Push every presence change to every member of every group.
  • Presence is cheap because each update is only a few bytes.
  • Make presence fidelity identical for all group sizes.
  • Poll every member's presence every few seconds from each client.
  • Hard-code the large-group threshold and never revisit it.