skip to content

In a fanout-on-write home timeline, what is gained and lost by skipping inbox writes for followers inactive for several weeks?

level: seniorimportance: nice to knowfreq 30%

answer

  1. work nobody reads
  2. cheap per-follower membership check
  3. missing inbox on return
  4. mark active before rebuilding
  5. mass return after a campaign

basics

~20 s

Skipping dormant followers removes a large share of fanout writes and lets their inboxes be evicted from memory. The cost is that a returning user has no inbox, so the first timeline load must be rebuilt by pulling from followees.

solid answer

~50 s

In many social networks a large share of follower edges point at accounts that rarely log in, so pushing into their inboxes is work nobody reads. Fanout workers check a cheap **activity marker**, such as a set of users seen in the last few weeks, and skip anyone not in it; their inboxes can also be expired to free memory. When a dormant user returns, their inbox is missing or stale, so the service first marks them active so normal fanout resumes, then **rebuilds the inbox** with a fanout-on-read pass over their followees' recent posts, deduplicating by post ID, serves that first page and writes the result back. The costs are a slower first load, a race where a post published between the rebuild and resumed fanout is missed if the user is marked active too late, and rebuild storms when many dormant users return at once, for example after a re-engagement campaign.

go deeper

for a junior

Recall that writing into inboxes nobody opens is wasted work, and that a returning user needs their timeline built another way.

for a middle

Explain the per-follower activity check during fanout and the fanout-on-read rebuild on return.

for a senior

Cover the rebuild ordering race and rebuild storms after mass re-engagement, with concrete mitigations for each.

for a principal

Choose the inactivity window from the return-visit distribution and weigh write savings against rebuild cost and first-load latency.

## Why skip anyone at all In **fanout-on-write**, every post is written into the precomputed inbox of every follower. That work only pays off if the follower reads the inbox. Follower lists accumulate dormant accounts over years: people who signed up, followed some accounts, and stopped visiting. Writing into their inboxes is pure cost: - **CPU and write throughput** spent by fanout workers; - **memory** held by inboxes that are never read; - **longer fanout jobs**, which raise delivery lag for active followers. Illustratively, if 40% of follower edges point at accounts inactive for over 30 days, skipping them removes about 40% of all fanout writes in one change. ## How the skip works 1. Maintain an **activity marker** per user, updated when they open the app, for example a last-seen timestamp or membership in an "active in the last N days" set. 2. Fanout workers check the marker for each follower batch and write only to active users. 3. Inboxes of users who fall out of the active set are **expired**, freeing memory. The check must be cheap because it runs for every follower of every post. A compact membership structure, such as a bitmap indexed by user ID, keeps the per-follower check fast. A bloom filter would also work if a small rate of false positives, meaning a few wasted writes, is acceptable; it never causes a skipped active user, because bloom filters have no false negatives, provided every user is added the moment they become active. Since a bloom filter cannot delete entries, it is rebuilt periodically to drop users who went dormant. ## What happens when the user returns The returning user's inbox is missing or out of date, so the first timeline request takes a **rebuild path**: 1. Mark the user active immediately, so new posts start being pushed to them. 2. Run a **fanout-on-read** pass: fetch recent posts from each followee's outbox and merge them by time. 3. Serve the first page from that merge. 4. Write the merged result back as the user's inbox, then continue normally. Ordering matters. If the user is marked active **after** the rebuild reads the outboxes, a post published in between is neither in the rebuild nor pushed to them. Marking active first, then rebuilding, and deduplicating by post ID closes that gap. ## Costs and failure modes | Gain | Loss | |---|---| | Fewer fanout writes | Slower first load for returning users | | Less inbox memory | A rebuild path that must be built and tested | | Lower fanout lag for active users | A race between rebuild and resumed fanout | | Cleaner capacity model | Rebuild storms when many users return together | The **rebuild storm** deserves attention. A mass notification or a re-engagement email can bring back millions of dormant users within minutes, each triggering an expensive pull merge. Mitigations include: - serving a lightweight pull-based page first and rebuilding the full inbox in the background; - rate-limiting rebuilds and queueing them; - pre-warming inboxes for the users targeted before a campaign is sent. ## Choosing the inactivity window - A **short window** saves more writes but sends more users down the slow rebuild path. - A **long window** makes rebuilds rare but saves less. - The right window comes from the return-visit distribution: pick a point after which returns are rare enough that the rebuild cost is smaller than the fanout saved. ## Where it fits This is a refinement, not a core design element. It matters once fanout volume is large enough that dormant followers show up as a real cost line, and it combines naturally with the hybrid model, which already has a pull path that the rebuild can reuse.

  • Why should a returning user be marked active before their inbox is rebuilt rather than after?
    The rebuild reads followees' outboxes at one moment. A post published after that read but before the user is marked active is neither in the rebuild nor pushed by fanout, so it silently goes missing. Marking active first means fanout starts pushing immediately, and deduplicating by post ID removes any overlap with the rebuild.
  • Why can a bloom filter serve as the activity check without ever skipping an active user?
    A bloom filter can report an item as present when it is not, but never reports an inserted item as absent. As long as users are added the moment they become active, its errors only cause a few extra writes to inactive users, never a missed write to an active one, which is the safe direction for this check.

It is like a postal service holding mail for a household that has been away for months instead of delivering daily, then handing over a bundle when they return.

saying these in an interview costs you the question

  • Inactive users can be skipped with no read-side change.
  • A returning user's first page is as fast as anyone else's.
  • Marking the user active after the rebuild is safe.
  • A bloom filter for active users can skip real active users.
  • The shortest possible inactivity window always saves the most overall.