skip to content

In a fanout-on-write home timeline, why does each user's precomputed inbox hold only post IDs capped at the most recent N entries?

level: middleimportance: should knowfreq 55%

answer

  1. what a read actually needs
  2. bytes per entry times users
  3. hydration by batched lookup
  4. edits without rewriting inboxes
  5. insert then trim

basics

~20 s

IDs keep each entry to a few bytes, so millions of inboxes fit in fast memory and edits need no inbox rewrite. Capping at N bounds storage because almost every read wants only the newest page; deeper history is rebuilt on demand.

solid answer

~50 s

An inbox entry only needs to say *which* post belongs in the timeline, so it stores a post ID, often with the author ID and a timestamp for filtering: roughly 16 bytes. At read time the service **hydrates** the page with one batched lookup against the post store. That makes inboxes small enough to keep in an in-memory store, and an edited post shows up correctly without touching any inbox. Inboxes are **trimmed to the last N entries** on each insert, because readers mostly look at the first page or two. Illustratively, 500 entries of 16 bytes for 200 million users is about 1.6 TB before replication, whereas storing 1 KB post bodies instead would need about 100 TB. Scrolling past N falls back to a slower path that pulls older posts from followees' outboxes.

go deeper

for a junior

Recall that the inbox holds references to posts, not the posts themselves, and only keeps the newest few hundred.

for a middle

Explain insert-then-trim, hydration with a batched lookup, and do the bytes-per-entry times user-count arithmetic out loud.

for a senior

Cover what happens past N, short pages caused by filtering, and why IDs make edits and deletes cheaper than copies.

for a principal

Treat N as a cost lever: memory for every user against the fraction of sessions that scroll deep, and choose the fallback path deliberately.

## What a precomputed inbox is In **fanout-on-write**, every user has an **inbox**: an ordered list of the posts that belong in their home timeline, maintained by background fanout workers. The inbox is read on every timeline request, so it must be fast, and it exists for every user, so it must be small. Those two pressures shape both what an entry holds and how many entries are kept. ## Store IDs, not posts A typical inbox entry holds: - the **post ID**, ideally one that sorts by creation time; - optionally the **author ID**, so the read path can filter unfollowed or blocked authors without loading the post; - optionally a **timestamp or score**, if the ID does not already encode ordering. That is on the order of 16-24 bytes. The post body, media references, counters and author profile live in their own stores and are fetched at read time by **hydration**: one batched multi-key lookup for the page of IDs, usually served from a cache. Reasons to keep bodies out of the inbox: - **Size.** A post body with metadata is easily 1 KB, roughly 50 times larger than an ID entry. - **Edits and deletes.** If inboxes held copies, editing a post would require rewriting it in every follower's inbox. With IDs, an edit changes one record and every reader sees it. - **Counters change constantly.** Like and reply counts would be stale the moment they were copied. ## Cap the inbox at N entries Fanout workers typically do an insert followed by a trim: append the new ID, then drop anything beyond the newest N. N is chosen to cover what readers actually consume: most sessions look at the first page or two, so a few hundred to about a thousand entries is a common illustrative range. Illustrative storage arithmetic, assuming 16 bytes per entry and 200 million users with inboxes: | Choice | Per user | All users | |---|---|---| | 500 ID entries | 500 x 16 B = 8 KB | 8 KB x 200M = 1.6 TB | | 500 full posts at 1 KB | 500 KB | 500 KB x 200M = 100 TB | | Unbounded IDs, 20,000 entries | 320 KB | 64 TB | With three replicas the first row becomes 4.8 TB: large but feasible in memory. The other rows are not. Truncation also keeps each insert cheap and each range read small, and it bounds the damage of a stale entry: anything that slipped past cleanup eventually ages out. ## What happens beyond N A reader who scrolls past the last stored entry needs another source. Common options: 1. **Fall back to pull** for older pages: fetch followees' outboxes older than the last inbox entry's timestamp and merge them. This is slower but rare. 2. **Stop the timeline** at a depth limit, which many products accept. 3. **Rebuild on demand** if the inbox is empty, for example for a user returning after a long absence. Paging through the inbox uses a cursor such as "posts older than this ID", which stays stable as new entries arrive at the head; how the client renders and requests those pages is a client-side concern. ## Trade-offs to state - **Hydration is a second round trip.** Batching it (one multi-get for the page) keeps it cheap; per-ID lookups would not be. - **Filtered entries shrink pages.** If some IDs resolve to deleted or hidden posts, the page comes back short, so the read path over-fetches slightly. - **N is a product decision too.** Raising it costs memory for every user, including those who never scroll that far. - **Ordering needs sortable IDs.** Time-ordered IDs let the inbox stay sorted by insertion and let merges compare IDs directly.

  • Why does a hydrated page from a truncated inbox sometimes contain fewer posts than requested?
    Some IDs may point to posts that were deleted, hidden, or written by authors the reader has since unfollowed or blocked, and the read path filters them during hydration. To still return a full page, the service reads a few more IDs than the page size and trims after filtering.
  • Why store the author ID alongside the post ID in each inbox entry?
    It lets the read path drop entries from unfollowed or blocked authors by checking the entry itself, before any post lookup. It also makes cleanup cheap: removing one author's posts from one inbox is a scan of small entries rather than a hydration of each post.

saying these in an interview costs you the question

  • Store full post bodies in each inbox to save a lookup.
  • Keep every post a user was ever sent in their inbox.
  • Editing a post requires rewriting it in every follower's inbox.
  • Readers can never see posts older than the inbox cap.