skip to content

In a cloud-drive service where one shared team folder has 50,000 members, would you copy its changes into every member's journal or keep a per-folder journal?

level: principalimportance: should knowfreq 35%

answer

  1. write amplification versus read merging
  2. one namespace per shared folder
  3. cursor as a map of positions
  4. membership events in the user's journal
  5. check access on every fetch

basics

~20 s

Keep one journal per shared folder and give each device a cursor per folder it can see. Copying each change into 50,000 personal journals makes every write a huge fan-out and makes joins and permission changes expensive.

solid answer

~50 s

Model each user's own tree and each shared folder as separate **namespaces**, each with its own journal; a shared folder is mounted into members' trees. A device's cursor becomes a map of namespace to position, and `list_changes` merges the namespaces the user can see. Writes cost one append however many members there are. Fan-out-on-write would give each user one simple journal, but 1,000 edits an hour into a 50,000-member folder is 50 million journal writes an hour, and a new member needs the folder's history backfilled. Membership changes become small events in the user's own journal: a **mount** event makes the device list the folder from a starting cursor; an **unmount** event makes it drop the folder, and the server enforces access on every fetch, not at fan-out time, so revocation is immediate. The cost of per-folder journals is on the read side: a user in hundreds of folders has hundreds of positions to track and watch, which some designs soften with fan-out for small folders only.

go deeper

for a junior

Recall that a shared folder's changes must reach every member, and that copying them to each member is one of two ways to do it.

for a middle

Explain namespaces, mounting, and why a per-folder design turns the device cursor into a map of positions.

for a senior

Show how join, revocation and view-only downgrade propagate as events, and why access is enforced on every fetch rather than at fan-out time.

for a principal

Decide from data: share-size distribution, edit rate and membership churn; quantify write amplification, and justify a hybrid only when both folder profiles are common.

## The decision A shared team folder in a cloud-drive service has many members, each with several devices. Every change to that folder must reach every member's devices through the change-journal mechanism. The design question is where the folder's changes are **recorded**: copied into each member's personal journal (**fan-out on write**) or kept once in a journal owned by the folder (**per-namespace journal**, read on demand). There is no universally right answer; the choice depends on folder sizes, membership churn and read patterns. ## Two models | Aspect | Fan-out into each member's journal | One journal per shared folder | |---|---|---| | Cost of one edit | One append per member | One append | | Device cursor | One position | One position per visible namespace | | Read path | Simple single-journal read | Merge across namespaces | | New member joins | Backfill history into their journal | Mount event plus snapshot of the folder | | Member removed | Stop fanning out; old entries remain | Unmount event; access checked on fetch | | Huge folders | Write amplification | No extra cost | ## The arithmetic Assume a folder with 50,000 members and 1,000 edits per hour. - Fan-out on write: 1,000 x 50,000 = **50,000,000** journal appends per hour, each an index write with retention cost. - Per-folder journal: **1,000** appends per hour, regardless of membership. The gap grows linearly with membership, and a single bulk operation, such as moving 10,000 files, multiplies it again. This is why large shared folders push designs toward per-namespace journals. ## How per-namespace sync works - Every user has a **root namespace** for personal files. Every shared folder is its own namespace, **mounted** at a path in each member's tree. - The device's cursor is an opaque token encoding a **map of namespace to position**. - `list_changes` reads each visible namespace's journal after its position and merges the results, translating namespace paths into the user's mount paths. - Long polling watches several namespaces. To keep that cheap, the commit path can publish a tiny per-user wake-up to members, which is a fan-out of a ping rather than of stored entries, or the notifier can subscribe to each namespace's signal. ## Propagating membership and permission changes Membership is recorded as events in the **user's own** journal, not the folder's: 1. **Join**: a mount event appears in the new member's root journal. Each device lists the folder at a snapshot cursor, adds that namespace and position to its cursor map, and syncs incrementally from there. Very large folders may be synced selectively. 2. **Revocation**: an unmount event appears. Devices stop syncing the folder and remove or detach their local copies. Crucially, the server checks access on **every** fetch and upload, so a device that has not yet processed the event is still refused immediately. 3. **Downgrade to view-only**: an event tells devices to mark the folder read-only; the server independently rejects uploads from that member. In a fan-out model, revocation is weaker by default: entries already copied into the member's journal remain readable unless the read path re-checks access, which reintroduces the per-fetch check anyway. ## Costs of the per-namespace model - **Many cursors.** A user in several hundred shared folders carries several hundred positions; cursor tokens grow and merge reads touch many journals. - **Cross-namespace moves.** Moving a file from a personal folder into a shared one becomes a delete in one namespace and a create in another, and the client must understand it as a move. - **Retention per namespace.** Each journal compacts independently, so a device can hold a valid position for one folder and an expired one for another, requiring a reset for just that namespace. ## A hybrid Some designs combine the two, similar to how home-timeline systems mix push and pull: - fan out on write for small, stable shares where per-user reads dominate; - keep per-folder journals for large or high-churn shares; - choose by member count and write rate, and migrate a folder between modes as it grows. The hybrid adds operational complexity, so it earns its place only when both kinds of folder are common. ## How a lead should decide Look at the distribution of shared-folder sizes, edit rates, how often membership changes, and how many shares a typical user has. Per-namespace journals win decisively when large shares exist; fan-out stays attractive only when every share is small. Whichever you pick, enforce permissions on the read path so revocation never depends on journal contents.

  • What does a newly added member's device do first?
    It sees a mount event in the member's own journal, requests a snapshot of the shared folder with a starting position, lists and downloads what it needs, and adds that namespace and position to its cursor map. From then on the folder syncs incrementally like any other namespace. Very large folders may be synced selectively to save space.
  • When would you still choose fan-out on write?
    When shares are small and stable, and per-user reads dominate: each device then reads one journal with one position, and long polling watches a single stream. The write amplification stays modest. Many designs use fan-out below a member-count threshold and per-folder journals above it.
  • How does long polling stay cheap when a user sees hundreds of namespaces?
    Instead of the notifier subscribing to hundreds of journals per device, the commit path can publish a tiny wake-up to each member's personal channel. That is still a fan-out, but of a ping, not stored entries. The device then fetches using its full cursor map, so correctness still comes from the journals.

saying these in an interview costs you the question

  • Always copy shared-folder changes into every member's journal, whatever the size
  • Check permissions only when changes are fanned out, not on each fetch
  • A revoked member's devices may finish syncing their backlog first
  • One global position number works across all shared folders
  • Membership changes need no journal entries of their own