In a cloud-drive sync service, why does the long-poll change-notification call return only a changed flag instead of the changes themselves?
answer
- hint versus source of truth
- idle connections should be cheap
- one path delivers the entries
- check the cursor before waiting
- jitter and backoff on reconnect
basics
~20 sSplitting notification from fetch keeps the long-poll tier cheap: it holds many idle connections and only says something changed, while the device reads the actual entries from its cursor through the normal paginated, authorized change-list call.
solid answer
~50 sThe device sends its cursor to a long-poll endpoint; the server holds the request open for up to a timeout (say about a minute, with random jitter) and answers `changes: true` as soon as the journal has moved past that cursor, or `changes: false` when the timeout expires. The device then calls the ordinary `list_changes(cursor)` and opens a new long poll. The notification servers hold a huge number of mostly idle connections, so they should keep almost no state and never read file metadata. The fetch path already handles pagination, permission checks and retries, so there is a single code path that delivers changes. And a lost notification is harmless: the cursor, not the notification, is the source of truth, so the next poll or fetch still returns the change. The server can also attach a `backoff` hint to shed load.
code
http · 9 linesPOST /sync/longpoll HTTP/1.1
Content-Type: application/json
{"cursor": "opaque-token-for-48215", "timeout": 60}
HTTP/1.1 200 OK
Content-Type: application/json
{"changes": true, "backoff": 0}go deeper
Know the difference between short polling, long polling and a persistent stream, and that the long-poll answer here just says whether to fetch.
Walk through the flow step by step, including the cursor comparison on arrival that closes the race, and explain why the fetch is a separate call.
Discuss operating the notifier: connection-count capacity, jittered timeouts, backoff hints after a restart, and a slow safety-net fetch when signals are lost.
Argue the split as an architectural boundary: a stateless wake-up tier scaled on connections, an authoritative read tier scaled on queries, and correctness owned only by the cursor.
## Getting changes to a device quickly A change journal with per-device cursors tells a device **what** changed once it asks. The remaining question is **when** to ask. A cloud-drive client should pick up an edit made on another device within seconds, without hammering the server when nothing is happening, which is most of the time. | Technique | How it works | Trade-off | |---|---|---| | Short polling | Ask every N seconds | Latency up to N; most calls return nothing | | Long polling | Request stays open until a change or a timeout | Plain HTTP; one held connection per device | | WebSocket or Server-Sent Events | Persistent stream from server | Lowest latency; long-lived connections need draining on deploys | Long polling is a common default because it is ordinary HTTP: it passes through proxies, balances like any request, and needs nothing special on reconnect. ## The notify-then-fetch flow 1. The device calls the notification endpoint with its current cursor and a timeout. 2. On arrival, the server first checks whether the journal is already past that cursor. If so, it returns `changes: true` immediately; this closes the race where a change committed just before the request arrived. 3. Otherwise the server subscribes to a lightweight **change signal** for that account, published whenever a commit appends to its journal, and holds the request. 4. It answers `changes: true` when a signal arrives, or `changes: false` when the timeout expires. 5. On `true`, the device calls `list_changes(cursor)`, applies the entries, saves the new cursor, and opens a new long poll. On `false`, it simply opens a new long poll. ## Why the response is only a flag - **Cheap connection holding.** The notification tier may hold one connection for every online device. If each held request carried a result set, those servers would need metadata access, memory for payloads, and pagination logic. A flag keeps them small, and they can scale separately from the metadata tier. - **One delivery path.** Changes reach the device only through `list_changes`, which already enforces permissions, pages large result sets, and handles retries. Duplicating that inside the notifier would create two paths that can disagree. - **Loss tolerance.** Signals may be dropped: a notifier restarts, a connection breaks, a pub/sub message is missed. Because the device's cursor has not moved, the next fetch still returns the change. The notification is a **hint**; the cursor is the truth. - **Coalescing.** Fifty edits in one second produce one wake-up and one fetch of fifty entries, not fifty responses. ## Load and failure handling - **Jittered timeouts.** If every client used exactly 60 seconds, a mass reconnect after a notifier deploy would re-align all their timeouts into a synchronized wave. Randomizing the timeout spreads later cycles out. - **Backoff hints.** The server can return a `backoff` value in seconds, telling clients to wait before polling again when it is overloaded. - **Safety-net fetch.** Some clients also run a slow periodic `list_changes` regardless of notifications, bounding the worst-case delay if signals are lost entirely. - **Connection limits.** Idle held requests still consume sockets and memory per connection, so notifier capacity is planned in concurrent connections, not requests per second. ## Common mistakes - Checking only for signals that arrive after the request starts, instead of comparing the cursor first, which misses changes committed during the gap between fetch and poll. - Treating the notification as the data, which turns any dropped message into a missed change. - Using identical fixed timeouts for all clients. ## Summary Notify-then-fetch separates a cheap, stateless wake-up channel from the authoritative, cursor-driven read. Latency comes from the long poll; correctness comes from the cursor.
- Would a WebSocket be a better channel than long polling here?It can be: a persistent stream lowers latency and saves reconnect overhead. But long-lived connections must be drained on every deploy and are harder to balance, while long polling is plain HTTP that any proxy handles. Either way, the design stays notify-then-fetch: the stream carries only a wake-up, and entries still come from the cursor-based fetch.
- How does the notification server learn that the journal moved?The commit path publishes a small signal keyed by account or namespace, for example on a pub/sub message bus, whenever it appends an entry. The notifier subscribes for the accounts it holds requests for. On each new request it also compares the latest journal position with the request's cursor, so a change committed before the subscription started is not missed.
saying these in an interview costs you the question
- The long-poll response should carry the full list of changed files
- A missed notification means the device permanently loses that change
- Long polling needs a special protocol beyond ordinary HTTP
- Every client should use the same fixed long-poll timeout
- The notifier only needs to watch for signals after the request starts