skip to content

In a cloud-drive sync service, what should happen when a device returns with a cursor older than the change journal's retention window?

level: seniorimportance: nice to knowfreq 30%

answer

  1. compaction outruns an offline device
  2. a signal that is not empty
  3. bookmark taken before listing
  4. per-file record of last sync
  5. never wipe unsynced local work

basics

~20 s

The server rejects the expired cursor with a distinct reset signal, and the device performs a full resync: list the current tree with a fresh cursor, reconcile it against local files without discarding unsynced local work, then continue incrementally.

solid answer

~50 s

Journals are compacted, so a device offline longer than the retention window holds a cursor pointing at entries that no longer exist. The server must answer with a distinct, machine-readable reset signal (for example `410 Gone` or a specific error code), never with an empty page that looks like no changes. The device then does the same thing a brand-new device does: take a **snapshot cursor**, list the tree as of that cursor, and reconcile. The snapshot cursor must be captured at the start of the listing, so changes made while listing are replayed afterwards. Reconciliation uses the device's per-file sync record: a local file missing on the server that was synced before and is unmodified was deleted remotely, so delete it locally; one that is modified or never synced is local work, so upload it; when both sides changed, apply the conflict rules. Resets should be rare and staggered, because a mass reset is a scan storm.

code

pseudocode · 22 lines
pseudocode
function full_resync(device):
    snap = server.start_snapshot()          # cursor C taken first
    remote = server.list_tree(snap)          # paginated view as of C
    for path in union(remote.paths, device.local.paths):
        r = remote.get(path)
        l = device.local.get(path)
        if r and not l:
            download(r)
        else if l and not r:
            if l.last_synced_rev and not l.modified_locally:
                delete_local(l)              # removed on server
            else:
                upload(l)                    # local work
        else:
            remote_changed = r.rev != l.last_synced_rev
            if remote_changed and l.modified_locally:
                make_conflicted_copy(l)
            else if remote_changed:
                download(r)
            else if l.modified_locally:
                upload(l)
    device.cursor = snap.cursor              # resume from C

go deeper

for a junior

Know that change logs are trimmed, so a device that was away too long must rebuild its view with a full listing.

for a middle

Explain the snapshot-cursor order and why taking the cursor after listing loses changes made during the listing.

for a senior

Walk through the reconciliation table and the per-file sync record that tells a remote delete from local work, and how to signal the reset.

for a principal

Treat resets as a capacity and safety event: pick a retention window against storage cost, stagger mass resets, and alert when the reset rate spikes.

## Why cursors expire A change journal cannot grow forever. The server keeps entries for a **retention window** and compacts older ones. That is fine for devices that sync daily, but a laptop left in a drawer for longer than the window returns with a cursor that points at compacted entries. The server can no longer produce the list of changes since that point, so incremental sync is impossible for that device. The same situation arises in other ways: a brand-new device has no cursor at all, a server-side restore can invalidate positions, and a migration that re-keys journals may deliberately expire old cursors. ## Signalling the reset The server must tell the device clearly that its cursor is unusable: - Use a **distinct, machine-readable** response, such as HTTP `410 Gone` or a 4xx with an explicit error code like `cursor_reset`. - Never return an empty page, which the device would read as nothing changed. - Never return a generic 5xx, which the device would retry forever. On receiving the signal the device discards the old cursor and switches to full-resync mode. ## Initial sync with a snapshot cursor A full resync is the same procedure a new device uses: 1. Ask the server for a **snapshot**: a cursor `C` plus a consistent view of the tree as of `C`. 2. List the tree at that view, page by page. 3. Reconcile the listing against local files. 4. Save `C` as the device's cursor. 5. Resume incremental sync from `C`, which replays every change committed while the listing ran. The order matters. If the device listed first and only afterwards asked for the current cursor, any change committed during a long listing would fall between the two and never be seen. Capturing `C` first means those changes are simply replayed; since applying entries is idempotent, a change already visible in the listing is harmless to apply again. ## Reconciliation rules The device keeps a local **sync record** per file: the last server revision it synced and whether the local copy was modified since. That record is what makes reconciliation safe. | Server listing | Local state | Action | |---|---|---| | Present | Absent | Download | | Absent | Synced before, unmodified | Deleted remotely: delete locally | | Absent | Modified or never synced | Local work: upload | | Newer revision | Unmodified | Download | | Same revision | Modified | Upload | | Newer revision | Modified | Conflict: apply conflicted-copy rules | The dangerous shortcut is to wipe the local folder and download everything. It is simple and destroys any work the user did offline, which is exactly the situation that produced the long offline gap. The opposite shortcut, uploading every local file the server lacks, resurrects everything deleted remotely while the device was away. A subtle case: a file the user deleted locally while offline has no local file, so the first row would download it again unless the client also remembers local deletions. Re-downloading is the safe failure (nothing is lost), but a good client records local deletes too. ## Operating resets at scale - **Mass resets are expensive.** Each is a full listing plus hashing of local files. An incident that expires many cursors at once can overload metadata servers. - **Stagger** resets with jittered backoff and serve snapshot listings from read replicas where the consistency model allows. - **Raise retention** before a risky migration rather than expiring cursors wholesale. - **Monitor** the reset rate; a sudden rise usually means a retention or cursor-format bug, not a wave of dusty laptops. ## Summary An expired cursor is not an error to retry; it is an instruction to rebuild state from a snapshot. Capture the snapshot cursor first, reconcile with per-file sync records, and never trade the user's offline work for simplicity.

  • Which HTTP response fits an expired sync cursor?
    410 Gone is a natural fit, since the requested position no longer exists, or a 4xx carrying an explicit error code. What matters is that the signal is distinct and machine-readable: an empty page would look like no changes, and a 5xx would trigger endless retries. The client must map it to full resync, not to a retry.
  • How do you stop a mass cursor reset from overloading the service?
    Stagger resets with jittered backoff, page the snapshot listings, serve them from read replicas where acceptable, and prioritise recently active devices. Better still, avoid the event: raise journal retention before a migration and alert on the reset rate, since a spike usually points at a bug rather than genuinely stale devices.

saying these in an interview costs you the question

  • Retry the expired cursor until the server accepts it again
  • Upload every local file the server lacks as a new file
  • Take the snapshot cursor after the listing finishes
  • Return an empty page when the cursor is too old
  • Wipe the local folder and download everything to be safe