skip to content

When a chat client reconnects after an hour offline, how does it catch up on missed messages without refetching every conversation's history?

level: middleimportance: must knowfreq 64%

answer

  1. cost what changed, not what exists
  2. one cursor per user
  3. seq greater than local cursor
  4. newest page first, holes later
  5. same ID, same seq

basics

~20 s

The client keeps a cursor of the last acknowledged sequence and sends it on reconnect. The server first answers which conversations changed, then returns only the messages after each cursor. Offline sends are retried with client-generated IDs so the server can recognise duplicates.

solid answer

~40 s

The client keeps the **last acknowledged sequence** for each conversation, persisted locally. On reconnect it runs a sync in two steps. First, it asks which conversations changed since a single **per-user cursor**. The server keeps a per-user change feed of entries like "conversation C reached seq N", so one number answers that question. Second, for each changed conversation it fetches `after=cursor` in pages. For a huge backlog it fetches the newest page first and lazy-loads the older part. Messages the user wrote while offline sit in a local outbox, each carrying a **client-generated message ID**. If the earlier send actually reached the server, the server sees the same ID, returns the original sequence, and does not create a second message. Cursors advance only after the client has written the messages to local storage.

code

json · 11 lines
json
{
  "request": { "type": "sync", "userCursor": 88120 },
  "response": {
    "userCursor": 88164,
    "changed": [
      { "conversationId": "c-17", "latestSeq": 412 },
      { "conversationId": "c-93", "latestSeq": 9031 }
    ],
    "hasMore": false
  }
}

go deeper

for a junior

Remember that the client stores where it stopped and asks only for messages after that point, instead of downloading everything again.

for a middle

Walk through the two cursors, per-user changes and per-conversation sequence, the paginated fetch, and why the client message ID lets the server recognise a resend.

for a senior

Cover persisting before advancing cursors, cursor expiry and full resync, newest-first paging with holes for huge backlogs, and duplicates where the fetch and the live stream overlap.

for a principal

Weigh how long to retain the change feed against how often clients must fall back to an expensive full resync, and set that from real offline-duration data.

## Why a plain "reload everything" does not work A user who has been offline for an hour may belong to hundreds of conversations, and only a few of them changed. Refetching every conversation's recent history on each reconnect multiplies the load by the number of conversations. Mobile clients reconnect constantly: on network changes, on app resume, after sleep. So the sync must cost roughly **what changed**, not **what exists**. ## The two cursors | Cursor | Scope | Answers | |---|---|---| | **Per-conversation sequence** | one conversation | "which messages in this chat have I not seen?" | | **Per-user change cursor** | one user, all conversations | "which of my chats changed at all?" | The per-conversation sequence is the number the server assigns to each message in a chat. The client stores the highest contiguous value it holds. The **per-user change feed** is a small, append-only list the server keeps for each user. Each entry records something like "conversation C advanced to seq N" or "you were added to conversation D". It has its own increasing position. Without it, the client would have to send one cursor for every conversation it belongs to, and the server would have to compare them all. ## The reconnect sequence 1. The client reconnects and sends its **per-user change cursor**. 2. The server returns the change-feed entries after that position, which gives the list of changed conversations and their latest sequence numbers. 3. For each changed conversation, the client requests messages with `seq > local_cursor`, **paginated**. 4. The client writes each page to local storage, then advances that conversation's cursor. 5. Once the catch-up is complete, the client advances its per-user change cursor and starts processing the live stream. 6. The client drains its **outbox** of messages written while offline. Messages that arrive live during the catch-up are stored and handled by the same gap rules. The fetch and the live stream converge on one contiguous range. ## Large backlogs - For a very busy group with tens of thousands of new messages, fetch the **newest page first**, record the older range as a known **hole**, and fill it only if the user scrolls up. - Unread counts come from the numbers alone (`latest_seq - last_read`), so the badge is right even before the history arrives. - The server retains the change feed for a limited time. If a client's cursor is older than the retained window, the server says so, and the client rebuilds its conversation list from scratch. ## Resends and client-generated message IDs While offline, or on a flaky link, the client cannot know whether an earlier send reached the server. The request might have been lost, or only the ack might have been lost. Before the first attempt, the client therefore attaches a **client message ID**: a unique value it generates itself. - The server records `(conversation, client_message_id) -> seq` when it accepts a message. - A retry with the same ID returns the **original** sequence instead of storing a second copy. - The ack carries the client ID back, so the client can match it to the optimistic "sending" bubble it already shows. A hash of the message text would be the wrong key. Sending "ok" twice on purpose is legitimate, and the two must stay separate messages. ## Correctness rules that keep sync honest - **Advance a cursor only after storing locally.** If the app is killed mid-sync, the next reconnect re-requests anything not yet stored, instead of skipping it. - **Cursors only move forward.** A delayed response must never move a cursor back. - **Treat duplicates as normal.** The overlap between the catch-up fetch and the live stream means a message can arrive twice, and the sequence check simply drops the second copy. ## What this leaves to others How the gateway fleet absorbs thousands of simultaneous reconnects, and how the client spaces out its reconnect attempts, are separate problems. This mechanism covers only *what* the client asks for once it is connected.

  • What should the client do when one group has 50,000 new messages?
    It should not page through all of them. It fetches the newest page so the user sees the current conversation, marks the older range as a known hole, and fills that hole only if the user scrolls back. The unread count is still exact, because it comes from the difference between the latest sequence and the read cursor.
  • Why is a hash of the message text a bad key for recognising resends?
    Identical text can be sent on purpose, for example "ok" twice, and those are two real messages. A client-generated ID identifies one send attempt, not the content, so the server merges retries of that attempt while keeping repeated identical messages separate.
  • What happens if the client's change cursor is older than the server keeps?
    The server cannot answer "what changed since X" anymore, so it replies that the cursor has expired. The client then does a full resync of its conversation list and each conversation's latest sequence, and resumes incremental sync from a fresh cursor.

saying these in an interview costs you the question

  • On reconnect, reload the last 50 messages of every conversation.
  • Advance the cursor as soon as the server response arrives.
  • Deduplicate resends by hashing the message text.
  • Send one cursor per conversation the user belongs to on every reconnect.
  • If a send times out, show it as failed and let the user retype it.