What delivery guarantees does a RethinkDB changefeed give, and what happens if the client disconnects?
answer
- Live only, nothing stored
- Nothing to resume from
- Per-document order, but merging allowed
- A disconnect is a resync, not a retry
- Unread changes cost server memory
basics
~20 sChangefeeds are live-only push, not a durable log. There is no offset to resume from, so anything that happens while a client is disconnected is lost, and consecutive changes to one document may be coalesced into a single before/after pair.
solid answer
~50 sA changefeed delivers changes for as long as the cursor is open, in order per document, but it is **not** a replayable log. There are no offsets, no retention and no acknowledgements: if the connection drops, the primary fails over, or the client is restarted, every change during that gap is simply gone, and the client must resubscribe and re-read. The feed may also coalesce — with `squash` enabled the server merges consecutive changes to the same document, so you see the net effect rather than every intermediate state, trading fidelity and latency for less traffic. Slow consumers cost server memory, because unread changes buffer on the server side. The design rule that follows: treat the feed as a notification channel for state you can always re-derive from the table, subscribe with `includeInitial: true` so reconnection resyncs automatically, and make client handlers idempotent. If you need replay or audit, put a durable log in the architecture — the changefeed is not one.
code
javascript · 11 linesasync function subscribe(conn) {
const feed = await r.table('jobs')
.changes({ includeInitial: true, squash: true })
.run(conn);
try {
for (;;) applyIdempotently(await feed.next());
} catch (err) {
await backoff();
return subscribe(conn); // includeInitial re-primes local state
}
}go deeper
Know the headline: a changefeed only delivers while you are connected, and there is no way to ask for changes you missed.
Explain the mechanics — no offsets or retention, per-document ordering, squashing that merges consecutive edits into one before/after pair, and server-side buffering for slow readers.
Demonstrate the production pattern: subscribe with includeInitial, treat a cursor error as a resync trigger, make handlers idempotent, and never let correctness depend on a notification arriving.
Own the boundary decision: which events require a durable, replayable log with retention and consumer offsets, and which are cheap notifications where re-deriving state from the table is acceptable.
## What the feed promises While a changefeed cursor is open and being read, RethinkDB pushes the changes matching your query, and for any single document the changes arrive in the order they happened. That is the guarantee. Everything an application usually wants beyond it — durability, replay, acknowledgement, exactly-once — is absent by design. ## No offsets, therefore no replay A log-based system gives each record a position, and a consumer resumes by remembering where it stopped. A changefeed has no such coordinate. The server holds the subscription in memory and dispatches to it; when the subscription ends there is nothing to point back at. So: - A dropped connection ends the feed. Changes during the outage are never delivered. - A failover of the shard's primary ends the feed, because the feed was being served by that replica. - A client crash and restart begins a brand-new subscription with no memory of the old one. The cursor surfaces the break as an error on the next read, which is the client's cue to reconnect. It is not a retryable hiccup you can ignore: correctness requires resynchronizing state, not just reopening the socket. ## Coalescing: you may not see every intermediate value Even with a healthy connection, you are not promised one emitted change per write. RethinkDB can squash multiple changes to the same document into a single `{old_val, new_val}` pair representing the net effect. `squash: true` asks for this explicitly, and a numeric value sets a window over which changes are merged before being sent. That is a deliberate throughput and latency tradeoff: a document being edited many times a second becomes a manageable stream instead of one round trip per edit. The implication for application logic is important. If your handler needs to observe *every* transition — say, to bill for each state change — a changefeed is the wrong instrument, because merging is allowed. If your handler only needs the current state, coalescing is a free optimization. ## Slow consumers are a server problem Because delivery is push, a client that reads slowly forces the server to hold undelivered changes. That memory belongs to the database process, so a handful of stalled subscribers on a hot table is an availability concern for everyone, not just for those clients. Squashing exists partly to bound this. Operationally, treat feed count times write rate as a capacity dimension, monitor it, and be suspicious of designs that open a feed per end user on a table that every user writes to. ## The reconnection pattern The idiom that makes changefeeds safe is: **subscribe with the initial state included, and treat every reconnection as a resync.** ```javascript async function subscribe(conn) { const feed = await r.table('jobs') .changes({ includeInitial: true, squash: true }) .run(conn); try { for (;;) { const change = await feed.next(); applyIdempotently(change); } } catch (err) { await backoff(); return subscribe(conn); // includeInitial re-primes local state } } ``` `includeInitial` sends the current matching documents before the live changes, so a fresh subscription rebuilds the client's view without a separate query and without a gap between reading and listening. `applyIdempotently` matters because the initial batch re-delivers documents the client already has, and because a change may represent several merged writes. ## What the feed is, and is not, for Good fits: driving UI updates for connected clients, invalidating caches, waking a worker when new rows appear, keeping a small in-process view of a table warm. All of these share one property — the truth lives in the table, and a missed notification costs a delay, not correctness. Bad fits: an event log of record, financial or audit trails, guaranteed job dispatch, cross-system integration that must not lose an event. Those need durability and replay, which means writing the event to a table (or a real log) and treating the feed only as the wake-up signal, so a consumer that missed the push can still find the work by querying.
- How would you use a changefeed for job dispatch without losing jobs?Write the job as a row with a status field and use the feed only as a wake-up signal. Workers claim jobs with a conditional update on status, and a periodic sweep queries for unclaimed rows. A missed push then costs latency until the next sweep, not a lost job.
- What does a numeric squash value change compared with squash: true?Both merge consecutive changes to the same document, but the numeric form sets a time window over which changes are accumulated before being sent. That bounds how often a very hot document can generate traffic, at the cost of adding that much latency to every change on the feed.
- Your feed dies during a primary failover. What should the client do beyond reconnecting?Resynchronize. The cursor error means unknown changes occurred, so reopening with includeInitial and reapplying idempotently is the correct recovery, or re-running the query and reconciling. Simply reopening a plain feed leaves the client silently stale for any row that changed during the gap.
saying these in an interview costs you the question
- Treating a changefeed as a durable replayable log
- Assuming missed changes are buffered until reconnect
- Expecting every intermediate write to be delivered
- Reconnecting without resynchronizing client state
- Ignoring that slow readers consume server memory