skip to content

RethinkDB

A document database whose signature feature is changefeeds: subscribe to a query and get every subsequent change pushed to you. Interviewers ask about it as the reference example of real-time push at the database layer, and about why betting on an unmaintained project is a risk.

on this pageshow

questions

5

What does a RethinkDB changefeed emit when you call .changes() on a table?

level: middleimportance: must knowfreq 70%

answer

  1. Cursor that never ends
  2. Two fields per emitted change
  3. Nulls tell you insert from delete
  4. Point, filter and range feeds exist too
  5. Aggregations generally cannot carry a feed

basics

~10 s

Calling .changes() returns a cursor that stays open and pushes one document per change, each shaped {old_val, new_val}. An insert has old_val null, a delete has new_val null, an update carries both.

solid answer

~50 s

`.changes()` turns a query into a **changefeed**: instead of a finite result set, `run(conn)` hands back a cursor that never ends and delivers a document every time a matching row changes. Each element has two fields, `old_val` and `new_val`. On an insert `old_val` is null; on a delete `new_val` is null; on an update both are present, so the client can diff before and after. This is RethinkDB's signature feature — the push is done by the database itself, so an application does not need a separate broker or polling loop to learn that data moved. Feeds attach to more than whole tables: `get(id).changes()` watches one document, and `filter`, `getAll` and `between` feeds report only rows matching the query. Adding `{includeTypes: true}` adds a `type` field labelling each change as an add, remove or change.

code

javascript · 7 lines
javascript
const feed = await r.table('games').changes().run(conn);
feed.each((err, change) => {
  if (err) throw err;
  if (change.old_val === null) handleInsert(change.new_val);
  else if (change.new_val === null) handleDelete(change.old_val);
  else handleUpdate(change.old_val, change.new_val);
});

go deeper

for a junior

Recall that .changes() gives you a live cursor and that each change carries old_val and new_val, with a null on the side that does not exist yet or no longer exists.

for a middle

Explain the mechanics: which queries can carry a feed, how insert, update and delete map onto the two fields, and that the server pushes on commit rather than the driver polling behind your back.

for a senior

Show that you have thought about cost — whole documents on the wire, feeds served by the shard primary, and a per-feed load that scales with both subscriber count and write rate.

for a principal

Frame the architectural choice: push in the database versus a broker or CDC pipeline, and what you give up in replay, fan-out and operational independence by putting the subscription in the database.

## The idea Every database can answer "what is the data now?". RethinkDB's distinguishing feature is that it can also answer "tell me every time that answer changes", pushed to the client as it happens. That is a **changefeed**, created by appending `.changes()` to a query. ```javascript const feed = await r.table('games').changes().run(conn); feed.each((err, change) => { console.log(change.old_val, change.new_val); }); ``` The cursor returned here behaves like an ordinary cursor except that it never reaches an end. It blocks until something happens, then yields the next change. ## The shape of a change Each element is a document with two fields, `old_val` and `new_val` (the field names are literally snake_case in every driver, because they are document fields rather than API methods): - **Insert** — `old_val` is `null`, `new_val` is the new document. - **Delete** — `old_val` is the document as it last existed, `new_val` is `null`. - **Update** — both are present: the document before and after the write. Because the whole before-and-after documents are shipped, the client can compute exactly what changed rather than re-reading the row. That is convenient for small documents and expensive for large ones, which is a real design consideration on a busy table. `{includeTypes: true}` adds a `type` field to each change so the client can branch on the kind of event without inferring it from null checks; the same option is what surfaces the special initial and state markers when those are enabled. ## What you can attach a feed to A feed is not limited to whole tables. The common forms are: - `r.table('t').changes()` — every write to the table. - `r.table('t').get(id).changes()` — a **point changefeed** on one document, the cheapest and most precise form. - `r.table('t').getAll(x, {index: 'i'}).changes()` and `r.table('t').between(a, b, {index: 'i'}).changes()` — index-scoped feeds. - `r.table('t').filter(...).changes()` — only documents matching the predicate; a document that stops matching produces a change too, which clients often forget. - `r.table('t').orderBy({index: ...}).limit(n).changes()` — a feed over a top-N window. What you generally cannot do is hang a feed off an arbitrary aggregation. Queries that fold many documents into a computed result — grouped counts, general reductions — are not supported as feeds, because there is no incremental way to express the change. The practical rule is: feed the raw rows and aggregate on the client, or maintain the aggregate as a document and watch that document with a point changefeed. ## Where the push comes from A changefeed is served by the replica responsible for the data — the primary for that shard. The server tracks the open feed and dispatches matching writes to it as they commit. Two consequences follow. First, the work is proportional to the number of open feeds and the write rate, so tens of thousands of chatty feeds are a capacity question, not a free lunch. Second, the feed's fate is tied to that server: a failover or a dropped connection ends the feed, and the client must resubscribe. ## Why interviewers ask about it RethinkDB is the reference example of *push at the database layer*. The usual alternative architectures — polling on a timer, dual-writing to a message broker, or reading a replication log — each add latency, staleness or an extra moving part. A changefeed collapses that into the query you were already writing, which is why the pattern keeps being cited even by teams who never deployed RethinkDB. Knowing the `{old_val, new_val}` shape, the null conventions, and which queries can carry a feed is the core of the answer; knowing what the feed does *not* guarantee is the senior half.

  • On a filter-based changefeed, what happens when an update makes a document stop matching the predicate?
    The feed reports it. You receive a change whose `old_val` is the document as it matched and whose `new_val` is null, signalling that the row left the result set — even though the document still exists in the table. Clients that treat a null `new_val` as "deleted from the database" get this wrong.
  • Why can't you attach a changefeed to a grouped aggregation such as a count per category?
    There is no general incremental form of an arbitrary reduction, so the server cannot say what the aggregate's before and after values are without recomputing. The usual workarounds are to feed the underlying rows and aggregate client-side, or to maintain the aggregate as its own document and put a point changefeed on it.

saying these in an interview costs you the question

  • Saying the client polls and .changes() just hides it
  • Thinking only inserts are reported
  • Expecting a diff of changed fields instead of whole documents
  • Assuming any query can carry a changefeed
  • Treating a null new_val as always meaning row deleted

context

open as a page

What delivery guarantees does a RethinkDB changefeed give, and what happens if the client disconnects?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Changefeeds are live-only push, not a durable log. There is no offset to resume from, so anything that happens while a client is disconnected is lost, and consecutive changes to one document may be coalesced into a single before/after pair.

open as a page

How does RethinkDB's ReQL query language differ from sending SQL query strings?

level: juniorimportance: should knowfreq 50%

basics

~20 s

ReQL is embedded in the host language: you chain driver methods to build a query object, then call run(connection) to execute it. Nothing is sent as a string, and nothing runs until run() is called.

open as a page

How do you build a live top-10 leaderboard in RethinkDB with a single changefeed?

level: middleimportance: should knowfreq 40%

basics

~10 s

Attach a changefeed to an ordered, limited query: orderBy on an index, limit(10), then .changes({includeInitial: true}). The feed sends the current top ten first, then one change each time the window's membership shifts.

open as a page

What is RethinkDB's maintenance status, and how would you weigh adopting it today?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

RethinkDB the company shut down in 2016; the code was relicensed under Apache 2.0 and moved to the Linux Foundation, maintained since by a small community with infrequent releases. Adopting it now is mainly a sustainability bet, not a technical one.

open as a page