skip to content

What does CouchDB's _changes feed return, and how do the normal, longpoll and continuous modes differ?

level: middleimportance: must knowfreq 70%

answer

  1. Reports documents, not every revision
  2. One row per changed document
  3. since gives you a resumable bookmark
  4. normal, longpoll, continuous, eventsource
  5. heartbeat keeps an idle socket open

basics

~20 s

The _changes feed lists documents that changed since a given sequence, one row per document with its current revision. feed=normal returns immediately, feed=longpoll holds the request open until a change arrives, feed=continuous streams rows indefinitely.

solid answer

~40 s

`GET /{db}/_changes` is CouchDB's public change log. Each row is `{"seq", "id", "changes": [{"rev"}]}` plus `"deleted": true` for a tombstone, and the feed is **coalesced per document** — a document updated fifty times appears once, at its latest sequence, with its current revision. It is not an operation log you can replay. The response ends with `last_seq` and `pending`; you store `last_seq` and pass it back as `since` to resume. `feed=normal` returns everything up to the current end and closes; `feed=longpoll` blocks until at least one change is available; `feed=continuous` streams one JSON object per line and never closes (`feed=eventsource` is the SSE flavour). For the two long-lived modes use `heartbeat` so proxies do not kill an idle socket. `style=all_docs` lists every leaf revision instead of only the winner, which is what replication uses.

code

bash · 2 lines
bash
# resume from a stored bookmark, blocking until something changes
curl 'http://host:5984/orders/_changes?feed=longpoll&since=23-g1AAAAB&timeout=60000&heartbeat=10000'

go deeper

for a junior

Know that CouchDB exposes a _changes endpoint listing recently changed documents and that you pass since to continue where you left off. Be able to name the normal, longpoll and continuous modes.

for a middle

Explain the coalescing rule (one row per document, current revision), why sequences are opaque strings in a clustered CouchDB, and what heartbeat and timeout are for on a long-lived feed.

for a senior

Show you have run a consumer in production: persisting last_seq after processing, idempotent handling of re-delivered rows, reconnect loops, and why include_docs is not a point-in-time snapshot.

for a principal

Own the decision of what should consume the feed at all — one shared indexer versus many consumers, backpressure and pending, and when a change feed is the wrong substrate for an audit or event-sourcing requirement.

## What the feed is `GET /{db}/_changes` returns the documents that have changed in a database, ordered oldest change first. It is CouchDB's change log, exposed over the same ordinary HTTP API as everything else, and it is the foundation of replication, of PouchDB sync, of external indexers, and of any change-driven service you bolt onto CouchDB. Nothing about it is privileged: anything that can issue an HTTP GET can follow it, which is exactly why a browser database can implement CouchDB replication. ## What a row actually contains A row looks like: ```json {"seq": "7-g1AAAA...", "id": "order:912", "changes": [{"rev": "4-cf9b"}]} ``` with `"deleted": true` added when the change was a deletion. Two properties surprise people. First, the feed is **coalesced per document**, not per revision. A document written fifty times shows up once, positioned at the sequence of its most recent write, carrying its current revision. You cannot use `_changes` to reconstruct history or to see intermediate states; it answers "which documents differ from what I already know", which is precisely what a replicator needs and is a poor fit for an audit trail. Second, `changes` is an array because a document can have more than one leaf revision at once. By default only the winning leaf is listed. ## Sequences are opaque bookmarks The response ends with `last_seq` and `pending` (an estimate of how many changes remain). In clustered CouchDB the `seq` value is a long opaque string that packs a sequence number for every shard of the database, which is why it looks like `"23-g1AAAAB..."` rather than an integer. Never parse it, never compare two of them to infer ordering, and never persist one and replay it against a different cluster or a database that has been deleted and recreated. Treat it as a bookmark CouchDB gave you and hand it back verbatim in `since`. `since=now` starts at the current end of the feed and skips history. ## The three modes - **`feed=normal`** (the default): returns every row from `since` to the current end of the feed, then closes. Good for batch catch-up and for a client that will poll on its own schedule. - **`feed=longpoll`**: identical, except that when there is nothing new the server holds the request open until a change occurs or `timeout` elapses, then returns an ordinary JSON response. This gives near-real-time updates with plain request/response plumbing and no streaming parser. - **`feed=continuous`**: streams one JSON object per line as changes happen and does not close. The client parses line by line. `feed=eventsource` is the same delivery model framed as Server-Sent Events so browsers can consume it with `EventSource`. ## Keeping long-lived feeds alive For longpoll and continuous, `heartbeat=<ms>` makes CouchDB emit a blank line periodically so that reverse proxies, load balancers and the client's own read timeout do not tear down a connection that is merely idle. `timeout=<ms>` bounds how long the server waits before responding empty. Even with heartbeats, any long-lived feed will eventually be dropped, so the client must persist `last_seq` and reconnect with `since=<that seq>`. That reconnect loop is the whole reliability story. ## Options that change what you see - `style=all_docs` emits every leaf revision for each changed document rather than only the winner. This is how a replicator discovers conflicting revisions so it can carry all of them to the target. - `include_docs=true` inlines the document body — but the body you get is the one current at read time, not necessarily the one that produced the row. Do not treat it as a point-in-time snapshot. - `conflicts=true` (with `include_docs`) adds the `_conflicts` array to the inlined documents. - `filter=ddoc/name` runs a JavaScript filter function; `filter=_selector` applies a Mango selector supplied in the request body; `filter=_doc_ids` with `doc_ids` restricts to a fixed list; `filter=_design` emits only design documents. - `limit` and `descending` apply as they do elsewhere, though `descending` on a changes feed is rarely what a client wants. ## How replication uses it A replicator asks the source for `_changes` with `style=all_docs` starting from the sequence recorded in its checkpoint, batches the ids and revisions it sees, and then asks the target which of those revisions it is missing. Everything the replicator knows about "what is new" comes from this feed, which is why the feed's coalescing and its opaque sequences shape the protocol so strongly. ## Common pitfalls Polling with `since=0` every time turns an incremental feed into a full scan. Persisting a sequence and replaying it against a rebuilt database gives you either an error or a full rewind. Running a long-poll behind a proxy without `heartbeat` produces mysterious 30-second disconnects. And expecting `include_docs` to show the state as of that sequence produces subtle bugs in downstream indexers.

  • A client stored last_seq, the database was deleted and restored from a backup, and the client reconnects with that since value. What happens?
    The sequence is meaningful only to the database instance that issued it. Against a rebuilt or restored database it is not a valid bookmark, so the client either gets an error or is rewound and re-reads changes it has already processed. Treat a restore as a reset: clear the stored sequence and re-scan, and make downstream processing idempotent so a rewind is harmless.
  • Why does style=all_docs matter to a replicator but rarely to an application?
    Default `main_only` lists only the winning revision of each changed document. A replicator must carry every leaf revision so the target ends up with the same revision tree, conflicts included, and so it asks for `all_docs`. An application usually only wants the winner, and asking for all leaves just adds rows it has to ignore.
  • Your continuous feed consumer misses changes after a network blip. What is the fix?
    Persist `last_seq` after each successfully processed batch and reconnect with `since=<that value>`, never `since=now`. `since=now` skips everything that happened while you were disconnected. Also add `heartbeat` so idle connections are not silently dropped, and make the consumer idempotent, because a reconnect can re-deliver the last rows you had already handled.

It is a mailbox flag rather than a diary: it tells you which documents have news since you last looked, not the story of everything that happened to them.

saying these in an interview costs you the question

  • Thinks _changes replays every revision of a document
  • Parses or numerically compares the since sequence value
  • Reconnects with since=now and loses changes made while offline
  • Runs a long-poll behind a proxy without a heartbeat
  • Believes rows include the document body by default

context