Why do CouchDB offline-sync deployments use a database per user instead of one filtered replication?
answer
- CouchDB authorizes whole databases
- No per-document read permission exists
- Filters shape transfer, not authority
- One database per user is the boundary
- Master database replicated down, users fanned in
basics
~20 sCouchDB authorizes readers per database, not per document, so a replication filter only shapes what a sync transfers and never restricts what a client may read. Confidentiality between users therefore requires giving each user their own database.
solid answer
~50 sCouchDB's `_security` object grants membership and admin rights on a **whole database**; there is no per-document read permission. A replication filter decides which documents a particular replication carries, but a client holding credentials for that database can bypass it entirely by reading `_all_docs`, `_changes` unfiltered, or any document by id. So a filter is a bandwidth optimisation, not an access boundary. The standard architecture is one database per user (CouchDB can even create them automatically as `userdb-*`), with shared reference data pushed into each user database by server-side replication from a master, and per-user data fanned back into a reporting database when you need cross-user queries. The cost is real: thousands of databases each carry their own view indexes, compaction schedule, file handles and continuous replications, and any query spanning users needs the fan-in. The alternative is to stop syncing directly and put an API in front of CouchDB — which means giving up offline replication.
code
json · 5 lines// PUT /userdb-616c696365/_security
{
"admins": {"names": [], "roles": ["_admin"]},
"members": {"names": ["alice"], "roles": []}
}go deeper
Know that CouchDB permissions apply to a whole database and that a replication filter controls only what gets copied, so it cannot keep one user from reading another user's documents.
Explain the mechanics: _security members grant read access database-wide, validation functions constrain writes only, and a client can bypass a filter with _all_docs or an unfiltered change feed.
Lay out the working architecture — a database per user, shared data replicated in from a master, per-user data fanned into a reporting database — and the failure modes each piece introduces.
Own the trade at scale: index and replication cost per database, rolling design-document deployments across a fleet, and when to abandon direct sync for an API layer that forfeits offline entirely.
## The security model is the whole reason CouchDB authorizes at database granularity. A database's `_security` document names admins and members; a member can read every document in that database and, subject to validation functions, write to it. There is no notion of a document-level read permission, and validation functions (`validate_doc_update`) constrain **writes** only — they cannot hide a document from a reader. This is not an oversight; it follows from replication. A replication peer must be able to enumerate the change feed and fetch arbitrary revisions. A database whose readable subset varied per client could not be replicated coherently, because two clients would compute different histories of the same database. ## Why a filter is not a boundary It is tempting to sync one big database to every device with `filter: "app/mine"` and assume each user only receives their own documents. The transfer is indeed limited. The **authority** is not. A client with credentials for that database can: - `GET /db/_all_docs?include_docs=true` and read everything; - open `_changes` without any filter; - fetch any document by id, if ids are guessable — and in a system with ids like `user:alice:profile`, they are. A filter is server-side code that shapes a replication. It is not consulted when a client makes an ordinary read. Treating it as a permission is the single most consequential mistake in CouchDB architecture, and an interviewer asking this question is usually probing exactly that. ## The database-per-user pattern So the boundary has to be the database: 1. **One database per user (or per tenant).** Its `_security` names exactly one member. CouchDB ships a feature that will create such a database automatically when a user is created, naming it from a hex encoding of the username (`userdb-...`); many teams instead provision explicitly from their own signup flow, which gives them control over naming and initial content. 2. **The device syncs only its own database**, usually with PouchDB and a live, retrying two-way sync. No filter is needed for isolation, because isolation is structural. 3. **Shared reference data** — catalogues, price lists, configuration — lives in a master database and is pushed into each user database by server-side replication, or is synced to the device as a second, read-only replication from a shared database the user can read but not write. 4. **Cross-user queries** — reporting, admin views, anything that must see all users — are served by replicating every user database *into* a central database (fan-in) and querying that. The fan-in copy is derived data; nobody edits it. ## What it costs This pattern scales to a lot of users, but the operational bill is real and a principal-level answer names it: - **Per-database overhead multiplies.** Every database has its own files, its own compaction, and — critically — its own copy of every view index in its design documents. Ten thousand users with three views means thirty thousand indexes to build and maintain. - **Replication count.** Fan-out from a master and fan-in to a reporting database means two server-side continuous replications per user, each holding a change feed open. The scheduler will queue what it cannot run concurrently, so replication latency becomes a capacity question rather than a constant. - **Design-document deployment is a fleet operation.** Changing a view means pushing a new design document to every user database and paying an index rebuild in each. Teams usually script this and roll it out in waves. - **Sharing between users is not free.** Anything two users must both see either gets duplicated into both databases or lives in a third database both can read. Modelling that up front matters more than any other decision in the design. - **Ids and file handles.** Thousands of open databases consume file descriptors and memory; this is a real limit to plan against rather than discover. ## When to choose something else If the product does not actually need offline operation, the alternative is straightforward and often better: keep CouchDB behind an application API that authorizes each request, and never expose the database to clients at all. You give up replication — no sync, no offline writes, no conflict-free local reads — and in exchange you get arbitrary authorization rules and a single deployment of query logic. The honest framing for an interview is that this is a trade between two coherent architectures. Direct sync buys genuine offline capability and pushes read authorization into the database's coarse model, which forces database-per-user and its operational tail. An API layer buys fine-grained authorization and a simpler operational story, and forfeits offline. Choosing filters as a middle path buys neither: you keep the operational simplicity of one database and you have no isolation at all. ## The tell in an interview A candidate who says "filter the replication per user" without qualification has not understood the security model. A candidate who says "database per user, and here is what that costs me at ten thousand users" has.
- Can validate_doc_update be used to hide documents from a reader?No. Validation functions run on the write path and can reject an update, so they enforce who may change what. They are never consulted when a document is read, listed in `_all_docs`, or emitted by the change feed. Read authorization in CouchDB exists only at database granularity, which is precisely why isolation has to be structural.
- How does shared reference data reach thousands of per-user databases?Either by server-side replication that pushes a master database's documents into each user database, or by having devices run a second read-only replication from a shared database every user may read. Push duplicates storage per user but keeps the device on one sync; the shared-database route avoids duplication and adds a second replication plus a second local store to manage.
- What breaks first as the number of per-user databases grows?Usually index and replication overhead rather than raw storage. Every database maintains its own copy of every view in its design documents, and fan-out plus fan-in means multiple continuous replications per user, each holding a change feed. Deploying a design-document change becomes a fleet-wide rolling rebuild. Plan for file descriptors and scheduler capacity before you plan for disk.
- When would you put an API in front of CouchDB instead of syncing directly to clients?When the product does not truly need offline writes, or when authorization rules are finer than any database partition can express — field-level redaction, role-dependent visibility, cross-tenant sharing. An API layer gives you arbitrary rules and one place to deploy query logic. The price is losing replication entirely, so it is a decision to make deliberately rather than drift into.
saying these in an interview costs you the question
- Says a replication filter keeps users from reading each other's documents
- Thinks validate_doc_update can hide documents from readers
- Assumes per-document read permissions exist in _security
- Proposes database per user without naming the index and replication cost
- Believes cross-user reporting works without a fan-in database