What does emit(key, value) in a CouchDB map function build, and when is that index updated?
answer
- Runs once per document, not per query
- Zero, one, or many rows per document
- The output is stored and ordered
- Ordering fixes what queries are possible
- Refreshed by what changed, at read time
basics
~20 sEach emit call adds one row to a persistent B-tree sorted by key, so a view is a precomputed, ordered index you query by key or key range. CouchDB updates it incrementally, by default when the view is next queried.
solid answer
~50 sA view lives in a design document as a JavaScript `map` function that CouchDB runs once per document. Every `emit(key, value)` call appends a row to the view's B-tree; a document may emit zero rows, one, or many. The B-tree is sorted by key using CouchDB's own collation order, which is why the only things you can ask a view are key lookups and key ranges — `key`, `keys`, `startkey`/`endkey`, plus `descending`, `limit`, `skip` and `include_docs`. Maintenance is incremental: when the view is queried, CouchDB maps only the documents that changed since the index was last built, then serves the query. `update=false` serves whatever is already built without waiting, and `update=lazy` serves the stale index but triggers a background build. Editing any function in a design document invalidates that design document's whole index group and forces a full rebuild.
code
json · 8 lines{
"_id": "_design/orders",
"views": {
"by_customer_date": {
"map": "function (doc) { if (doc.type === 'order') { emit([doc.customerId, doc.date], doc.total); } }"
}
}
}go deeper
Know that a view is JavaScript stored in a design document, that emit adds a row, and that you query it by key or key range rather than by arbitrary fields.
Explain the B-tree and its key ordering, why that limits queries to lookups and ranges, and how incremental maintenance works with the default, update=false and update=lazy.
Show operational judgment: warming views after bulk writes, blue/green design-document swaps to avoid rebuild stalls, view cleanup, and choosing between emitted values and include_docs on real payload sizes.
Own the modelling consequence: with no ad-hoc planner, every read path must be designed in advance, so treat view definitions as part of the schema contract and budget rebuild time into deployment plans.
## A view is a materialized, sorted index CouchDB has no ad-hoc query planner for views. You declare, in advance, a JavaScript `map` function stored inside a design document (`_design/…`), and CouchDB runs it against every document in the database. The function's only output channel is `emit(key, value)`. Each call produces one row `(key, value, id)`; a function that calls `emit` three times for one document produces three rows, and a function that calls it zero times excludes the document entirely — which is how you build filtered indexes. Those rows are stored in a B-tree ordered by key. That single fact determines everything a view can do. ## Querying is range scanning Because the index is a sorted tree, the query surface is deliberately small: `GET /db/_design/{ddoc}/_view/{name}` accepts `key` for an exact match, `keys` for a set, `startkey` and `endkey` for a range, and modifiers `descending`, `limit`, `skip`, and `include_docs`. There is no arbitrary filtering, no OR across unrelated fields, no sorting by anything except the key. If you need a different access path, you emit a different key — often a **compound key**, an array such as `[customerId, orderDate]`, which lets you range-scan all of one customer's orders in date order with `startkey=["c1", "2026-01-01"]` and `endkey=["c1", "2026-02-01"]`. Keys sort by CouchDB's collation, which orders across JSON types in a fixed sequence: `null`, then booleans, then numbers, then strings (compared with Unicode collation, not raw byte order), then arrays, then objects. Two consequences bite people: string ordering follows ICU collation rather than ASCII byte order, and `descending=true` reverses the traversal, so you must also swap `startkey` and `endkey`. ## Incremental maintenance The index is not rebuilt from scratch on each query. CouchDB records how far through the database's change sequence the index has been built. On the next query, it maps only the documents that changed since that point, updates the affected B-tree rows, records the new sequence, and answers. Cost is therefore proportional to churn since the last read, not to database size — with one nasty corollary: a view that is queried once a day pays for a day of changes on that one query, and the first query after a bulk import can take a very long time. The update behaviour is controllable per query. The default builds before answering. `update=false` answers immediately from the index as it stands, which may be missing recent writes — the right choice for dashboards and autocomplete. `update=lazy` answers from the current index and kicks off a background refresh, which keeps a rarely-read view roughly warm. Production deployments commonly add a small "warmer" job that queries each view with `limit=0` after big writes, so no user request ever eats the build. ## Design documents are the unit of indexing All views inside one design document form a single index group, built and stored together. Adding a view to an existing design document therefore forces the whole group to rebuild. Worse, changing so much as whitespace in a map function changes the design document's content, which changes the index identity and triggers a complete rebuild for every view in it. The standard mitigation is to deploy a *new* design document, warm its index in the background, then switch reads over and delete the old one — a blue/green swap. `POST /db/_view_cleanup` removes the index files orphaned by that swap. ## Map functions must be pure Because the index is incremental, a map function must be a deterministic function of the single document passed to it. It cannot read other documents, cannot call out to the network, and must not use the clock or a random number: a document mapped last Tuesday will keep last Tuesday's output until that document itself changes, so `Date.now()` in a map function bakes in a stale value and produces an index that is wrong in ways no error will report. ## emit(key, doc) versus include_docs Emitting the whole document as the value makes reads a single B-tree scan with no extra lookups, but duplicates the document into the index file and inflates it. `include_docs=true` keeps the index small and fetches each document by id at query time, at the cost of a lookup per row and a small race: the document returned is the current one, which may be newer than the row that pointed at it. Emit `null` and use `include_docs` for large documents; emit the two or three fields you actually render for hot, high-volume views. ## Interview framing One sentence: `emit` writes rows into a persistent B-tree sorted by key, so a CouchDB view is a precomputed index whose only query shape is a key range, refreshed incrementally at read time unless you tell it otherwise.
- Why must a CouchDB map function avoid Date.now() or random values?Because the index is incremental. A document is re-mapped only when that document changes, so a clock or random value is frozen into its rows and silently drifts out of date. The map function must be a deterministic, side-effect-free function of the one document it is given.
- When should you emit the whole document as the value instead of using include_docs=true?Emit the fields you need only for hot, high-volume views where the extra per-row document lookup matters and the documents are small. For large documents prefer emitting null with include_docs, which keeps the index file small — at the cost of one lookup per row and possibly returning a newer body than the row reflected.
- How do you change a view's map function without a user-facing stall?Deploy the new function in a new design document rather than editing the old one, query it with limit zero to warm its index in the background, then switch reads to it and delete the old design document. Editing in place invalidates the whole index group and forces the next reader to wait for a full rebuild.
A view is the index at the back of a printed book: someone had to decide in advance which terms get listed, and once printed you can only look things up the way it was alphabetized.
saying these in an interview costs you the question
- Thinks the map function runs at query time
- Expects a view to filter on fields it never emitted
- Assumes a document produces exactly one view row
- Edits a design document in place on a large production database
- Uses the current time inside a map function