skip to content

Why does a CouchDB reduce function receive a rereduce flag, and what must the function guarantee?

level: seniorimportance: nice to knowfreq 33%

answer

  1. Reduction is not one pass over all rows
  2. Partial results are stored in the tree
  3. The function eats its own output
  4. Grouping shape is not under your control
  5. Growing output is refused outright

basics

~20 s

CouchDB stores partial reduce results inside the view B-tree's inner nodes, so it must combine already-reduced values as well as raw emitted ones. The rereduce flag tells the function which it is being given, and the output must stay small no matter how much input it summarizes.

solid answer

~50 s

A view's reduce function is not run once over all rows. CouchDB reduces each leaf group of the B-tree, stores that partial result in the parent node, and then combines partials as it climbs — which is what makes a range query over a million rows cheap. So the function is called two ways. With `rereduce` false, `keys` holds `[key, docid]` pairs and `values` holds the raw emitted values. With `rereduce` true, `keys` is `null` and `values` holds the outputs of previous reduce calls. The function must handle both and must be combinable in any grouping — effectively associative — because the tree shape is not something you control. It must also shrink: a reduce whose output grows with its input (returning `values` itself, or accumulating a list) blows up the inner nodes and CouchDB aborts it with a reduce overflow error. Prefer the built-ins `_count`, `_sum`, `_stats` and `_approx_count_distinct`, which run natively and are far faster.

code

javascript · 8 lines
javascript
function (keys, values, rereduce) {
  if (rereduce) {
    var total = 0;
    for (var i = 0; i < values.length; i++) { total += values[i]; }
    return total;
  }
  return values.length;
}

go deeper

for a junior

Know that a view can carry a reduce function alongside its map, and that built-in names like _count and _sum give you totals without writing JavaScript.

for a middle

Explain the two call shapes and what keys and values contain in each, and why a reduce that ignores the rereduce flag gives wrong answers once the index spans more than one tree node.

for a senior

Demonstrate the constraints in practice: outputs that must shrink, the reduce overflow failure, why averages need sum-and-count, and choosing built-ins over JavaScript for query-server cost.

for a principal

Own the read-path strategy: decide which aggregates deserve a permanently maintained reduced view versus ad-hoc querying, and keep the count of design documents from sprawling into unmaintainable near-duplicates.

## Where reduction actually happens A CouchDB view index is a B-tree of `(key, value)` rows sorted by key. When a view defines a `reduce` function alongside its `map`, CouchDB does not keep the reduce result in one place. It reduces the rows in each leaf node, stores that partial result **in the leaf's parent**, then reduces those partials into their parent, all the way to the root. The tree carries precomputed summaries at every level. That design is what makes reduced range queries fast. Asking for the sum over a key range does not scan the range; CouchDB walks down to the edges of the range, grabs the already-computed subtotals of the whole inner nodes that fall inside it, and combines only those. The work is proportional to the depth of the tree, not to the number of rows. ## The two call shapes The consequence is that the function is asked to combine two different kinds of input, and the third argument tells it which: ```javascript function (keys, values, rereduce) { if (rereduce) { return sum(values); // values are previous reduce outputs } return values.length; // values are raw emitted values } ``` - **`rereduce === false`** — the initial pass over actual view rows. `keys` is an array of `[key, docid]` pairs and `values` is the array of values your map function emitted. - **`rereduce === true`** — combining partials. `keys` is `null`, and `values` is an array of things this same function returned earlier. A function that ignores the flag and assumes it always has raw emitted values will produce wrong answers silently as soon as the view grows past a single B-tree node — a bug that never appears in a small test database and appears in production. ## The rules a reduce function must obey **It must be combinable in any grouping.** You do not control how the tree splits, and splits change as documents are added. Sum, count, min, max and their combinations are fine. An average is not — average of averages is wrong — so you compute it by reducing to `{sum, count}` and dividing on the client, which is exactly what `_stats` does for you. **Its output must be small and shape-stable.** The result is stored inside inner tree nodes, so if it grows with the number of rows summarized, the index degenerates: the nodes bloat, the rereduce work grows, and the whole point of the structure is lost. CouchDB detects this and refuses, raising a **reduce overflow** error. The two classic offenders are `return values;` (returns the whole input) and building an array or object keyed by every distinct value seen. If you want the rows, you do not want a reduce — you want `reduce=false`, which returns the underlying map rows. **It should be idempotent under re-application.** Reducing a set of partials must give the same answer as reducing the raw values directly, or grouped queries and full-range queries will disagree. ## Built-in reducers CouchDB ships native reduce functions written as strings: `"_count"`, `"_sum"`, `"_stats"` and `"_approx_count_distinct"`. They execute inside the database rather than in the JavaScript query server, so they avoid the per-call serialization cost of shipping values out to an external process. For counting and summing they are dramatically faster than the equivalent hand-written JavaScript, and they are correct by construction. Reach for a custom reduce only when the built-ins genuinely cannot express what you need. `_sum` also handles arrays and objects of numbers element-wise, which covers many multi-metric cases without custom code. ## Grouping By default a reduced view query collapses the entire range to one value. `group=true` reduces per distinct key instead. With array keys, `group_level=N` reduces per distinct N-element prefix — so keys shaped `[year, month, day]` give yearly totals at `group_level=1`, monthly at `2`, daily at `3`, and one grand total with no grouping at all. Building that one time-hierarchy view replaces what would otherwise be several separate rollups, and it is the single most common reason to use array keys. ## When not to reduce If what you want is the matching documents rather than a summary, do not write a reduce function at all, or pass `reduce=false`. And note that a reduce answers only over key ranges of that view — it cannot filter on anything you did not encode in the key. Trying to force a general-purpose aggregation engine out of reduce functions is the classic way to end up with a dozen near-duplicate design documents; often the right answer is Mango for ad-hoc predicates and a small number of well-chosen reduced views for the aggregates you serve constantly. ## Interview framing The expected answer: reduction is tree-structured, so the function must combine its own previous outputs, and it must shrink its input or the index cannot hold the intermediate results.

  • Why can't a reduce function compute an average directly?
    Because averages are not combinable: the average of a set of averages is not the average of the underlying values unless every group is the same size, and group sizes depend on how the B-tree happens to split. Reduce to a sum and a count — which is what the built-in _stats returns — and divide on the client.
  • What does group_level=2 do on a view whose keys are arrays like [2026, 3, 15]?
    It reduces separately for each distinct two-element prefix, so you get one row per year-and-month combination. The same view answers yearly totals at group_level 1 and daily totals at 3, which is why time-hierarchy array keys are the standard pattern for rollups.
  • When would you write a custom reduce instead of using _count or _sum?
    Only when the aggregate genuinely is not expressible with the built-ins — a bitwise union, a custom min-max-by-field pair, a bounded top-N. The built-ins execute natively rather than in the JavaScript query server, so a hand-written equivalent is slower and offers no benefit beyond familiarity.

saying these in an interview costs you the question

  • Assumes reduce runs once over the full result set
  • Ignores the rereduce argument entirely
  • Returns the values array to collect rows
  • Computes an average by averaging partial averages
  • Hand-writes a sum instead of using the built-in

context