Why does a derivative over a date_histogram misreport change when intervals have no documents?
answer
- It subtracts the bucket before it
- 'Before' means the previous returned bucket
- Empty intervals may carry no metric value
- One setting decides skip versus zero
- Dropping empty buckets changes what 'adjacent' means
basics
~20 sA derivative subtracts consecutive buckets, so what counts as consecutive depends on which buckets came back. Metrics like avg yield no value on an empty interval, and min_doc_count above zero drops those intervals entirely, so the derivative silently spans a gap.
solid answer
~50 s`derivative` computes each bucket's value minus the previous bucket's, so its correctness rests on the bucket series being complete and evenly spaced. Two things break that. First, an interval with no documents still produces a bucket by default, but the metric inside may have **no value** — `avg` yields null where `sum` yields 0 — and `gap_policy` decides what happens: the default `skip` emits no derivative there and continues from the next available value, while `insert_zeros` substitutes 0. Second, if you set `min_doc_count` above 0 on the `date_histogram`, the empty intervals disappear from the response entirely, and the derivative then compares two buckets that may be weeks apart while still calling it a day-over-day change. The fixes are to leave `min_doc_count` at 0, use `extended_bounds` so the series covers the intended range, and choose `gap_policy` deliberately: zeros for counts and sums, `skip` for averages and rates.
code
json · 19 lines{
"size": 0,
"aggs": {
"daily": {
"date_histogram": {
"field": "@timestamp",
"calendar_interval": "day",
"min_doc_count": 0,
"extended_bounds": { "min": "2026-08-01", "max": "2026-08-31" }
},
"aggs": {
"orders": { "value_count": { "field": "order_id" } },
"daily_change": {
"derivative": { "buckets_path": "orders", "gap_policy": "insert_zeros" }
}
}
}
}
}go deeper
Recall that a derivative reports the change from the previous bucket and that the first bucket has no value to report.
Explain the mechanics: empty intervals appear by default but their metric may be null, and gap_policy chooses between skipping the bucket and inserting zero.
Diagnose the silent version — min_doc_count above zero removes intervals, so 'adjacent' buckets can be weeks apart and the chart looks plausible while being wrong.
Own the modelling call across dashboards: whether zero is a truthful value for a metric decides the gap policy, and inconsistent choices across teams make the same number mean different things.
## What derivative actually computes The `derivative` pipeline aggregation is a parent pipeline that must be nested in a `histogram` or `date_histogram`. For each bucket it emits the metric's value minus the **previous bucket's** value. It has no notion of time beyond bucket order — "previous" simply means the bucket before this one in the returned series. That is why the composition of the series decides whether the number means what you think it means. The first bucket never gets a derivative value at all: there is nothing before it to subtract. Chart code that assumes every bucket carries the field will break on the first point. ## Where empty intervals come from A `date_histogram` fills gaps between the first and last matching document with empty buckets by default, because its `min_doc_count` defaults to 0. It does **not** extend beyond the data unless you ask: `extended_bounds` (or `hard_bounds` to clamp) forces the series to cover a range you specify, which matters when a dashboard shows a fixed 30-day window but the data starts on day 12. So by default the intervals are there. What is missing is sometimes the **metric value inside them**. A `sum` or `value_count` over an empty bucket returns 0 — a real number the derivative can use. An `avg`, `min`, `max` or `percentiles` over an empty bucket returns null, because there is nothing to average. That null is the "gap" the pipeline machinery talks about. ## gap_policy Every pipeline aggregation accepts `gap_policy`: - **`skip`** (the default) treats a missing value as though the bucket did not exist. No derivative is emitted for that bucket, and the following bucket's derivative is computed against the last value that did exist. - **`insert_zeros`** replaces the missing value with 0 before computing. Choosing between them is a modelling decision, not a formatting one. For a count or a revenue total, a day with no orders genuinely *is* zero, so `insert_zeros` produces the honest series and a real drop-then-rebound in the derivative. For an average basket size or an error rate, zero is a lie: no orders does not mean the average basket was £0. There `skip` is correct, and the chart should show a hole rather than a crash to zero. ## The min_doc_count trap The more damaging failure is setting `min_doc_count` to 1 on the `date_histogram`, often done to shrink a response with many sparse intervals. Now the empty intervals are gone from the series entirely, and `derivative` happily subtracts bucket *n-1* from bucket *n* — except those two buckets may be a fortnight apart. The response still labels the value as the change for that day. Nothing errors; the numbers are simply wrong, and they are wrong in a way that looks plausible on a chart. If you must drop empty buckets for response size, do not run a `derivative`, `serial_diff`, `cumulative_sum` or `moving_fn` over the result. ## Normalising by time Calendar intervals are not equal in length: months have 28 to 31 days. A monthly `derivative` therefore mixes "change over 31 days" with "change over 28 days". Setting the `unit` parameter — for example `"unit": "day"` — adds a `normalized_value` alongside `value`, expressing the change per unit of time, which makes consecutive months comparable. The raw `value` is still there; `normalized_value` is the one to chart when interval lengths differ. ## The related series pipelines The same considerations govern the neighbouring parent pipelines. `cumulative_sum` produces a running total across the buckets in order and is the standard way to draw a cumulative curve. `moving_fn` applies a function over a sliding `window` of previous bucket values — a moving average, a maximum, an exponentially weighted average — via a script such as `MovingFunctions.unweightedAvg(values)`. (The older `moving_avg` aggregation was deprecated and removed; on Elasticsearch 8.x and later `moving_fn` is the supported form.) All of them read the returned bucket series, so all of them inherit the same dependence on that series being complete and evenly spaced, and all of them respect `gap_policy`. ## How to answer in an interview Say that the derivative is a difference between adjacent *returned* buckets; that empty intervals are returned by default but their metric may be null; that `gap_policy` decides between skipping and zero-filling and the right choice depends on whether zero is meaningful for the metric; and that raising `min_doc_count` silently makes "adjacent" mean something else. That sequence covers the whole trap.
- When is insert_zeros the wrong gap_policy?Whenever zero is not a truthful value for the metric. For an average basket size, an error rate or a percentile, an interval with no documents does not mean the value was zero — inserting zeros drags the series down and creates fake swings in the derivative. Use `skip` there and let the chart show a hole; reserve `insert_zeros` for counts, sums and other additive metrics.
- What does the derivative's unit parameter add to the response?On a `date_histogram`, setting `unit` (for example `"day"`) adds a `normalized_value` beside `value`, expressing the change per unit of time rather than per bucket. It matters because calendar intervals are unequal — a monthly series mixes 28- and 31-day buckets — so `normalized_value` is what makes consecutive buckets comparable on a chart.
- How do you compute a running total and a smoothed trend over the same date_histogram?`cumulative_sum` over the metric gives the running total; `moving_fn` with a `window` and a script such as `MovingFunctions.unweightedAvg(values)` gives the smoothed series. Both are parent pipelines declared inside the histogram, both honour `gap_policy`, and both can coexist with a `derivative` in the same block. On Elasticsearch 8.x the older `moving_avg` aggregation no longer exists; `moving_fn` replaces it.
saying these in an interview costs you the question
- Sets min_doc_count to 1 and still trusts the derivative
- Uses insert_zeros for averages and rates
- Expects a derivative value on the first bucket
- Thinks a date_histogram covers a fixed range without extended_bounds
- Assumes calendar months are equal-length intervals