In PromQL, which queries are expensive to evaluate, and how do you decide what a recording rule should precompute?
answer
- Series matched times samples read
- Wide selectors, long ranges, subqueries
- Move cost off the read path
- Pay once on a schedule, not per viewer
basics
~20 sCost is roughly series matched times samples read per series, times graph steps. Wide selectors, long ranges and subqueries dominate. A recording rule moves that cost off the read path, paying it once per evaluation rather than per viewer.
solid answer
~50 sOne evaluation costs about **series matched multiplied by samples read from each**, and a range query multiplies that by its step count. Three shapes dominate: selectors wide enough to touch a large slice of the store, such as a regular expression on the metric name; long bracketed ranges, since samples per series scale with range divided by scrape interval; and subqueries, which evaluate their inner expression once per inner step. A recording rule does not make an expression cheaper — it changes **where and how often the cost is paid**, evaluating once on a schedule and storing a small result that dashboards and alerts then read directly. Precompute what alerts evaluate, what incident dashboards read repeatedly, and expressions collapsing large fan-out into a small stable label set. Do not precompute exploratory queries or anything already cheap: every rule is permanent cost, storage and a name others depend on.
code
promql · 2 linessum by (courtroom) (rate(hearings_booked_total[5m]))
max_over_time(sum by (courtroom) (rate(hearings_booked_total[5m]))[24h:1m])go deeper
Recall that a query matching many series over a long window is slow, and that some dashboard queries are precomputed and stored under a new name so that opening the dashboard is fast.
Explain the cost model in concrete terms: how many samples a range pulls per series at a given scrape interval, why a wide selector matters, and what changes for the reader once an expression is recorded.
Diagnose a slow dashboard by finding the expensive shape in the query, then choose between narrowing the selector, shortening the range, removing a subquery, and recording the expression. Know that recording costs evaluation time forever.
Own the policy: what earns a rule, who owns its name and label set, how they are deprecated, and the honest tradeoff between precomputing aggregates and reducing cardinality at the source. Recorded series are an interface other teams build on.
## What a PromQL query costs The engine's work for one evaluation is close to **the number of series the selectors match, multiplied by the number of samples read from each**. A range query multiplies that again by the number of steps in the graph. Nothing else in the language moves the number as much as those two factors. Take a courtroom-scheduling platform on a 7-node cluster holding 2,340,000 active series, scraped every 15 seconds. - `rate(hearing_search_duration_seconds_bucket{job="search"}[5m])` matches 3,472 bucket series and reads about twenty samples from each: roughly 69,000 samples for one evaluation. Unremarkable. - `rate({__name__=~"hearing_.+"}[1h])` matches 384,000 series and reads about 240 samples from each: over 92,000,000 samples for one evaluation. A query engine carries a configurable ceiling on how many samples a single query may load, precisely so that this aborts instead of taking the server down with it. - Drawn as a six-hour graph, that same expression is not one evaluation but hundreds, one per step. Three shapes account for nearly all runaway queries. 1. **Wide selectors.** A regular-expression matcher on the metric name, or a selector with no `job` matcher at all, invites the engine to touch a large fraction of the store. Cardinality that is affordable at write time can be ruinous at read time. 2. **Long ranges.** Samples read per series scale with the bracketed duration divided by the scrape interval. Moving from `[5m]` to `[24h]` is a 288-fold increase in samples read per series, usually for a smoother line nobody asked for. 3. **Subqueries.** `max_over_time(sum by (courtroom) (rate(hearings_booked_total[5m]))[24h:1m])` evaluates its inner expression once per inner step — 1,440 times for that window — and each of those evaluations expands a five-minute range across every matching series. It is the one construct whose cost multiplier the reader has to work out for themselves. ## What a recording rule actually moves A recording rule evaluates one expression on a schedule and writes the result back as a new series. What changes is not the cost of the expression but **where and how often it is paid**. | | Ad-hoc expression | Recorded series | |---|---|---| | When it is evaluated | on every dashboard load and every alert evaluation | once per evaluation interval | | Series touched at query time | thousands to hundreds of thousands | one per output label combination | | Cost of ten people opening the dashboard | paid ten times | no extra cost | | History available | as far back as the raw data | only from when the rule was created | The second row carries most of the value. `sum by (courtroom) (rate(hearings_booked_total[5m]))` expands 3,472 series and collapses them to 41; recorded, the dashboard reads 41 series and a handful of samples each. The read path stops being proportional to the size of the fleet. What it does not do is make anything free. The rule still runs on the same server at every evaluation interval, forever. If the expression is heavy enough the rule falls behind its schedule, and a rule that falls behind delays everything that reads it. Recording an expensive query converts an unpredictable cost per viewer into a predictable cost paid once. It does not remove the cost. ## Deciding what to precompute For a platform whose on-call rotation is three people, the discipline that survives contact with reality is narrow. 1. **Anything an alert evaluates.** Alert expressions run on a fixed schedule forever, so a slow one is a reliability problem rather than a performance annoyance. 2. **Anything the incident dashboards read.** Those are opened repeatedly, under time pressure, by the same three people, often simultaneously. That multiplier is exactly what a recorded series removes. 3. **Expressions that collapse enormous fan-out into a small, stable label set.** The saving is the ratio between input series and output series, so recording something that stays at full cardinality saves almost nothing and doubles the storage. And the things not to record: - Exploratory queries. If nobody has run it twice, it does not need a rule, and you will accumulate rules nobody dares delete because nobody knows what reads them. - Anything already cheap. Every rule has a permanent cost in evaluation time, in storage, and in one more name people must learn. - The raw detail itself. A rule produces an aggregate; keeping the underlying series is a retention decision, not a query one. Two organisational points outweigh the query mechanics at this level. First, a recorded series is a **published interface**: other teams build dashboards and alerts on it, so its name and its label set deserve the same care and the same deprecation manners as an API, which is the whole reason for the convention of naming a recorded series after the aggregation level, the source metric and the operations applied. Second, every recorded series is a second copy of a truth that also exists in raw form, and when the two disagree — because the expression changed, or because the rule was created after the incident someone is now trying to explain — people trust the wrong one. Keep the set small enough that the team can hold all of it in their heads.
- What does a recording rule not fix?It does not make the expression cheap, reduce the cardinality of the underlying series, or provide history from before the rule existed. The work still happens on the same server at every evaluation, and if the expression is heavy enough the rule falls behind its schedule, delaying everything that reads it. It converts an unpredictable per-viewer cost into a predictable scheduled one, nothing more.
- How do you decide between recording an aggregate and reducing what is collected in the first place?Recording helps when the raw detail is genuinely needed for other questions and only the common aggregate is hot. If nobody ever queries the underlying dimensions, the honest fix is upstream: stop producing or stop ingesting the labels that create the fan-out. Recording on top of cardinality nobody uses pays for the same data twice and leaves the write path just as loaded.
- How do you keep a growing set of recorded series from becoming unmanageable?Treat each one as a published interface with an owner, a consistent naming convention describing the aggregation level, source metric and operations, and a way to find out what reads it before deleting it. Keep the set small enough for the team to hold in their heads, and prefer removing a rule nobody reads over adding another that duplicates a truth already available in raw form.
saying these in an interview costs you the question
- Believing a recording rule makes an expensive expression cheap
- Expecting recorded series to be backfilled with history
- Recording an aggregate that stays at full cardinality
- Assuming query cost depends on the graph's pixel width
- Treating subqueries as no more expensive than range selectors
- Adding rules for one-off exploratory queries