In PromQL, what do the by and without modifiers do to a result's labels, and why must rate() come before sum()?
answer
- Grouping decides the output labels
- One keeps, the other drops
- The metric name never survives
- Reduce per series before summing
basics
~20 sby keeps only the labels listed; without keeps everything except them; both drop the metric name. rate() must be applied per series first, because summing counters across replicas turns a restart into a drop that rate misreads as a reset.
solid answer
~40 sAn aggregation operator groups the input series and emits one series per group, and the grouping clause decides the output labels. `by (courtroom)` keeps only `courtroom`; `without (instance)` keeps every label except `instance`; with no clause everything collapses to one unlabelled series. In all cases the metric name is dropped, so the result cannot be selected by name afterwards. Order matters just as much: `sum(rate(x[5m]))` is correct, `rate(sum(x)[5m:1m])` is not. A counter is only monotonic within one process, so summing counters across replicas produces a series that falls whenever a replica restarts or scales away; `rate` reads that fall as a single reset and adds back the pre-drop value of the whole sum, producing a phantom spike. Applying `rate` per series first yields per-second values that add up cleanly.
code
promql · 2 linessum by (courtroom) (rate(hearings_booked_total[5m]))
sum without (instance, pod) (rate(hearings_booked_total[5m]))go deeper
Recall that by lists the labels to keep and without lists the labels to drop, and that a query returns one series per group. Being able to read sum by (courtroom) (...) out loud is the bar here.
Explain what happens to labels the clause does not mention, why the metric name disappears, and why rate has to be applied before an aggregation rather than after. Be ready to describe the phantom spike a restart creates.
Diagnose the failure in real data: a dashboard showing an impossible spike at a deploy, and the reasoning that traces it to counters aggregated before the rate. Also weigh the stability tradeoff between by and without as a metric gains labels.
Own the conventions that stop this recurring: which grouping labels are canonical across teams, how queries stay valid when metrics gain dimensions, and how a reviewable house style keeps aggregation order right without relying on individual recall.
## What an aggregation operator does to a label set PromQL's aggregation operators — `sum`, `min`, `max`, `avg`, `count`, `stddev`, `stdvar`, `group`, `count_values`, `topk`, `bottomk` and `quantile` — take an instant vector, group its series, and emit one output series per group. The grouping clause decides what those output series are called. - `by (labels)` keeps **only** the listed labels on the result. Everything else is discarded. - `without (labels)` keeps **everything except** the listed labels. - With no clause at all, every input series collapses into a single output series carrying no labels. - The metric name is dropped either way. `sum by (courtroom) (hearings_booked_total)` produces series identified by `courtroom` alone, with no metric name, which is why you cannot select the result by name afterwards. | Query | Labels on the result | |---|---| | `sum(hearings_booked_total)` | none, one series | | `sum by (courtroom) (hearings_booked_total)` | `courtroom` only | | `sum without (instance) (hearings_booked_total)` | every original label except `instance` | | `sum by (courtroom, job) (hearings_booked_total)` | `courtroom` and `job` | The practical difference between the two clauses shows up when somebody adds a label later. A query written `by (courtroom)` is stable: a new `region` label appears on the source series and the result is unchanged. A query written `without (instance)` is deliberately permissive: the new label survives and the result silently splits into more series. Use `by` when you know exactly the shape you want, and `without` when the intent really is "collapse away the per-target labels and keep the rest", with `without (instance, pod)` as the archetype. ## Why order matters: rate first, then aggregate `sum(rate(x[5m]))` and `rate(sum(x)[5m:1m])` look like one idea with the operations swapped. They are not, and the second is wrong. A counter belongs to one process. Its meaning — increases only, resets to zero on restart — is a property of that process's lifetime. Aggregating counters across processes destroys that property. 1. **A restart becomes a cliff.** In a courtroom-scheduling platform with 41 booking replicas on a 7-node cluster, one pod restarting takes its counter from 128,904 to 0. The sum across replicas drops by 128,904 within a single scrape. Applied to that summed series, `rate` sees a decrease, treats it as one reset, and adds back the pre-drop value of the entire sum, producing a single enormous phantom spike. 2. **Scaling moves the sum for no real reason.** When the deployment goes from 41 replicas to 47, six counters start at zero and join the sum; when it scales back down, six series disappear and the sum falls. Neither event has anything to do with how many hearings were booked. 3. **The type is wrong as written.** `sum(x)` is an expression rather than a selector, so `[5m]` cannot be attached to it without a resolution, which forces a subquery. That subquery re-evaluates the inner aggregation at every inner step, so the wrong answer is also the expensive one. Do it the other way round and every problem disappears. `rate` is applied per series, inside each process's own lifetime, where a reset is meaningful and is handled. Its output is a per-second value that behaves like a gauge, and gauges add cleanly across replicas. `sum by (courtroom) (rate(hearings_booked_total[5m]))` is bookings per second per courtroom, and it stays correct across restarts and rescaling. ## The general rule and its edges The rule generalises past counters: **reduce to a per-series value first, aggregate second.** Concretely: - `sum(rate(...))`, never `rate(sum(...))`. - `sum(increase(...))` when you want totals across replicas. - `avg` of rates is meaningful, but it is average per-replica throughput; `sum` of rates is total throughput. Choose deliberately, because the two differ by the replica count and both look plausible on a dashboard. - `max_over_time` and its family follow the same shape: the range function collapses time within a series, then an aggregation operator collapses series. - `count` is the exception that catches people out: `count(rate(x[5m]))` counts series that produced a rate, which is a way of counting live replicas rather than measuring traffic. - The rule does not hold for anything already reduced to an estimate. Estimated quantiles cannot be re-aggregated by averaging them, because the distribution they came from is no longer in the data. If one sentence survives from all of this: the brackets belong next to a selector, and everything that consumes them belongs inside the aggregation, not outside it.
- When would you prefer without over by?When the intent is to collapse away specific per-target labels and keep whatever else exists. `without (instance, pod)` keeps job, courtroom and any label added later, which is what you want for a query that should follow the metric as it gains dimensions. `by` is the safer choice when you want a fixed, stable output shape, because a new label on the source series cannot silently split the result into more series.
- What is the difference between sum and avg of a rate across replicas?sum gives total throughput for the group; avg gives per-replica throughput. They differ by the number of replicas, and both plot as plausible-looking lines, so choosing by accident is a real source of wrong dashboards. Total is what you want for capacity and traffic; the average is what you want when comparing how evenly load is spread and it moves when replica count changes even if traffic does not.
- Why can you not select the result of an aggregation by its metric name?Because aggregation drops the metric name along with any label not named in a by clause. The output series are identified only by their remaining labels. If you need a named series to build on, record the expression under a new name so later queries have something to select; otherwise the expression has to be repeated in full wherever it is needed.
saying these in an interview costs you the question
- Thinking by and without differ only in syntax
- Expecting the metric name to survive an aggregation
- Summing counters across replicas and then taking a rate
- Assuming a restart cannot distort an aggregated counter
- Reaching for avg when total throughput was intended