skip to content

In PromQL, how do rate(), irate() and increase() differ, and how should the range relate to the scrape interval?

level: middleimportance: must knowfreq 76%

answer

  1. Three ways to read one counter
  2. How many samples each one uses
  3. Resets handled, edges extrapolated
  4. Range versus scrape interval

basics

~20 s

rate() averages the per-second increase across every sample in the range; irate() uses only the last two samples; increase() gives the total rise. All handle counter resets and extrapolate, so use a range of at least four scrape intervals.

solid answer

~50 s

All three take a range vector over a counter. `rate` averages the per-second increase across every sample in the window, `irate` uses only the final two samples and so reacts instantly but ignores everything before them, and `increase` returns the total rise over the window — arithmetically `rate` multiplied by the window in seconds. `rate` and `increase` both treat a drop between consecutive samples as a process restart, adding the pre-drop value back rather than reporting a negative rate, which is why neither belongs on a gauge. Both also extrapolate to the window edges, so `increase` on an integer counter can return 42.7. The range must contain at least two samples; the usual rule of thumb is at least four scrape intervals, which survives a missed scrape. Use `rate` by default, `irate` only for high-resolution graphs of volatile signals.

code

promql · 3 lines
promql
rate(hearings_booked_total[5m])
irate(hearings_booked_total[5m])
increase(hearings_booked_total[1h])

go deeper

for a junior

Recall that all three read a counter over a bracketed window and that rate gives a per-second figure while increase gives a total. Knowing rate is the everyday default is enough at this level.

for a middle

Explain which samples each function consumes, how counter resets are absorbed, and why extrapolation makes increase return fractions. Be able to size a range from a stated scrape interval.

for a senior

Justify a range choice against real failure modes: missed scrapes, rolling restarts, spike visibility versus smoothing, and why irate belongs on a graph a human is watching rather than in an expression that pages someone.

for a principal

Own the convention. A fleet-wide default range interacts with scrape interval, alert responsiveness and read cost at once, so decide it deliberately and make it consistent rather than leaving each dashboard author to guess.

## The three functions, side by side All three take a **range vector**, a bracketed window such as `[5m]`, and all three are meant for counters: series that only increase and reset to zero when the exporting process restarts. | Function | Samples it uses | What it returns | Typical use | |---|---|---|---| | `rate(x[5m])` | every sample in the window | average per-second increase over the window | graphs, alert expressions, anything aggregated | | `irate(x[5m])` | the last two samples only | per-second increase between those two points | high-resolution graphs of volatile signals | | `increase(x[5m])` | every sample in the window | total increase over the window | answering "how many in the last hour" | `increase(x[5m])` and `rate(x[5m]) * 300` express the same calculation; `increase` is the readable form when a human wants a count rather than a rate. ## Counter resets and extrapolation Two behaviours are shared by `rate` and `increase`, and they explain most surprising results. **Reset handling.** When a sample is lower than the one before it, the engine treats the drop as a restart of the exporting process: the counter is assumed to have fallen to zero and climbed again, so the value before the drop is added into the total instead of producing a negative rate. This is exactly why these functions belong only on counters. Applied to a gauge, a value that legitimately goes down, every genuine decrease is silently reinterpreted as a reset and the result is nonsense. For a gauge the corresponding tools are `delta` and `deriv`. **Extrapolation.** Samples almost never land exactly on the edges of the window. `rate` and `increase` measure the rise between the first and last sample they can see and then project it out toward the window boundaries, with a bounded adjustment so the projection cannot run away when samples are sparse. The visible consequence is that `increase` on an integer counter returns fractional values: 42.7 rather than 43. Engineers meeting this for the first time usually conclude the data is broken. It is not; it is the price of a window that does not align with scrape times. ## Choosing the range The range is not cosmetic. It decides how many samples the function has to work with. - `rate` needs **at least two samples** inside the window. A range shorter than one scrape interval frequently contains one sample or none, and those series vanish from the result rather than erroring. - A single failed scrape removes a sample, so a window of exactly two scrape intervals is one missed scrape away from producing nothing. - The common rule of thumb is a range of **at least four scrape intervals**, which tolerates a missed scrape while still reacting quickly. For a courtroom-scheduling platform scraped every 15 seconds that means `[1m]` at the aggressive end and `[5m]` as the safe default. - Longer windows smooth. A `[30m]` rate of a spiky booking counter shows a gentle hill where a `[2m]` rate shows the spike. Neither is wrong; they answer different questions. - The range must also outlive the gaps you expect. If a booking pod is only scraped while it is up and it restarts a few times an hour, a short window leaves holes in the graph that look like outages. ## Where irate belongs, and where it does not `irate` discards everything in the window except the final two samples and divides by the time between them. That makes it extremely responsive: a spike lasting one scrape interval appears at full height rather than averaged down. It also makes it fragile. 1. **On a graph rendered at low resolution it lies by omission.** If the display step is five minutes but `irate` only inspects the last two samples before each step, the many samples in between are never examined. Spikes are not smoothed away; they are skipped entirely. 2. **In an alert expression it is the wrong instrument**, because the firing decision then rests on two data points, one of which may be a scrape that landed slightly late. 3. **After aggregation the noise compounds.** `sum(irate(...))` across 41 replicas adds up 41 independent two-point estimates, and the resulting line jitters far more than the underlying traffic does. The rule most teams settle on is short: `rate` by default; `irate` only when a human is staring at a high-resolution graph of a fast-moving signal and wants the shape rather than the trend; `increase` when the question is genuinely "how many" and not "how fast".

  • What happens if the range you give rate() contains only one sample?
    The series is dropped from the result. A rate needs two points to subtract, so with fewer than two samples in the window there is nothing to compute and the series is silently omitted rather than raising an error. This is why a range narrower than two scrape intervals produces gaps, and why one failed scrape can empty a marginal window during exactly the incident you are investigating.
  • Why should irate() not be used in an alert expression?
    Because the firing decision would rest on two samples. irate discards the rest of the window, so a single late or missing scrape swings the value dramatically, and the resulting alert flaps. Alerting wants the stability of an average across the whole window, which is what rate provides. irate earns its place only when a human is watching a high-resolution graph and wants the true shape of a fast-moving signal.
  • What should you use instead when the series is a gauge rather than a counter?
    delta() for the change across the window and deriv() for the per-second slope fitted over it. Both are built for values that legitimately go down. Applying rate or increase to a gauge silently misreads every genuine decrease as a counter reset and adds the pre-drop value back, which inflates the result rather than failing visibly.

rate() is your average speed across the whole journey, irate() is the speedometer reading between the last two mileposts, and increase() is the difference on the odometer.

saying these in an interview costs you the question

  • Claiming irate() is simply a more accurate rate()
  • Using rate() or increase() on a gauge metric
  • Treating a fractional increase() result as corrupted data
  • Choosing a range shorter than two scrape intervals
  • Assuming a negative counter drop produces a negative rate