What makes a metric a counter rather than a gauge, and why is a counter read as a rate over a window?
answer
- Which direction may the value move?
- Think about a deploy mid-window
- Totals are not comparable across processes
- Increase per second, with resets absorbed
basics
~20 sA counter only ever increases and returns to zero when its publishing process restarts; a gauge is a level that moves either way. Because a counter's absolute value depends on process uptime, you read its rate over a window.
solid answer
~50 sA **counter** is monotonic: it only increases while the publishing process lives, and it returns to zero when that process restarts. A **gauge** is a level that can move either way. Because a counter's absolute value depends on how long the process has been up, it is not comparable between processes and means little on its own — what you want is its slope. A rate over a window gives that: the engine walks the observations in order, treats any drop as a restart and adds the pre-drop value instead of recording a negative step, then divides the accumulated increase by the window length. Subtracting two readings does none of that — it goes negative across a deploy, is not expressed per unit of time, and depends on exactly which two observations you caught. A gauge is read directly; rating a gauge is meaningless.
code
pseudocode · 8 linescounter orders_placed_total # in-memory accumulator, only ever +=
gauge orders_in_flight # a level, may rise and fall
# wrong: goes negative across a restart, and has no time unit
throughput = orders_placed_total(now) - orders_placed_total(now - 15m)
# right: reset-aware increase per second over the window
throughput = rate_over(orders_placed_total, 15m)go deeper
Be ready to define both types in one sentence each and to say which of a list of examples is which. Requests served is a counter; open connections is a gauge. Knowing that a counter is read as a rate is the part interviewers actually listen for.
Explain the mechanics: a counter is an in-memory accumulator scoped to one process, so a restart returns it to zero, and a rate function absorbs that drop as a reset rather than reporting a negative change. Say why a difference between two readings is not throughput.
Show that you have debugged a dashboard built on the wrong arithmetic. Talk about what a rolling deploy does to counter-based panels, why the window length changes what the graph shows, and when you would publish paired counters instead of a gauge for a level.
Own the convention, not the single metric. Argue for a house rule that events are always counted and levels are always gauged, so that dashboards and alerts across many teams compose; the cost of a fleet using both shapes for the same idea is unreviewable queries.
## A type is a contract about valid arithmetic In every mainstream metrics system a time series carries a declared **type**, and that type is not a storage detail — it is a contract with whoever queries the data later about which arithmetic is valid on it. A **counter** promises that the value only ever moves upward for as long as the process publishing it stays alive. A **gauge** promises nothing of the kind: it is a level that may rise, fall and rise again between one observation and the next. Almost every quietly wrong graph in a monitoring stack begins with a value published under the wrong one of those two contracts. ## Monotonic, and what a restart does *Monotonic* here means non-decreasing: `4`, `4`, `9`, `9`, `1204` is a legal run of counter observations; `9` followed by `7` is not. The promise is scoped to **one publishing process**, and that scope is the part people forget. A counter is normally an in-memory accumulator, so when the process restarts the accumulator is gone and the counter begins again at zero. Nothing is corrupted, but a naive reader sees the value fall off a cliff. That is why a query engine reading a running-total counter treats a decrease as a **reset** rather than as a negative change: 1. It walks the observations inside the requested window in timestamp order. 2. Wherever an observation is lower than the one before it, it assumes the publishing process restarted, and adds the pre-drop value into the running increase instead of recording a negative step. 3. It divides the accumulated increase by the real elapsed time of the window, producing an increase-per-second figure. The result stays continuous across a deploy. You get that behaviour only because the series was declared a counter; publish identical numbers as a gauge and no reset handling happens anywhere. ## Why a rate, and not a subtraction A counter's absolute value is close to meaningless on its own. It says how many events have happened since the publishing process started, so two processes started three weeks apart report wildly different totals for identical traffic. What is meaningful is the **slope**. Subtracting the value now from the value an hour ago looks like a shortcut to that slope, and it fails in four ways: - **It breaks on restart.** Any reset inside the window makes the subtraction negative or absurdly small, exactly when the system is least healthy. - **It is not expressed per unit of time.** A difference is "events since then", not "events per second", so two panels covering different ranges are not comparable to each other. - **It depends on which two observations you happen to catch.** One missed observation at either end silently changes the answer, and nothing on the graph says so. - **It hides the shape.** A rate computed over a sliding window shows *where inside* the window the work actually arrived; two endpoints cannot. | | Counter | Gauge | |---|---|---| | Direction | Non-decreasing while the process lives | Free to move either way | | On process restart | Returns to zero; readers absorb the drop as a reset | Simply reports whatever the new value is | | How you read it | Rate of increase over a window | The value itself, or min/max/mean over a window | | Typical subjects | Requests served, bytes sent, errors, retries | Queue depth, open connections, memory in use | | Classic mistake | Subtracting two readings and calling it throughput | Using one for something that is really an event count | ## Reading a gauge A gauge is read directly: the value carried at that timestamp is the answer, and min, max and mean over a window are all sensible. Applying a rate to one is meaningless, because the difference between two gauge observations is not an accumulation of anything. The compensating weakness is that a gauge is a **point sample**: whatever the level did between two observations is simply not recorded anywhere. A queue that drained and refilled twice between observations looks perfectly flat. ## A worked example A seed-catalogue ordering service was driven to a 5,400-request-per-second peak by a campaign, and afterwards nobody could explain the incident from the existing dashboards. The order-throughput panel had been built as *orders total now minus orders total fifteen minutes ago*. Three of the fourteen service processes had restarted mid-peak under memory pressure, and each restart contributed a large negative term, so the panel showed throughput sagging to roughly 2,900 per second at the exact moment the service was busiest. The counter was correct throughout; the arithmetic layered on top of it was not. Rebuilt as a reset-aware rate over the same window, the panel showed the peak plainly. ## Rules of thumb - If you are counting **events**, use a counter, even when you expect the number to stay small. - If you are reporting a **level**, use a gauge, and accept that you cannot see between samples. - If a level moves both ways but you also care how often it moved, publish two counters — total in and total out — and derive the level; counters survive a missed observation, a gauge does not. - Never reset a counter on a schedule: it turns a well-defined restart signal into noise no reader can tell apart from a crash.
- A counter observation reads 4,180,000 and the very next one reads 12. What should the query engine do with that?Treat the drop as a restart, not as a negative change. It adds the pre-drop value (4,180,000) into the accumulated increase for the window and then counts the 12 on top, so the reported increase-per-second stays continuous across the restart instead of showing a huge negative spike or a hole.
- Why is a gauge that happens to only ever increase still not a counter?Because the type is a contract with the reader, not a description of the values seen so far. Nothing applies reset handling to a gauge, dashboards and alerts will read its value directly rather than its slope, and the moment the process restarts the series drops to zero with no reader able to tell that apart from a genuine fall.
- When would you publish two counters instead of a single gauge for something that is really a level?Whenever you care how often the level moved, not only where it is. Publishing total-enqueued and total-dequeued counters lets you derive both the throughput in each direction and the depth as their difference, and counters tolerate a missed observation because the next one still carries the accumulated total.
A counter is a car's odometer and a gauge is its speedometer. The odometer only climbs, and it is useful for the distance covered between two glances rather than for its own reading; swap the engine and it starts from zero again.
saying these in an interview costs you the question
- Says a counter may decrease as long as it climbs again afterwards
- Subtracts two counter readings and calls the result throughput
- Expects a restart to show up as a negative rate on the graph
- Publishes a request count as a gauge and rates it later
- Believes a counter's absolute value is comparable between two processes
- Treats resetting a counter on a nightly schedule as normal practice