How do you pick a metric's type at instrumentation time, and what can you never recover afterwards?
answer
- Start from the question, not the value
- What did the recording never capture?
- A gauge is blind between observations
- A new name beats changing a type
basics
~20 sPick from the question you will ask later: counters answer how often, gauges answer how much now, distributions answer what the slow end looks like. Nothing unrecorded can be recovered, so fixing a wrong type helps only going forward.
solid answer
~50 sChoose from the question you intend to ask months later, not from the value that happens to be in front of you. If you will ask *how often* or *how fast*, count the events with a counter. If you will ask *how much, right now*, publish a level as a gauge. If you will ever want more than a mean — what the slow end looks like — you must record a distribution at the time, because a shape cannot be reconstructed afterwards. That is what makes the choice partly irreversible: a gauge observed every twenty seconds has already discarded everything that happened between observations, and changing the instrument improves the data only from the change forward, never behind it. Dimensions, names, dashboards and thresholds are all reversible; the shape of what you recorded is not.
code
pseudocode · 8 lines# the same backlog, three ways - and only one answers "how often"
gauge pending_orders # level now; nothing between observations survives
counter orders_enqueued_total # every event counted; rate answers how fast it grew
counter orders_dequeued_total # the pair also reconstructs the level
distribution order_wait_seconds # the shape, at B + 2 series per process
# changing a type later: new name, run both, migrate, retire
# never: same name, different type, mid-rolloutgo deeper
Recall the mapping in one line each: count events with a counter, report levels with a gauge, record a distribution when you care about more than the typical value. Knowing that the wrong choice cannot be fixed retroactively is the important half.
Explain what each type lets you ask later and what it discards. Be able to say why a gauge observed every twenty seconds cannot show a spike that started and ended in between, and why a mean does not show the slow end.
Show the production judgement: which decisions are frozen for the length of the retention window, how you migrate a wrong type without breaking existing queries, and why a counter beside a gauge is the cheapest insurance you can buy.
Own the defaults across teams — count by default, distribute deliberately, publish parts rather than derived values, and require a new distribution to state its series cost. The failure you are preventing is an organisation discovering mid-incident that the question is unanswerable.
## Choose from the question, not from the value The instinct at instrumentation time is to look at what is in front of you — a variable, a queue object — and publish it. That is backwards. Instrumentation is a **recording**, not a query: months later you will ask questions of what you recorded, and never of what you did not. So the decision starts from the question, and the type follows from it. - *How often, or how fast?* — an event happened, so **count** it. A counter's rate answers "per second" at any window you choose later. - *How much, right now?* — a level, so publish a **gauge**, and accept that nothing between two observations is preserved. - *What does the slow end look like?* — you need the shape of many observations, so record a **distribution**. Nothing else can answer it after the fact. - *What is the typical value?* — a mean is enough, and a mean is two counters: a running sum and a count. Far cheaper than a distribution, and legitimate when nobody will look past the typical case. | The question you will ask later | Type to record | What it costs | |---|---|---| | How many, how often, how fast | Counter | One series; cheapest thing you can publish | | How much right now, how deep, how full | Gauge | One series; blind between observations | | What the slow end looks like | Distribution | A family of series per publishing process | | What the typical value is | Two counters (sum and count) | Two series; no shape | ## What is reversible, and what is not This is the part that makes the decision worth a minute of thought rather than a reflex: 1. **The dimensions you attach are reversible.** You can add or remove them, and while that costs storage and produces a discontinuity in the affected series, the underlying measurement survives the change. 2. **Names, dashboards and alert thresholds are reversible.** They sit on top of the recording and can be rewritten at any time. 3. **What you never recorded is gone.** A gauge observed every twenty seconds has already discarded everything that happened in between; no query, no backfill and no reprocessing recovers it. If the level spiked and recovered inside one interval, that spike never existed as far as the store is concerned. 4. **Data already collected keeps the shape it was recorded with.** Changing an instrument improves data from the change forward, never behind it, so today's decision is frozen for the length of your retention window. That asymmetry is why experienced engineers over-instrument slightly at the start: the marginal cost of one extra counter is small and knowable, and the cost of discovering mid-incident that the question cannot be answered is not. ## Changing a type on a live metric name Sometimes you were simply wrong. Changing the type **in place, under the same name** is the worst option available: - Every query written against the old type keeps running and silently returns different arithmetic — a rate over something that no longer accumulates, or a raw reading of something that now climbs forever. - During a rolling deploy both shapes are published under one name at once, so the series is internally inconsistent for as long as the rollout lasts. - The history becomes uninterpretable, and alerts built on it either stop firing or start flapping. The safe migration is to publish the new instrument under a **new name**, run both long enough to cover the retention window that matters, move dashboards and alerts across deliberately, then retire the old one. It is more work, and it is the only version that leaves the history readable. ## A worked example A seed-catalogue ordering service published its pending-order backlog as a gauge, observed every twenty seconds. During a 5,400-request-per-second peak the service degraded, and nobody could explain the incident afterwards from the existing dashboards: the backlog panel showed a maximum of 812 across the whole event, well under the alert threshold of 2,000, yet customers reported minutes of failed checkouts. The gauge was not lying — it was answering a different question from the one being asked. Every observation had genuinely found the backlog under 812; what nobody could see was whether it had climbed far higher and drained in between, because a gauge records nothing between samples. Had the same code also published enqueued and dequeued counters, their rates would have shown the backlog's true growth at any resolution, and the peak would have been visible to a query written after the incident. That would have cost two series per process. ## The cheap insurance - **Count the events, always.** A counter beside a gauge is one extra series and it survives missed observations, which makes it the most retrospectively useful thing on the list. - **Record a distribution only where the shape will really be read**, and count the rest — the family multiplies per publishing process, so this is where a budget goes. - **Write down what question each instrument answers**, next to it in the code. It is the only artefact that makes the choice reviewable by the next person. - **Never publish a value you have already divided** if you will want it aggregated later; publish the parts and divide at query time.
- You inherit a service where the request count is published as a gauge. What is the migration?Add a counter under a new name and publish both. Let them run long enough to cover the retention window the dashboards read, move panels and alerts onto the counter deliberately, then remove the gauge. Changing the type under the existing name would silently change the arithmetic of every query already written against it and would publish two shapes at once during the rollout.
- Which instrumentation decisions are genuinely cheap to revisit later?Dimensions, names on dashboards, alert thresholds and the queries built on top. Adding or dropping a dimension costs storage and puts a discontinuity in the affected series, but the underlying measurement survives it. The expensive decisions are the ones about what gets recorded at all, because those bound what any future query can ask.
- How do you avoid over-instrumenting everything as insurance against this?Make the cheap insurance cheap and the expensive insurance deliberate. Counters are one series and can be added liberally; distributions multiply per publishing process and should be reserved for operations whose shape someone will really read. Requiring a new distribution to state its expected series count in review keeps that line honest without discouraging counting.
Instrumentation is a recording rather than a query. You can crop, relabel and re-edit later, but you cannot go back and film a moment the camera was not pointed at.
saying these in an interview costs you the question
- Publishes an event count as a gauge and tries to rate it later
- Believes a missed spike can be reconstructed from stored samples
- Changes a metric's type in place while keeping the same name
- Instruments whichever value is handy rather than the question
- Assumes a mean from sum and count is enough for latency
- Treats adding a distribution as free because it is one line of code