In Micrometer, the same counter can be registered from several places without creating duplicate series, a counter on a step-based export registry can read as zero moments after you incremented it, and an export that fails can lose a whole interval. Explain the registry mechanics behind all three: how Micrometer decides that two registrations refer to the same meter, what accumulate-and-reset means for step meters and which interval a step meter's reported value refers to, and what a step-based registry does when a publish fails.
answer
- Meter.Id = name + normalised tag set; re-register returns the existing meter
- same id, different meter type → rejected, not duplicated
- step meter reports the LAST COMPLETED interval, then resets
- publish schedule derived from step — no second interval knob; failed batch dropped
- SimpleMeterRegistry: in-memory, cumulative, plus a mock clock for roll-over
basics
~20 sRegistration is idempotent: the same name plus tag set is the same Meter.Id, so you get the existing meter back. Step meters report the last completed interval and then reset. Publishing runs on that same step interval, and a failed batch is dropped, not retried.
solid answer
~50 sThree mechanics explain most Micrometer surprises. **Identity.** A meter is keyed by its `Meter.Id` — the name plus the tag set, with tags normalised (sorted, de-duplicated by key), so tag order never splits a series. Registering the same id again returns the *existing* meter rather than creating a second one, which is why calling `Counter.builder(...).register(registry)` repeatedly is safe. Same name and tags with a different meter type is rejected outright, not silently duplicated. **Step semantics.** On a step-based export registry, a counter or timer accumulates into the current interval and rolls over at each step boundary; its reported value is the *last completed* interval, not the in-progress one. So an increment made a millisecond ago is invisible until the roll-over — that is normal, not a lost write. **Publishing.** The publish schedule is derived from the configured step, so the two cannot diverge. Export is best-effort: a failed batch is logged and dropped, leaving a permanent hole for that interval.
code
text · 11 linesstep = 60s
t=0s ---- interval A opens
t=5s increment -> A accumulator = 1
t=6s read count -> 0 (reports previous interval, still empty)
t=50s increment -> A accumulator = 2
t=60s ---- boundary: A closes (=2), accumulator resets, B opens
publish sends A = 2
t=61s read count -> 2 (A, the last completed interval)
t=70s increment -> B accumulator = 1
t=71s read count -> 2 (still A; B not closed yet)go deeper
Recall the two rules that stop real bugs: the same name plus the same tags is the same meter, so registering twice is harmless; and on a push-style registry a counter shows the last completed interval, so a zero right after an increment is expected.
Explain the Meter.Id key including tag normalisation, and describe accumulate-reset roll-over — what the accumulator holds versus what the meter reports. Know that the publish schedule comes from the step and that a failed batch is dropped.
Diagnose with it: separate an identity mismatch from step semantics when a value reads wrong, know that an export outage leaves a permanent hole rather than a delayed spike, and use an in-memory registry plus a mock clock to test roll-over honestly.
Own the fleet-level call: one step interval everywhere, chosen to match the backend's aggregation window so rates are comparable across services, with the accepted consequence that export is best-effort and lost intervals must be detected via the exporter's own signals or absorbed by a buffering collector rather than by the application.
## What a registry actually stores A `MeterRegistry` is a map from meter identity to meter state, plus a strategy for getting that state to a backend. Everything surprising about Micrometer at runtime comes from one of three properties of that map: how identity is computed, how the stored value moves through time on step-based registries, and what happens on the way out. ## Identity: registration is idempotent by `Meter.Id` Every meter is keyed by a `Meter.Id`, which is the meter **name** plus its **tag set**. Tags are normalised — sorted and de-duplicated by key — before the id is formed, so `tags("a","1","b","2")` and `tags("b","2","a","1")` are the *same* meter, and tag order can never accidentally create two series. The consequence is that registration is **idempotent**: asking the registry for a counter that already exists returns the existing instance instead of creating a second one. This is what makes it safe for several components — your code, a library, an auto-instrumentation agent — to register the same meter independently, and it is why the builder-then-register call on a warm path is correct rather than a leak. (Caching the returned handle in a field is still the conventional style; it avoids the lookup, not a correctness problem.) Two edges follow from the same rule. First, if a meter with that id already exists but is a **different meter type**, the registry refuses the registration with an error rather than creating a parallel series — one name cannot be both a counter and a timer. Second, for meters that are bound to an object at registration time, re-registering the same id does **not** rebind: the registry hands back the meter that is already there and quietly ignores the new state object you passed. A second registration pointing at a fresh object therefore keeps reporting the first one, which is a classic "my gauge is stuck" bug. Duplicate series in a backend are almost never caused by repeated registration; they are caused by *different* tag sets, i.e. different ids. ## Step meters: accumulate, roll over, reset Export registries that push data to a backend are **step-based**. A step counter does not hold a lifetime total. It accumulates recordings into the current interval, and at each step boundary the current interval becomes the previous interval and the accumulator resets to zero. The part that catches people out is which number is visible: **a step meter reports the last completed interval**. Increment a step counter and read its count immediately and you see the previous interval's value — often zero on a freshly started process. Nothing was lost; the in-progress interval is simply not observable until it closes. Cumulative registries behave differently: an in-memory or scrape-rendered registry can carry a monotonically increasing total, so the same test written against a cumulative registry passes and against a step registry fails. ## The publish schedule *is* the step A push/step registry schedules its own publish, and that schedule is derived from the configured step interval — they are the same value by construction, and there is no second "publish every N seconds" knob to set against it. So the misconfiguration worth teaching is **not** two intervals drifting apart. It is: - a step that **disagrees with the backend's own aggregation window**, so intervals land unevenly inside the backend's buckets and rates look spiky or diluted; and - **different steps across services**, which makes their rates non-comparable and any cross-service dashboard subtly wrong. Pick one step for the fleet, and pick it to match what the backend aggregates on. ## Export is best-effort A publish is a fire-and-forget attempt. Failures are caught and logged so the scheduler survives, and the batch is **dropped** — there is no durable queue, no retry of the previous interval on the next tick. Because a step meter has already reset, that interval is gone permanently: you get a hole, not a delayed spike. Treat the exporter's own failure counters and logs as first-class signals, and remember that a clean shutdown gives the registry a chance to flush a final publish while a hard kill does not. ## Testing against these rules `SimpleMeterRegistry` is the in-memory registry with no export at all — the right tool for unit tests. It counts cumulatively by default, so counts do not vanish under you mid-test. Combine it with the library's mock clock when you specifically want to exercise step behaviour: advance the clock past a step boundary and then assert, which is the only honest way to test roll-over. Reading values back through the registry's search API also surfaces identity mistakes immediately — a lookup that fails to find a meter with the tags you expected is usually telling you that a tag typo produced a different `Meter.Id`.
- Two services in one fleet are configured with different step intervals for their metrics export. What actually goes wrong?Nothing errors, which is the problem. Each service's counters are aggregated over a different window before they leave the process, so the two series carry different implicit resolutions and any rate or ratio computed across them compares unlike quantities. Dashboards and alert thresholds tuned on one service read as spikier or flatter on the other. Worse, if a step does not line up with the backend's own aggregation window, individual intervals land unevenly inside the backend's buckets, producing periodic dips that look like real load variation. Standardise one step across the fleet and choose it to match the backend's aggregation.
- A publish to the metrics backend fails because the endpoint is down for two minutes. What does the backend see afterwards, and what should you do about it?It sees a hole. Export is best-effort: the failed batch is logged and dropped, and because a step meter already reset at the boundary, that interval's data no longer exists in the process to resend. There is no durable buffer and no replay on the next tick, so when the endpoint recovers you get the next interval's values, not a catch-up spike. Mitigations live outside the meter API — monitor the exporter's own failure signal so you know the gap is an export failure rather than an idle service, and put a collector or gateway in front if you need buffering.
- A colleague registers a gauge, later re-registers a gauge with the same name and tags pointing at a new object, and the reported value never changes. Why?Because registration is idempotent by Meter.Id. The registry already has a meter under that name and tag set, so the second call returns the existing meter and the new state object is simply not used — the gauge keeps observing the original object. The fix is to give the new instance a distinguishing tag if it is genuinely a different thing, or to remove the old meter from the registry before re-registering if it is a replacement.
- You increment a counter in a unit test and the assertion on its count fails with zero. How do you tell an identity bug from a step-semantics bug?Look up the meter through the registry's search API by name and tags. If the lookup itself fails, it is an identity problem — the tags you asserted with do not match the tags you registered with, so you are querying a different Meter.Id. If the lookup succeeds but the value is zero, it is step semantics: you are on a step-based registry reading the last completed interval. Use the in-memory registry, which counts cumulatively, for ordinary assertions, and drive a mock clock past a step boundary when the roll-over itself is what you are testing.
A step meter is a shift tally board: you chalk marks during the shift, but the number posted for everyone to read is last shift's total. Ask mid-shift and you get the old number — nothing was lost, the shift just hasn't ended.
saying these in an interview costs you the question
- Claiming the publish interval and the step interval are separate knobs that can be tuned against each other — a step registry derives its publish schedule from the step, so they cannot diverge.
- Saying failed publishes are buffered and retried, so no data is lost. They are dropped, and the step has already reset, so the interval is gone.
- Believing that registering the same counter from two places creates duplicate series. Registration is idempotent by Meter.Id; duplicates come from different tag sets.
- Treating a step counter's reported value as the lifetime total, and calling a zero read immediately after an increment a bug rather than the last-completed-interval rule.
- Assuming a re-registration with the same name and tags rebinds the meter to the new object it was given — the registry returns the pre-existing meter and ignores it.
- Thinking the in-memory test registry exports somewhere, or that its cumulative counting is how the production export registry behaves.