A metrics gauge counts entries by ranging a sync.Map — why is that count never exactly right?
answer
- the type offers no size method
- the walk never blocks other writers
- each key at most once, that is all
- a number mixing several moments
- count where membership actually changes
basics
~20 ssync.Map has no Len, and its Range takes no snapshot: it never blocks the map's other methods, so entries stored or deleted mid-walk may or may not be visited. The number mixes several moments — an estimate, not a count.
solid answer
~50 sThere is deliberately no `Len` on `sync.Map`, so counting means calling `Range` and incrementing a variable. `Range` guarantees only that no key is visited more than once; it explicitly does not correspond to any consistent snapshot, it does not block other methods, and if a key is stored or deleted concurrently — including by the function `Range` calls — the walk may reflect any mapping for that key from any point during the call. For a live-connection registry, that means the gauge counts some connections that closed mid-walk and misses some that opened, and the error grows with churn and with the size of the map. The fix is not a better `Range`: track membership where it changes. Increment an `atomic.Int64` only when `LoadOrStore` reports `loaded == false`, decrement it only when `LoadAndDelete` reports `loaded == true`, and publish that counter. `Range` stays for debug dumps, where an approximation is fine.
code
go · 7 linesn := 0
conns.Range(func(key, value any) bool {
n++
return true // returning false would stop the walk early
})
// n mixes several moments: entries added or removed during the
// walk may or may not have been visitedgo deeper
Remember two facts: sync.Map has no Len, and counting means walking it with Range. Know that returning false from the function stops the walk.
Explain the actual guarantee — each key at most once, no consistent snapshot, no blocking of other methods — and why that makes a derived count approximate.
Show how you would produce a trustworthy gauge: keep an atomic counter updated from the booleans LoadOrStore and LoadAndDelete return, funnel every mutation through one wrapper, and reserve Range for debug and shutdown paths.
Own the reliability question: decide which numbers a service publishes may be approximate and which must be exact, and make sure nothing downstream — an alert, a limit, an invoice — is quietly built on an estimate.
## Why there is no Len The missing `Len` is not an oversight. Ask what such a method could honestly return on a map that any number of goroutines may be mutating: a number that was true at some instant during the call and is potentially false by the time the caller sees it. Maintaining it would also mean a shared counter updated on every insert and delete — a contention point on the very operations `sync.Map` exists to keep cheap, imposed on every user whether they want a size or not. So the type simply does not offer one, and the API forces you to be explicit about what kind of count you actually want. ## What Range does and does not promise `Range(f func(key, value any) bool)` calls `f` for each entry, and stops early if `f` returns `false`. The guarantees are narrow and worth memorising: - **No key is visited more than once.** That is the strong guarantee. - **There is no consistent snapshot.** The walk does not correspond to the contents of the map at any single moment. - **Range does not block the map's other methods.** Other goroutines keep storing and deleting while you iterate, and `f` itself may call any method on the same map, including deleting the entry it was just handed. - **A key mutated during the walk may be observed in any state it held during the call** — the old value, the new value, or not at all if it was deleted before the walk reached it. This is a deliberate trade. A snapshot would require either copying the whole map or locking out writers for the duration of the walk, and both would defeat the point of the type. ## How the gauge goes wrong A registry of live client connections publishes `connections_live` by walking the map every fifteen seconds: ```go n := 0 conns.Range(func(key, value any) bool { n++ return true }) ``` On a quiet service this looks correct forever, which is what makes it dangerous. Under churn the walk takes real time, and during that time connections open and close. Some that closed early in the walk were already counted; some that opened late are not visited. The result is a number that is plausible, never exactly right, and biased in whichever direction traffic happens to be moving — so it drifts most precisely when you most want to trust it, during a spike or a mass disconnect. The second failure is cost. `Range` is O(entries), and it walks live structures rather than a copy. A gauge that used to walk two hundred entries and now walks two hundred thousand, on a fifteen-second timer, is suddenly a measurable slice of a CPU profile — and if someone wires the same walk into a request handler because "we need the current size", it becomes a per-request full scan. ## Counting where the change happens Membership changes at exactly two points, and both of those operations already report whether they changed anything: ```go var live atomic.Int64 if _, loaded := conns.LoadOrStore(id, meta); !loaded { live.Add(1) // this call actually inserted } if _, loaded := conns.LoadAndDelete(id); loaded { live.Add(-1) // this call actually removed } ``` The discipline is to route every mutation through those two helpers and never call bare `Store` or `Delete` on the map, because `Store` on an existing key must not move the counter and `Delete` of an absent key must not either — that is exactly the information the two `loaded` booleans carry and the bare methods throw away. Wrap the map and the counter in one small unexported struct so the invariant is enforced by the type rather than by review. The counter is itself a moving target — it is read without stopping the world, so it reports a moment, not a truth — but it is *consistent* with the sequence of insertions and removals, and it costs one atomic add per membership change instead of a full walk per scrape. ## Where Range is still the right tool `Range` earns its place where approximation is acceptable and a walk is what you actually want: a debug endpoint dumping every live entry, a shutdown path closing everything still registered, a sweeper that deletes entries older than a deadline (deleting from inside `f` is explicitly allowed), or a test where you can prove the map is quiescent. What it must not become is the source of a number some other system treats as exact — a billing count, a capacity check, a comparison against a hard limit. The rule of thumb: `Range` answers "roughly what is in here", never "how many are in here right now".
- What exactly does sync.Map's Range guarantee?That no key is visited more than once, and that iteration stops when the function returns false. It explicitly does not give a consistent snapshot, does not block the map's other methods, and for a key mutated during the walk may reflect any mapping it held at any point during the call. The function may itself call any method on the map.
- Where is a Range-derived count still acceptable?Wherever approximation is fine or you can prove quiescence: a debug dump, a shutdown sweep that closes whatever is still registered, a test after all producers have stopped, a rough log line. It stops being acceptable the moment another system treats the number as exact — a limit check, a billing figure, an alert threshold.
- Why hang the counter off LoadOrStore and LoadAndDelete rather than Store and Delete?Only those two report whether membership actually changed. Store on a key that already exists is an overwrite, not an insert, and Delete of a key that has already gone removes nothing — counting either would drift. The loaded booleans carry precisely the information the bare methods discard.
- Is walking the whole map on a timer a performance concern?It can be. Range is linear in the number of entries and walks live structures rather than a copy, so a registry that grew from hundreds to hundreds of thousands turns a cheap scrape into visible CPU time. The atomic counter replaces a full walk per scrape with one atomic add per membership change.
Counting heads by walking through a party while guests keep arriving and leaving by other doors: you will not count anyone twice, but the total belongs to no single instant.
saying these in an interview costs you the question
- Believes Range iterates a snapshot taken at the start
- Treats a Range count as an exact current size
- Assumes Range locks the map against other goroutines
- Asks for the missing Len as an obvious oversight
- Increments a counter on every Store including overwrites
- Calls a full Range from inside a request handler