skip to content

Why is an exact per-user counter a poor fit for a tier of many small edge locations?

level: seniorimportance: nice to knowfreq 30%

answer

  1. many locations, each a partial view
  2. exactness needs agreement
  3. agreement is a round trip
  4. local generous, central exact
  5. overshoot bounded by slack times sites

basics

~20 s

Each location sees only its own traffic, so an exact shared count needs the locations to agree - and agreeing costs the round trip the tier existed to avoid. Exact, durable counting belongs in a region.

solid answer

~40 s

The thin tier's shape is many independent locations each with a partial view. A counter that must be exact is shared state: every increment has to be visible to whoever reads it next, wherever that read lands. Getting that guarantee means coordinating across locations, and coordination is a round trip - exactly the cost you moved the work there to avoid. So the practical options are a local, per-location approximation that is fast and deliberately inexact, or an authoritative counter in a region that is exact and pays the crossing. Real designs usually combine them: a generous local check that sheds obvious excess cheaply, with the exact accounting done where the data durably lives.

go deeper

for a junior

Remember that each small location sees only its own traffic, so a count kept there is local. Anything that has to be exact across all users belongs where the authoritative record is.

for a middle

Explain why exactness over shared state needs the copies to agree, and that agreement costs a crossing - the very cost the thin tier was chosen to avoid. Name the three options and what each gives up.

for a senior

Demonstrate the combined design: a generous uncoordinated local check plus exact accounting centrally, with the overshoot bound stated explicitly rather than discovered downstream.

for a principal

Decide which rules may be approximate at all. A limit protecting a scarce resource can tolerate bounded overshoot; a figure promised to a customer cannot, and that distinction should be written down rather than left to each team.

## The shape that causes the problem The thin tier is wide and thin on purpose: many small locations, each close to some users, each with its own local view. A given user's requests do not necessarily all land at the same location, and no location is told what the others are seeing. That is fine for work answerable from the request itself. It is a genuine obstacle for **shared mutable state**, of which an exact counter is the smallest interesting example. ## Why exactness costs a round trip An exact count means: after an increment, any subsequent read anywhere returns a value that includes it. With one authoritative copy that is trivial - everyone asks the copy. With many copies it requires the copies to agree before the answer is given, and agreement between machines that are far apart costs at least one crossing between them. So the cost cannot be designed away, only moved: - **Ask a central authority on every request.** Exact, and pays the crossing on every request - which removes the reason for doing the work near the user at all. - **Keep a local copy and reconcile later.** Fast, and inexact by construction during the window before reconciliation. - **Partition the allowance.** Give each location a slice of the budget up front; each enforces its slice locally with no coordination. Exact against the slice, approximate against the global total, and wasteful when traffic is uneven. | Approach | Exact? | Cost paid | |---|---|---| | Central authority consulted per request | Yes | A crossing on every request | | Local copy with later reconciliation | No, within a window | Overshoot and a reconciliation path | | Pre-partitioned allowance per location | Against the slice only | Unused slices where traffic is uneven | ## What this rules in and out Ruled out at the thin tier: an exact balance, a strictly enforced global limit, a uniqueness guarantee, a lock, an audit-grade count, and anything whose acknowledgement must survive the loss of the machine that gave it. Ruled in: decisions that tolerate approximation. Shedding obviously excessive traffic, refusing a caller already known locally to be abusive, holding a small read-mostly rule set, and returning a cheap answer while the authoritative accounting happens elsewhere. The distinguishing question is not `is this state?` but **`what does being wrong for a short while cost?`** ## The combination that is usually right 1. Enforce a **generous** local threshold at each location, with no coordination. This sheds the bulk cheaply, near the user, and is allowed to overshoot. 2. Enforce the **exact** limit in the region, where the authoritative record lives and where the acknowledgement is durable. 3. Accept explicitly that a caller may briefly exceed the global figure by roughly the local slack multiplied by the number of locations they reach, and decide whether that is acceptable for this particular rule. That third step is where the design is actually made. If the limit protects a scarce downstream resource, a bounded overshoot is often fine. If it is a commitment to a customer about exactly how much they may consume, it is not, and the check belongs where the truth is even though that costs the crossing. ## The failure mode to name The common mistake is assuming that because the local check is fast and returns a number, the number is the global one. Traffic then arrives spread across many locations, each of them well under its own threshold, and the aggregate sails past the figure the rule was supposed to enforce - with every location's own logs showing it behaving correctly. Nothing looks broken anywhere, which is precisely why the overshoot is discovered downstream. ## A quick sanity question Before placing any counter at the thin tier, ask what it costs to be wrong for a few seconds in every location at once. If the answer is a slightly noisy graph, the local copy is fine and the approximation is the right trade. If the answer is a bill, a broken promise to a customer or a downstream system overwhelmed, the count belongs where the authoritative record is, and the crossing is simply the price of the guarantee. ## Saying it well State the property, not the anecdote: exactness over shared state requires coordination, coordination costs a crossing, and a crossing is the thing the thin tier exists to avoid. Then name the trade you chose and the overshoot you accepted. An answer that claims you can have wide proximity and an exact global count with no coordination has simply not said where the round trip went.

  • How large can the overshoot be with a purely local check at each location?
    Roughly the per-location slack multiplied by the number of locations a caller manages to reach, plus whatever arrives during a reconciliation window. Because the bound scales with the width of the tier, a rule that matters should never rely on local checks alone - state the bound explicitly and decide whether it is tolerable.
  • Does pre-partitioning the allowance across locations solve it?
    It removes the coordination but not the inexactness. Each location enforces its own slice exactly, so the global figure is only respected if traffic arrives in the proportions you assumed. Uneven traffic means some slices are exhausted while others sit unused, rejecting callers who were globally within budget.
  • What kinds of state are genuinely comfortable at the thin tier?
    Small, read-mostly material that is the same everywhere and changes rarely - verification keys, a compact rule set, a shared response held for a while. Anything per-user, per-object, exact or durable belongs where the authoritative record lives, because that is the only place a guarantee can be given about it.

It is a noticeboard in every branch office. Fine for posting what this branch has seen, useless as the single ledger - to know the company-wide total, someone still has to phone head office.

saying these in an interview costs you the question

  • Assumes an increment at one location is visible at the others
  • Believes later reconciliation makes an approximate limit exact
  • Treats coordination between locations as free because the locations are near users
  • Describes best-effort local state as durable storage
  • Claims exact global enforcement with no round trip anywhere