A stored number must never exceed a cap, but the store's arithmetic is unconditional; how do you enforce the bound?
answer
- the arithmetic knows nothing about your cap
- apply first, then inspect the reply
- the bound is breached transiently
- a crash before the undo leaks capacity
- money and stock do not live here
basics
~20 sApply the change first, inspect the total the operation hands back, and undo it if it went over. The bound is therefore breached transiently, concurrent readers can see the overshoot, and a caller that dies before undoing leaks capacity.
solid answer
~50 sThe arithmetic a store offers knows nothing about your maximum, so checking before changing is exactly the read-then-write that loses changes. The workable shape is **apply, then inspect**: add one, look at the number handed back, and if it is above the cap, apply the opposite change and reject the request. Three consequences follow, and an interviewer is listening for them. The number **goes over the bound transiently**, so nothing may treat the raw value as truth. If the process dies between the change and the undo, that unit is **leaked** until something reconciles it or the entry's lifetime ends. And the undo is unconditional too, so a retried undo takes the number below where it belongs. Where the bound must hold even transiently, or where losing the number is unacceptable, this tier is the wrong home for the authority — it can drop the entry, and the bound resets with it.
go deeper
Recall that the store's arithmetic has no notion of a maximum: it applies whatever amount it is given, so any cap has to be enforced by the code around it.
Explain why checking before changing fails under concurrency, and describe applying the change first and deciding on the number the operation hands back.
Volunteer the costs unprompted: a bound that is exceeded for one round trip, capacity leaked when a caller dies before undoing, and an undo that subtracts twice on retry.
Decide what the bound is allowed to cost. Say which limits may be enforced approximately on a tier that can lose the number, and which must be held where they can be recovered.
## Why the obvious answer is the wrong one "Read the number, compare it to the cap, and if there is room, increase it" is the answer that fails an interview, because the gap between the read and the change is exactly the window in which another caller does the same thing. Both see room, both proceed, and the total ends above the cap with neither caller having done anything visibly wrong. The bound is not enforced by the check; it is enforced by whatever makes the check and the change one thing. The server's arithmetic is **unconditional**. It applies the amount to whatever is there. It does not know your maximum, and asking it to stop at one is not part of the shape. ## The apply-then-inspect shape 1. Apply the change — add one. 2. Read the resulting total out of the operation's own reply, where the store hands it back. 3. If the total is within the bound, proceed. 4. If it is over, apply the opposite change to undo, and reject the request. This works because the decision is made on a number **no other caller can have influenced between the change and the reply** — it is this caller's own result. Two callers racing at the cap both apply, both get a number back, and they get different numbers; the one over the line is the one that backs out. ## What you are actually signing up for - **A transient breach.** For the duration of one round trip, the stored number is above the cap. Anything that reads the raw number — a dashboard, a second service, an operator — can observe a value that was never supposed to exist. Nothing may treat the raw number as the truth about how much has been consumed. - **A leak on failure.** If the caller crashes, times out or is killed between the change and the undo, the unit is consumed and never released. Capacity erodes silently, and the erosion is permanent unless something reconciles it or the entry's lifetime ends and takes the whole number with it. - **An undo that is itself unconditional.** A retried or duplicated undo subtracts again and pushes the number below the real consumption, handing out capacity that does not exist. The undo path needs the same care as the request path. - **A number that can vanish.** This is a volatile tier. Losing the entry resets consumption to zero, which is the most generous possible failure — every caller is suddenly under the cap. ## The mirror shape: count down from the cap Seed the entry with the cap and subtract, rejecting when the number handed back goes below zero. The mechanics are identical, and the differences are ergonomic: | | count up to a cap | count down from the cap | |---|---|---| | The test | is the returned total above the cap? | is the returned number below zero? | | Setup | none needed where a missing key starts at zero | the seeding write is load-bearing | | A lost entry means | consumption forgotten, everyone is under the cap | capacity forgotten, and a missing entry may read as none left | | Reading remaining capacity | subtract from the cap | read it directly | The seeding step is the trap in the second shape. Where a missing key is created at zero, a counter that was supposed to be seeded and was not reads as no capacity at all — a hard outage rather than an over-issue. ## Moving the check to the server Both shapes above put the decision in the caller, one round trip after the change. The alternative is to make the conditional part run where the value is, so the number never goes over at all: either a compare-and-set loop that re-reads and retries when another caller got in first, or logic the store executes for you as one unit. Both exist in this class and what a given store offers varies, so name them as the direction rather than as a specific mechanism. The trade is real: the number is never transiently wrong, and you pay in retries under contention or in operational weight. ## Where this shape does not belong at all Ask what the bound protects. If it protects something recoverable — seats shown as available, a soft allowance, work in flight — an occasional over-issue is a business decision and this design is fine. If it protects money, physical stock or anything that cannot be un-issued, then a tier that can evict the entry, lose it on a restart and briefly exceed the bound by design is not where the authority lives. Keep the authoritative number where it can be recovered, and let this tier hold the fast, advisory copy in front of it. ## Answering it well Reject the check-then-change answer and say why in terms of the window. Give apply-then-inspect. Then volunteer the three things it costs — the transient breach, the leak on a crash, the unconditional undo — and finish by naming the class of bound that should not be enforced here at all.
- Why is the number handed back safe to decide on, when the stored number is changing constantly?Because it is the result of this caller's own change, produced inside the operation the server ran to completion. No other caller can have slipped between the change and its reply. Two callers racing at the cap therefore get different numbers, and exactly the ones above the line back out. What is unsafe is a number obtained by a separate read, which is stale the instant it arrives.
- How do you stop a leaked unit from eroding capacity forever?Give the counter a lifetime so the whole number resets on a known cadence, if the bound is per-period; or reconcile it against an authoritative source that knows what was really consumed. A third option is to stop leaking: make the undo idempotent by keying it to the request, so a retry cannot subtract twice and a missed undo can be replayed safely.
- Does counting down from the cap avoid the transient breach?No, it relocates it. The number goes below zero for one round trip in exactly the same way, and the caller that sees a negative result is the one that backs out. What changes is the reading: remaining capacity is directly visible, and the failure mode inverts, because a counter that was never seeded reads as nothing left rather than as nothing used.
saying these in an interview costs you the question
- Reads the total, compares it to the cap, then applies the change
- Assumes the undo always runs, so the cap is never really exceeded
- Calls the transient overshoot harmless without asking who reads the number
- Expects the store's arithmetic to accept a ceiling
- Keeps an authoritative money or stock limit on a tier that can evict it
- Retries the undo without making it safe to apply twice