Turnstile scans arrive faster than the counting consumer can take them — how do you choose the adapter's boundary policy?
answer
- ask what the consumer computes
- count, freshness, or cadence
- every value means a bound and an alarm
- newest wins means overwrite
- aggregating loses events, not information
basics
~20 sLet the consumer's need decide. If every event must be counted, hold a bounded buffer and fail when it fills, or aggregate in the adapter; if only the newest value matters, overwrite; if a steady cadence is enough, release the latest value on a clock.
solid answer
~50 sThe source will not slow, so something at the boundary must go — and which thing is decided by what the consumer is for, not by what is convenient. A consumer producing a total admissions count cannot tolerate a dropped scan, so the honest options are a bounded hold sized to the largest realistic burst with a loud failure past it, or folding the scans into a running count inside the adapter and emitting that when demand arrives. A consumer driving a current-occupancy display wants only the freshest value, so overwriting a single stored value is exact for its purpose and cheap. A consumer that samples a trend wants a regular cadence, which means releasing whatever the latest value is at fixed instants, even if it repeats. Choose before the pipeline exists, write the choice down, and count what it costs.
code
pseudocode · 21 lineson arrival(value):
if outstanding > 0:
emit(value)
outstanding = outstanding - 1
return
# no demand: the declared policy decides, and only here
if policy is BOUNDED_HOLD:
if size(hold) < limit:
append(hold, value)
else:
fail("boundary hold exceeded")
else if policy is AGGREGATE:
total = total + 1
else: # KEEP_LATEST and SAMPLE both store
latest = value
hasLatest = true
on cadenceTick(): # runs only when policy is SAMPLE
if outstanding > 0 and hasLatest:
emit(latest) # may repeat if nothing new arrived
outstanding = outstanding - 1go deeper
Know that when a source cannot be slowed, something at the boundary is dropped, merged or held, and that this is a deliberate choice rather than a bug.
Explain how the policies differ in what they keep: every value, the newest value, or a value per interval, and what each costs the consumer.
Derive the policy from what the consumer computes, size any bound from a real burst, and make discards and high-water marks visible in production.
Treat the loss question as a product decision: get an answer from whoever owns the number before designing, and refuse a policy that nobody has agreed to pay for.
## The choice is fixed by the consumer, not by the source When a source cannot honour demand, the adapter around it must decide what happens to values that arrive with no demand outstanding. Engineers reach for the wrong decision rule here: they look at the source ("it is fast, so I will buffer") when the only question that answers it is **what the consumer is computing**. Three consumers over the same gate feed want three different things. - A **total admissions count** is a sum over every event. One lost scan is a permanently wrong number, and no later value repairs it. - A **current occupancy display** is a function of the newest value only. Every intermediate value is already worthless by the time it could be drawn. - A **trend chart** is a function of values at regular instants. It wants a value per interval, not a value per event, and it does not care which event within the interval it got. ## The policies, and what each one actually preserves | Boundary policy | Preserves | Sacrifices | The consumer it fits | |---|---|---|---| | Bounded hold, fail on breach | Every value while the bound holds | Availability at the breach, loudly | A count that must be exact | | Fold into a running aggregate | The information, not the events | The individual values downstream | A total, a count, a maximum | | Overwrite the stored latest | The newest value at release time | Every value superseded before release | A current-state display | | Release the latest on a cadence | A value per interval, evenly spaced | Intra-interval detail; may repeat a value | A trend or a sampled reading | The last two are easy to conflate and are not the same policy. Overwriting releases at most one value per value that arrived, and it releases when demand appears — the consumer sets the pace. Releasing on a cadence emits at instants chosen by a clock, and may emit the same value twice if nothing new arrived, or skip several that did. If a downstream calculation assumes evenly spaced samples, only the second gives it that; if it assumes each emission is a distinct event, only the first. ## Sizing the bound is part of choosing it A bounded hold is the policy people pick when they cannot bring themselves to lose anything, and it is legitimate — but only if the bound is derived rather than guessed: 1. **Find the burst.** What is the largest arrival burst the source really produces — a full gate opening, a scheduled batch, a reconnection replay? 2. **Find the drain.** How fast does the consumer take values once it is running normally? 3. **Size for the burst that the drain can still clear** within the latency the consumer will tolerate, and treat a breach as a real failure signal rather than a nuisance. A bound nobody alarms on is a slower version of no bound: it converts a memory problem into a silent failure at an unpredictable moment. The virtue of the bounded hold is precisely that exceeding it is *news*. ## Decide before, not during The cost of deciding late is that the pipeline has already inherited a decision by default. A stage that quietly accumulates makes the choice "lose nothing until we lose everything"; an adapter that emits regardless makes the choice "let the slowest stage deal with it". Both are policies — just not ones anyone selected, and neither can be found by reading the pipeline's definition. Deciding up front also forces a useful conversation with whoever owns the consumer. "Can we drop scans during a surge?" has a real answer, and it is usually not the one an engineer would have assumed. Where the answer differs by mode — exact counts normally, freshness during a surge — the switch itself must be explicit and recorded, because a stream that silently changes from complete to lossy is worse than one that is always lossy: every downstream total is then valid only over intervals nobody can identify. ## Make the cost observable Whichever policy is chosen, the adapter should publish what the policy cost: values discarded, values overwritten, the high-water mark of the hold, the number of breaches. Those numbers are how anyone later answers "is this count trustworthy?" without reading the adapter's source. They are also what turns a policy from a code comment into an operational fact, and they are the first thing a good interviewer asks about after the policy itself.
- Can the adapter avoid the choice entirely by aggregating instead of dropping?Often, yes, and it is the option people forget. If the consumer needs a total rather than each event, the adapter folds arrivals into a running count and emits that when demand arrives: the emission rate is bounded by demand and the number is exact. The cost is that individual events are no longer recoverable downstream, so anything needing per-event detail must be served by a second path.
- What if the consumer needs exact counts normally but only freshness during a surge?Then make the switch explicit and record when it happened. A stream that silently changes from complete to lossy poisons every downstream total, because nobody can say which intervals are trustworthy. Better still, publish both: an exact aggregate computed in the adapter, and a freshest-value stream beside it, so the surge behaviour costs detail rather than correctness.
saying these in an interview costs you the question
- Picks keep-latest for a consumer that is counting events
- Sizes the boundary hold by guesswork and never alarms on it
- Believes a bounded hold removes the need for a loss decision
- Thinks releasing on a cadence and keeping the latest are one policy
- Chooses the policy only after symptoms appear in production