You must set one clock-driven reduction rule for hundreds of heterogeneous panel feeds - how would you decide and defend it?
answer
- publish a bound, not an operator
- measure gaps before choosing
- one rule will hurt half the fleet
- replay recorded traces offline
- alert on delivered-value age
basics
~20 sDecide from measured gap distributions and the consumer's freshness need, not from taste. Classify feeds into bursty and continuous, give each class a default reduction, and publish a staleness bound as the contract so the operator choice can change later.
solid answer
~40 sThe honest answer starts by refusing the premise: one rule across heterogeneous feeds will starve the continuous ones or alias the bursty ones. Measure the distribution of gaps between values per feed, and split on whether it has a valley - a bursty feed gets a quiet-period reduction with a maximum wait, a continuous feed gets periodic sampling, and a feed whose consumer must see every value gets neither. Then publish the promise as a bound on how stale a delivered value may be, plus whether intermediate values may be dropped, rather than as an operator name. That way each class can be tuned, and the guarantee stays testable: alert on output age, replay recorded traces before changing a default, and give teams a documented escape hatch with the evidence they must bring.
go deeper
Note the shape of the answer: the choice depends on how the source behaves and what the consumer needs, so measurement comes before any operator name.
Be able to map each feed shape to a reduction and say what each one discards, since that is the raw material the policy is built from.
Show the operational side: recorded traces replayed offline, delivered-value age as the alert, and the failure mode you are explicitly accepting.
Own the contract and its reversal condition - the classes, their defaults, the evidence an exception must bring, and what would move a feed between classes.
## What the decision really is The request is phrased as a choice of operator, but the durable artefact is a **contract**. Operators get replaced; the promise the consumer builds on should not. So the output of this exercise is two numbers and one flag per feed class: - the **maximum age** of a delivered value (how stale the freshest thing downstream may be), - the **maximum output rate** the consumer must absorb, - whether **intermediate values may be dropped**. Every clock-driven reduction in this leaf is then simply an implementation that satisfies some triple. A quiet-period stage with no ceiling satisfies none of them, which is by itself an argument against it as a default. ## The evidence to collect before choosing 1. **The gap distribution per feed**, not the mean rate. The mean hides exactly the property that decides the choice: whether there is a valley between short in-burst gaps and long between-burst gaps. 2. **The consumer's real need**, stated as one of three: the settled value after activity stops, a recent value at a predictable rate, or every distinct value the source passed through. 3. **The cost of being wrong in each direction** - a starved feed that shows nothing during heavy use, versus an aliased feed that silently misses short-lived excursions. 4. **The spread across feeds.** If the distributions cluster, one default is defensible; if they are bimodal across the fleet, insisting on one rule is a decision to hurt half of them. ## Classify, then default | feed class | evidence | default reduction | what it commits you to | |---|---|---|---| | bursty with clear quiet gaps | gap histogram has a valley | quiet period sized into the valley, with a maximum wait | one value per burst, and a stated ceiling on how long a value is held | | continuous, never silent | gaps cluster below any usable window | periodic sampling at a fixed interval | bounded staleness, and the loss of anything between two ticks | | low rate, each value meaningful | gaps already long | no reduction; pass through | full fidelity, and a consumer that must handle bursts | | needs every value, at a manageable size | any | grouping by size and age together | amortised per-item cost, at the price of the age bound in latency | A policy of three defaults plus a documented escape hatch beats one rule, and beats "teams choose" - the latter produces a fleet nobody can reason about. ## Defending it - **Replay, do not argue.** Record real traces from a sample of feeds and run the candidate defaults over them offline. Report, per feed, the output rate, the distribution of delivered-value age, and the fraction of source changes that never appeared downstream. Those three numbers settle most of the debate. - **Alert on age, not rate.** An output-rate alert cannot tell a quiet source from a starved stage. An age alert fires on exactly the condition that harms the consumer, and it is the same metric the contract is written in. - **State the failure you are accepting.** Sampling accepts invisible short-lived excursions. A quiet period with a ceiling accepts that the consumer sometimes sees a mid-burst value rather than the settled one. Pretending a reduction is lossless is the thing that gets discovered in an incident. - **Write the reversal condition.** Name in advance what evidence would move a feed to another class - for example, a gap histogram that loses its valley, which turns a bursty feed into a continuous one and makes the quiet-period default wrong for it. ## What not to fold into this decision Two neighbouring concerns arrive uninvited and should be kept separate, because they have different owners and different failure modes: - **A consumer that cannot keep up** is a flow-control question, answered by what the pipeline does when demand runs out, not by which clock stage sits in front of it. A clock reduction happens to lower the rate, which makes it a tempting and unreliable substitute. - **Protecting a service from excess inbound load** is an admission-control question about rejecting or delaying work at the boundary, and it lives in the architecture layer rather than in a pipeline stage. Keeping those out makes the reduction policy testable on its own terms: it is about the freshness contract for values already accepted, and nothing else.
- What three numbers would you report from replaying candidate defaults over recorded traces?Output rate per feed, the distribution of delivered-value age, and the fraction of source changes that never appeared downstream. The first tells the consumer what it must absorb, the second is the contract itself, and the third quantifies the fidelity you are giving up. Together they usually settle the argument without appeal to taste.
- Why alert on output age rather than output rate?Because a low output rate is ambiguous: a quiet source and a starved stage look the same. Age fires on exactly the condition that harms the consumer - the freshest value downstream being too old - and it is measured in the same unit the published contract is written in, so the alert and the promise cannot drift apart.
- A team asks to opt out of the default for their feed - what do you require?The gap distribution for that feed, the consumer need stated as settled value, recent value or every value, and the replay numbers for both the default and their proposal. If the evidence shows their feed is in a different class, the right outcome is usually a new class with its own default rather than a one-off exception.
saying these in an interview costs you the question
- Picks one reduction for the whole fleet without measuring anything
- Uses mean event rate, which hides whether bursts exist
- Publishes an operator name as the contract instead of a bound
- Claims the chosen reduction loses nothing
- Reaches for a clock reduction as a substitute for flow control