Two teams' stream pipelines have now died from silent buffering; what would you require of every pipeline before it ships?
answer
- make the failure a design-time decision
- declare a bound, choose a policy
- age bound usually beats memory bound
- alert on depth that never recovers
- prove it with a deliberate deficit soak
basics
~20 sRequire four things: no unbounded retention anywhere, a capacity derived from the freshness budget and checked against memory, depth and delivery age exported per retaining stage, and a soak at a deliberate deficit proving memory stays flat.
solid answer
~50 sThe standard is that unlimited retention is not a default a team may take silently. Every stage that can hold items declares a capacity and an explicit policy for being full — which policy is the team's call, since it depends on what stale data is worth. Capacity is derived from the delivery-age budget and then checked against memory, with the smaller bound winning, rather than sized from spare heap. Every retaining stage exports depth and delivery age, and the alert is on depth that never returns to baseline, not on an absolute number. Before shipping, the team demonstrates a soak run under a deliberate sustained deficit: flat memory, a bounded delivery age, and the full-hold policy actually firing. What you accept in exchange is that pipelines now fail earlier and more visibly, and somebody has to be ready to act on that signal.
go deeper
Take away the habit rather than the policy: when you add a stage that holds items, write down how many it can hold and what happens when it is full.
Be able to produce the sizing table — item size, spare memory, drain rate, freshness budget — and explain why the tighter of the memory bound and the age bound is the capacity you declare.
Argue for the evidence rather than the rule: a soak that actually reaches the bound, depth and delivery age exported per stage, and an alert keyed on a backlog that never recovers.
Own the trade you are imposing. Bounding converts rare catastrophic outages into frequent visible events, so the standard only works if it ships with an agreement about who responds to those events and what they are allowed to discard.
## What the standard has to guarantee After two outages with the same root cause, the goal is not to ban a construct. It is to make one class of failure impossible to reach by accident: **a pipeline whose memory is bounded only by the machine, discovered at the moment the machine runs out**. Everything in the standard should serve that, and anything that does not is process for its own sake. The framing that makes it enforceable: retention is a **budget**, and a budget that nobody wrote down is a budget nobody can review. ## The four requirements 1. **No unbounded retention by construction.** Every stage that can hold items declares a capacity and what happens when it is full. The choice of what to sacrifice when full is deliberately left to the team, because it depends on whether a stale item is still worth anything — but making *some* choice is mandatory. The value of the rule is not the number chosen; it is that a human chose it at design time instead of the allocator choosing at 3 a.m. 2. **Capacity derived from the freshness budget, checked against memory.** Two bounds apply and the smaller wins: `capacity x bytes per item` must fit the memory the process can spare, and `capacity / drain rate` must be under the delivery age consumers need. In practice the age bound is usually far tighter, which is why sizing from spare memory produces pipelines that survive and deliver hours-old data. 3. **Depth and delivery age exported per retaining stage.** Both, not one. Depth without a drain rate cannot be interpreted; delivery age states the harm directly. The alert fires on **depth that fails to return to baseline across successive periods**, because a burst pushing depth up is normal and an absolute threshold either fires constantly or never. 4. **Evidence before the first deploy.** A soak run with a deliberately induced sustained deficit, showing flat memory, a bounded delivery age that matches the declared capacity divided by the drain rate, and the full-hold policy actually firing. A run that never reached the bound proves nothing about what happens at the bound. ## Deriving capacity, worked | Input | Value | Bound it gives | |---|---|---| | Retained bytes per item | 400 bytes | — | | Memory the process can spare | 3 GB | about 7,500,000 items | | Drain rate of the slowest stage | 11,500 items per second | — | | Delivery age the consumers need | 60 seconds | 690,000 items | | **Declared capacity** | | **690,000 — the tighter bound** | The table is the artefact the review asks for. It is five numbers, it fits in a pull request description, and producing it forces the team to know its item size, its drain rate and who consumes the output — three things that are usually unknown in exactly the pipelines that fail. ## The audit that goes with it Declared holds are the easy half. The standard also asks each team to list its **implicit** retention points and say what bounds each: a time-based grouping stage bounded by input rate times window length, a flattening stage bounded by its concurrency limit or by nothing at all, a de-duplicating stage bounded by the cardinality of its key space, a handoff between workers bounded by its queue capacity. Any line of that list whose bound is proportional to the input rate has to be converted to a constant before the pipeline ships. ## What you deliberately do not mandate - **One overflow policy for everyone.** Whether to keep the oldest, the newest, or to fail depends entirely on what the data is for, and a central mandate would be wrong in half the cases. - **One capacity number.** Rates and item sizes differ per pipeline; a shared constant is a number that is wrong everywhere. - **A specific implementation.** The standard is about declared bounds and exported signals, which every reasonable implementation can provide. ## The cost you are accepting Bounded pipelines fail **more often and earlier**. A condition that used to be invisible for four hours now produces an explicit event within minutes, sometimes including deliberate loss. That is the trade: you are converting rare catastrophic failures into frequent small visible ones, and it is only an improvement if the organisation is prepared to act on the small ones instead of silencing the alert. Say that out loud when you introduce the standard, and pair it with the operational agreement about who responds and how, or the new signal will be tuned out and the next collapse will look exactly like the last two.
- Why not simply mandate one full-hold policy across all teams?Because the right sacrifice depends on what the data is worth. A pipeline where only the newest state matters and one where every record must survive want opposite behaviour, and a central rule would be wrong for half of them. Mandate that a policy is chosen and documented, not which one.
- What evidence would you reject as insufficient?A soak that stayed under the declared bound. It shows the pipeline works when nothing is wrong, which was never in doubt. The run must reach the bound so that the full-hold behaviour, the alerting and the recovery after the deficit ends are all observed rather than assumed.
- How do you keep the standard from decaying as pipelines are edited?Tie it to the exported signals rather than to the document. If every retaining stage must publish depth and delivery age, a new stage without them is visible in the dashboard and in review, and the retention table becomes something a reviewer can check against reality instead of trusting.
saying these in an interview costs you the question
- Setting one organisation-wide capacity number for all pipelines
- Sizing capacity from spare memory and ignoring delivery age
- Accepting a soak run that never reached the declared bound
- Alerting on an absolute depth threshold instead of failure to recover
- Assuming bounding is free rather than trading rare outages for frequent alerts