skip to content

Across an estate of streams, what should a standing policy say about who may discard a record to unblock a frozen reader, and what must be captured first?

level: principalimportance: nice to knowfreq 34%

answer

  1. the decision arrives with a clock
  2. content, not severity, sets the rule
  3. pre-waive the reversible exits
  4. capture first, discard second
  5. refusing to discard also has a bill

basics

~20 s

Settle it before the incident: classify each stream by tolerable loss, pre-authorise the cheap exits, require the record to be captured durably before any discard, name an owner for anything set aside, and set a deadline before the retained window decides for you.

solid answer

~50 s

Every fast way out of a frozen reader loses something, so the decision is about authority and evidence rather than mechanism. A workable policy fixes four things in advance. First, **classification by stream**: what a stream carries decides whether a record may be discarded at all, and that attribute belongs with the stream's definition, not with whoever is on call. Second, **pre-authorisation of the cheap exits** — setting a record aside, restarting a reader, adding capacity for a genuinely slow one — so nobody waits for a signature to do something reversible. Third, **capture before discard**, without exception: identity and content written somewhere durable, or the loss is undescribable later. Fourth, **an owner and a deadline** for anything set aside, because deferral without a name is loss with a delay. Note also that a blanket refusal to discard is not a safe default: at the edge of the retained window it trades one known record for everything behind it.

go deeper

for a junior

Recall that discarding a record to restore flow is a decision with consequences, and that teams agree in advance who is allowed to make it rather than improvising during an incident.

for a middle

Explain why capture must precede any discard: without the identity and content written down, nobody can later say which work went missing or reconstruct it.

for a senior

Show that you would arrive at an incident already knowing the stream's loss classification and the retained-window deadline, and that you would pre-waive the reversible exits so containment never waits on approval.

for a principal

Argue the trade explicitly: a bounded captured loss against the unbounded loss of a whole backlog at the retention boundary, and make the classification an attribute of the stream so the estate answers consistently.

## Why this cannot be settled at three in the morning The person in front of a frozen reader has a share that is not moving, a backlog that is ageing, and a menu in which the fast options destroy or reorder work. They are also, structurally, the person least able to judge what the record is worth: they know the stream's name and not its meaning. Deciding authority during the incident produces one of two failures — an engineer who discards something that mattered, or an engineer who waits for approval until the retained window makes the decision for everyone. A policy here is not ceremony. It is a pre-computed answer to one question: *what is this team permitted to lose, and what must exist afterwards to prove what was lost?* ## What the policy must fix in advance 1. **Which streams may lose a record at all.** This is a property of content, not of incident severity, and it should be recorded wherever the stream is defined so it is visible before anyone needs it. 2. **Who may authorise a discard for each class.** For a tolerant stream, the on-call engineer. For an intolerant one, a named owner who can be reached, with an explicit fallback for when they cannot be. 3. **What must be captured first.** Identity and content, written durably, before anything irreversible happens. This is the clause that turns a silent loss into a repairable one. 4. **Which exits need no authorisation.** Setting a record aside, restarting a reader, adding readers for a genuinely slow member — reversible actions should be pre-waived so nobody improvises a discard because the approval path was the only path available. 5. **Who owns what was set aside, and by when.** A side destination with no owner and no deadline is a discard that took longer and generated more paperwork. 6. **The deadline the retained window imposes.** Every deferred decision expires when the oldest unread record does. The policy should say so, and the incident should carry that time as a fact rather than a surprise. ## Classify the streams, not the incidents | what the stream carries | may a record be discarded? | what must happen first | |---|---|---| | sampled telemetry and metrics feeds | yes, by the on-call engineer | note the volume discarded and the reason | | derived or recomputable projections | yes, if the source can rebuild it | record the range so it can be rebuilt | | customer-visible business events | only with the named owner's decision | capture the record, and notify the owner | | financial instructions and audit records | no, at any hour | set aside, escalate, and treat the window as a hard deadline | The table is the artefact that makes the rest work. Without it, every incident re-argues the same question with different people and reaches different answers. ## A blanket refusal is also a decision "We never discard anything" sounds like the conservative position and is frequently the expensive one. A frozen share holds a growing backlog whose oldest entries age towards the edge of the retained window; when they cross it they are gone without anyone choosing that. The trade a principal is actually making is between: - the **known, bounded** loss of one record, captured and describable, and - the **unknown, unbounded** loss of everything behind it if the wall stands too long. Stated that way, a policy that permits a captured discard on tolerant streams is the conservative one, and a policy that forbids all discarding must come with enough retention that waiting is genuinely safe. ## What has to survive the incident - The captured record itself, somewhere durable and findable by whoever asks later. - A short record of the decision: which stream, which exit, who authorised it, what was captured and where. - Whether the record was reprocessed afterwards or not, because that single fact decides whether a downstream gap exists. - A named owner for anything deferred, with the date the deferral expires. None of that is expensive during the incident if the form exists beforehand, and all of it is impossible to reconstruct afterwards. ## Where this policy stops It covers who may destroy work and what must be preserved when they do. It does not set alert thresholds, page severities or reliability targets — those are different documents with different owners, and merging them produces a policy nobody reads. Keep this one to the decision it exists for: the authority to lose a record, and the evidence that must survive it.

  • Why pre-authorise the reversible exits rather than requiring approval for everything?
    Because an approval path that covers every action makes the fastest irreversible action the tempting one. If setting a record aside or restarting a reader needs a signature at 03:00, an engineer under pressure will reach for whatever they are already permitted to do. Pre-waiving reversible steps removes that pressure and keeps the approval where it belongs — on destruction, not on containment.
  • Where should the per-stream loss classification actually live?
    With the stream's own definition, alongside its owner and its purpose, so it is discoverable before anyone needs it and is reviewed when the stream changes. A classification that lives only in an incident document is found after the decision it was meant to inform. The point is that the operator can read it in the first minute without waking the team that produced the data.

saying these in an interview costs you the question

  • Decides during the incident who is allowed to discard a record
  • Treats a blanket refusal to discard as a cost-free default
  • Applies one loss rule to every stream regardless of content
  • Leaves what was set aside with no named owner or deadline
  • Counts a chat message as the durable record of what was lost
  • Folds alert thresholds and reliability targets into the same policy