A month-end job broke three weeks after a retired stream was deleted: what did the retirement schedule get wrong?
answer
- the slowest consumer sets the clock
- periodic work is invisible for weeks
- one full cycle inside the window
- the announcement reaches who traffic cannot
- cheap before the delete, a project after
basics
~20 sThe cut-over window and the quiet period were shorter than the interval between two runs of the slowest consumer. A monthly reader never appeared while the old stream was still there to catch it, so the first symptom arrived after the irreversible step.
solid answer
~40 sA retirement's clock is set by the **longest interval between two reads**, not by the sprint or the change calendar. Continuous consumers reveal themselves within minutes; periodic and occasional ones — month-end reconciliation, a quarterly extract, a rebuild that only runs after an incident — are invisible to any plan that spans days. The fix is to schedule the whole retirement against that slowest known cycle: run dual publication through at least one full cycle plus margin, announce the retirement with a dated deadline so the humans who own periodic work can answer, and make the quiet period long enough to contain another cycle. The asymmetry is the point: before the delete, a late discovery costs one re-point; after it, the job has to be rebuilt from whatever else still holds the data.
go deeper
Remember that some jobs only run monthly or quarterly, so a stream that looks unused all week can still have a reader. The retirement has to last long enough for that reader to appear.
Explain the scheduling rule: dual publication through at least one full cycle of the slowest known consumer plus margin, an announcement with a dated deadline, and a quiet period before the delete.
Show the asymmetry explicitly — minutes to fix a late discovery before the delete, a rebuild afterwards — and say how you would handle a stream whose slowest consumer nobody can name.
The judgment is how much delay an estate should buy for how much residual risk, who is permitted to accept that risk, and how it is recorded so the decision is defensible when a job fails months later.
## Why a monthly reader is invisible to a weekly plan Every retirement is a bet that everyone who needs the stream will show up before it goes away. The bet is settled by a **calendar**, and the calendar that matters is not the one the retirement was planned on — it is the one its consumers run on. Consumers fall into rough classes by how often they touch a stream: - **Continuous** — attached all the time. Any change is visible within minutes. - **Daily or hourly batch** — visible inside a normal cut-over window. - **Periodic and slow** — month-end reconciliation, a quarterly regulatory extract, an annual recompute. Invisible for weeks at a stretch. - **Occasional and event-driven** — a rebuild after an incident, a replay run by an analyst to answer a question, a disaster exercise. These may not run for a year and are not on anyone's schedule at all. A retirement that runs dual publication for a week and waits two days before the delete has tested only the first two classes. The checks that preceded the cut may have been entirely sound for the period they covered; the defect is in the period, not in the checks. ## Setting the clock from the slowest cycle The scheduling rule is short and unpopular, because it makes retirements slow: 1. **Find the longest cycle** among anything plausibly reading this stream — from the humans who own the periodic work, not from live traffic, because live traffic is exactly where a monthly job is absent. 2. **Run dual publication through at least one full cycle plus margin.** One full cycle means a month-end job actually runs, at least once, while both streams carry traffic — which is the only condition under which it can fail safely. 3. **Announce with a dated deadline.** The announcement is part of the mechanism, not courtesy: it is the only signal that reaches a consumer which is not currently running. It must name the replacement stream and the date of the delete, and it must ask for an answer rather than assume silence is consent. 4. **Make the quiet period contain another cycle** where the stakes justify it. After dual publication stops, a periodic consumer that runs once more will fail loudly — and be repaired for the price of a re-point. 5. **Record what you could not establish.** "No known consumer with a cycle longer than one month" is a defensible statement. "Nothing reads it" usually is not. ## The asymmetry that makes this worth the delay | When the periodic consumer surfaces | What it costs | |---|---| | During the cut-over window | Nothing. It is reading a stream that still carries everything; schedule its move. | | During the quiet period | A re-point, and a decision about where on the replacement it starts. | | After the delete | A rebuild. The records are gone; the work must be reconstructed from whatever other system still holds the data, or the period is simply missing. | The first two rows cost minutes. The third costs a project. That gap is the entire justification for a retirement schedule measured in weeks instead of days, and it is what a candidate is being asked to show they understand. ## When the slowest cycle is genuinely unknowable Sometimes nobody can say whether anything reads a stream on a yearly rhythm. Two honest responses exist, and they are not the same as waiting forever: - **Make the failure loud and reversible before making it permanent.** Withdraw the ability to read, leave the records in place, and wait. A consumer that appears now fails with an access error that is undone in seconds — a rehearsal of the delete with none of its consequences. - **Accept the residual risk explicitly**, with the owning team named and the date recorded. An organisation that will not accept residual risk never deletes anything, and an estate that never deletes anything is its own problem. ## Where platforms differ The schedule reasoning is universal; what a late discovery can be given is not. On a platform where records remain readable after delivery, a consumer found during the quiet period may still read the history it missed straight from the old stream, for as long as those records are held. Where a delivered message is gone, there was never a backlog of history to hand back, and the periodic consumer's exposure is limited to whatever it would have consumed live. Either way the delete ends the conversation, which is why the calendar has to be right before it, not after.
- The slowest known consumer runs quarterly. Does that mean dual publication must run for three months?It means one of two things must be true: either the window spans a full quarter so the job runs at least once with both streams live, or its owner confirms in writing that it has been moved and tested against the replacement. The second is much cheaper and is the reason the announcement with a dated deadline exists.
- Why is silence after a retirement announcement not evidence that nobody reads the stream?Because the people who own periodic work are frequently not the people who read the announcement, and an unattended job has no way to answer at all. Silence establishes only that nobody objected. It is a reason to keep the schedule generous, not a reason to shorten it.
- What should be written down when the delete is finally executed?Enough to make the decision auditable later: the replacement's name, the dates the window opened and closed, who was notified, the last record time and rough volume removed, and the residual risk that was accepted. When a job fails six months later this is the only remaining account of what happened to its input.
saying these in an interview costs you the question
- Sizes the retirement against the change calendar, not the consumer's
- Treats a quiet week of traffic as proof nothing reads it
- Takes silence after the announcement as agreement
- Believes re-creating the deleted name restores the records
- Plans no margin beyond exactly one cycle of the slowest job