A node is refusing writes because its data volume is full — in what order do you try the fixes, and which one cannot be undone?
answer
- order by cost of being wrong
- space first, history last
- moving data needs space too
- the fastest rung is the final one
- never delete files under the node
basics
~20 sWork from reversible to irreversible: add space, move the pressure elsewhere, clear non-store files, release bytes something pins, and only then shorten the retention window. That last rung restores writes fastest and destroys history permanently.
solid answer
~40 sOrder the ladder by what it costs if you are wrong. First **buy space** — grow the device or add one where the platform and host allow it online; nothing is lost. Then **move the pressure**: relocate a stream to a node with room, or pause the writer producing the growth, remembering that moving data consumes space and I/O on both ends. Then **clear bytes that are not history** — archived diagnostics, dumps, artefacts sharing the device. Then **release what pins history**: retire a subscriber nobody owns where its unread bytes are ineligible for removal, or lift a rule holding records that no longer needs to. Only then **shorten the retention window**, which frees space fast and destroys the replay budget for good. Never hand-delete segment files under a running node.
go deeper
Recall the two ends of the ladder: more space costs money and loses nothing, while cutting how long records are kept is instant and permanent. Never remove the store's files yourself.
Explain why relief that keeps history is preferred, and why the cut may not free space immediately — the pass is scheduled, works on whole closed segments and never touches the one being appended to.
Show the ordered ladder with its caveats: online growth is not universal, relocation costs space on both ends, retiring a subscriber loses its recorded progress, and the cut is the one irreversible rung.
Decide the policy before the night it matters: who may authorise destroying history, what the post-incident record must state about it, and what bounding the device would have cost instead.
## The principle: reversible relief first Every rung of this ladder restores writes. They differ in what they cost if the decision turns out to be wrong, and under incident pressure that is the only ordering that matters. The rung that works fastest is also the rung that cannot be taken back, which is exactly the trap: the pressure to act rewards the one action you can never reverse. ## The ladder 1. **Buy space.** Grow the data volume, or give the store an additional one where it can append to more than one device. Where the platform and the host support online growth this restores writes with nothing lost. It is not universal — some designs require moving data rather than resizing beneath a running node — so know in advance which you have. 2. **Move the pressure off this node.** Relocate a stream, or part of one, to a node with room, or pause the writer producing the growth while you work. Caveat worth saying aloud: **movement consumes space and I/O on both ends**, and starting a large relocation onto a nearly-full source can make the incident worse before it makes it better. 3. **Clear bytes that are not history.** Archived diagnostics, crash dumps, old artefacts and anything else sharing the device. This is often the fastest safe win and it costs nothing anybody will miss. 4. **Release what pins history.** Retire a subscriber that nobody owns, where its unread bytes are what keeps segments ineligible; lift a rule suspending removal that no longer needs to apply. Be honest that this rung is only *partly* reversible — a retired subscriber's recorded progress is gone, and re-creating it does not restore where it had got to. 5. **Shorten the retention window.** The span of history still readable *is* the replay budget, and cutting it converts that budget into free space. It is the fastest rung and the only permanently destructive one. ## Why the cut is last | Rung | How fast | Reversible | What it costs | |---|---|---|---| | Grow or add a device | Minutes to hours | Yes | Money, and support for online growth | | Move a stream or pause a writer | Slow, and worse first | Yes | I/O and space on both ends, writer downtime | | Clear non-store files | Immediate | Effectively | Nothing of value | | Release pinned bytes | Fast once found | Partly | A retired subscriber's recorded progress | | Shorten the retention window | Fastest | **No** | History, permanently, for every reader | Three reasons the cut sits at the bottom: - **It destroys exactly what the incident needs.** The investigation that follows will want the records from the period that caused it, and any rebuild will want to re-read history. Cutting the window removes both, and removes them before anyone has written the timeline. - **It reduces everyone's margin going forward.** The new, smaller window is how long any reader may now be down and still catch up. You have relieved today's incident by shortening the fuse on the next one. - **It may not even act promptly.** Removal is a background pass over whole closed segments, it never takes the active segment, and it cannot touch bytes something pins. Paying the permanent price and still watching the volume stay full is the worst outcome available, and it is common enough that it is a standard interview probe. ## The thing to never do **Do not delete segment files by hand under a running node.** The store keeps its own view of what exists, which segment is active and where the earliest readable position is. Removing files beneath that view invites a node that serves history with holes in it, a node that crashes on the next read of a file it believes exists, or — the worst case — the removal of the file being appended to. Depending on how the platform replicates, the divergence may also propagate: some designs will notice and re-copy, others will treat the node as inconsistent. And it may not even free space, because a file that is still held open returns nothing to the filesystem when it is unlinked. Use the platform's own removal path, which updates that view and moves the earliest readable position as part of the same operation. ## Afterwards Restoring writes is not the end of the incident. Raising the bound back afterwards does not bring removed records back — they are gone — so the post-incident record should say exactly what history was destroyed and when, because someone downstream will assume it is still there. Then fix the cause rather than the symptom: a device whose streams are individually bounded, with the sum plus headroom fitting the device, is the difference between this being an incident and being a recurring one.
- Why is shortening the retention window the last rung rather than the first?Because it is the only rung that destroys something permanently. It removes the history the incident review and any replay will want, it shortens how long every reader may now be down and still catch up, and it may not even free space promptly given segment granularity.
- What actually goes wrong if an operator deletes segment files by hand?The store's own view of what exists, which segment is active and where the earliest readable position sits no longer matches the device. The node can serve history with holes, fail on the next read, or lose the file it was appending to — and space may not return if the file is still held open.
- Relocating a stream to a node with room is listed second — what is the catch?Data movement consumes space and I/O on both ends. Starting a large relocation from a nearly-full node can push it further over before it relieves anything, so check headroom and move the smallest thing that buys enough room first.
- Does raising the bound back after the incident undo the cut?No. Records the pass already removed are gone from this store. Raising the bound only changes what survives from now on, which is why the post-incident record must state precisely which history was destroyed.
A flooding basement gives you three options: bail water, open a drain, or throw out the boxes of paper records stacked on the floor. Throwing out the boxes clears the most floor space fastest and is the only choice you can never take back — and it is always the box you needed.
saying these in an interview costs you the question
- Reaches for the retention cut first because it is fastest
- Deletes segment files by hand to free space quickly
- Assumes raising the bound afterwards restores removed history
- Starts a large relocation from a nearly-full device without checking headroom
- Believes the cut frees space the instant it is saved
- Treats retiring a subscriber as fully reversible