A retired pipeline's nightly snapshot job is still running - why does that waste grow every month, and what bounds it?
answer
- a stock, not a rate
- every run adds and nothing removes
- flat bills mean something is pruning
- steady state is delta times retention
- stop the flow before judging the stock
basics
~20 sBecause the charge is a stock, not a rate: every run adds stored bytes and nothing removes them. A retention rule bounds it, at roughly the nightly change volume times the days retained. Without one, the stock rises for as long as the job runs.
solid answer
~40 sAn idle machine costs the same each month; a running snapshot schedule costs more each month. The charge is on **bytes stored**, so it is a stock built from a flow: every run adds the blocks that changed since the last one, and deleting nothing means the stock only rises. Two things bound it. A **retention rule** caps the stock at roughly the nightly change volume times the days kept - that is the only reason a healthy backup bill is flat rather than rising. And stopping the schedule caps the flow at zero, though the stock already built stays billable until it is deleted. Note the direction here: a snapshot with retention protects against a mistake, which is exactly why you decide whether it should exist rather than deleting the chain reflexively.
code
pseudocode · 16 linesstoredBytes = baseSize
kept = empty list
for each night in schedule:
delta = blocksChangedSinceLastRun()
storedBytes = storedBytes + delta
append(kept, { night: night, size: delta })
if retentionDays is set:
for each snap in kept where age(snap) > retentionDays:
storedBytes = storedBytes - snap.size
remove(snap, kept)
# no retention set: nothing is ever removed,
# so storedBytes only ever increases
monthlyCharge = storedBytes * ratePerGiBStoredgo deeper
Recall that snapshot charges are for bytes stored, so a schedule nobody stopped keeps adding to the bill every night even though each run looks trivial.
Explain the stock-and-flow arithmetic: the steady state is roughly the daily change volume times the days retained, and with no retention rule there is no steady state at all.
Show the order of operations - stop the schedule, establish what the chain protects, set retention and let it converge, keep one restorable copy while the purpose is in doubt.
Argue for retention as a default at creation rather than a cleanup: any artefact produced on a schedule with an unchosen retention becomes a rising cost with no operational signal attached to it.
## Waste that accumulates is a different animal Most orphaned spend sits still. An abandoned volume costs the same in December as it did in June. A snapshot schedule does not: it costs more every month it survives, because you are being billed for a **stock of stored bytes** and something is still adding to it. This is the class of waste that is smallest when you first see it and largest when you finally deal with it, which is precisely backwards from how attention gets allocated. ## The arithmetic of a chain Snapshots are usually incremental in effect: the first one stores essentially the whole source, and each later one stores only the blocks that changed since the previous one, while the platform keeps whatever older blocks are still needed for any snapshot you still hold. Providers differ in how they account for this, but the shape is the same everywhere. Take a source that changes by some roughly constant amount each night - call it the **daily delta** - and a chain that is never pruned: - Month 1: the base, plus thirty deltas. - Month 2: the base, plus sixty deltas. - Month n: the base, plus thirty n deltas. It is linear and it does not flatten while the source keeps changing and nothing is deleted. Now add a **retention rule** of R days: every night one snapshot is created and the one that has aged past R is deleted, so the stock settles at roughly the base plus R deltas and stays there. That steady state is the whole reason a healthy backup line is flat. The numbers here are invented to make the shape visible, not taken from anyone's price list, and the assumption they rest on is a constant daily delta - a source that suddenly churns harder moves the steady state up with it. One edge worth stating honestly: if the retired source stops changing at all, the deltas approach zero and the stock stops growing by itself. That is the exception rather than the reassurance - most sources that are still attached to something still change. ## Why nobody notices | Property | Effect | |---|---| | Each run is individually tiny | No single event is large enough to investigate | | The job succeeds every night | Monitoring stays green; success is what is being reported | | The charge lands in a large shared storage line | It is spread across a dimension that is legitimately large | | Nothing fails when the stock grows | There is no operational signal at all, only a financial one | The combination is why this is found by reconciliation - matching schedules and stored artefacts against the workloads they were created for - rather than by anyone noticing. ## The same shape elsewhere The pattern is not specific to snapshots. Any artefact produced on a schedule and kept under a default nobody chose behaves this way: build outputs, exported extracts, and telemetry retained at whatever period the platform set when it was switched on. The point this leaf owns is narrow and worth stating exactly: **the cost exists because a decision was never made.** What that telemetry is actually worth keeping, and at what resolution, is a genuine judgement that belongs with the people who own the signal - the reconciliation only establishes that nobody has ever made it. ## What to do with the stock you find 1. **Stop the flow first.** Disable the schedule. This is reversible, it is the part that grows, and it needs no decision about the existing data. 2. **Establish what the chain protects.** If its source is gone and nothing has ever restored from it, it protects nothing. If the source is live, you are looking at an unbounded backup, not an orphan, and the answer is a retention rule rather than a deletion. 3. **Set a retention rule and let it converge.** The stock falls to the steady state on its own, which is far easier to get agreement on than a bulk delete. 4. **Keep one restorable copy while you decide**, if anything about the chain's purpose is unclear. Deleting the only backup of something that mattered is a much more expensive mistake than a month of the storage charge. One boundary to keep straight: deciding that surviving data should **age into a cheaper class** is a separate question with its own owner. This leaf's question is whether the bytes should exist at all, and a chain that nothing restores from should be deleted rather than made cheaper to hoard.
- Deleting the source volume - does that shrink the snapshot charge?No. Snapshots are separately stored objects, not pointers into a live volume, so the chain keeps billing after its source is gone. It also stops growing at that point, because nothing is changing any more, which is why an orphaned chain typically sits at whatever size it reached.
- Why not just delete every snapshot the sweep finds with no live source?Because a snapshot with retention is the thing that protects against a mistake, and a chain whose source was deleted last week may be the only copy of it. Confirm that nothing has restored from it and that no owner claims it, keep one copy while the question is open, and delete the rest.
saying these in an interview costs you the question
- Incremental snapshots are too small to matter.
- Deleting the source volume deletes its snapshot chain.
- The backup line is flat, so retention must be fine.
- A green nightly job means nothing is being wasted.
- A replica of the volume makes the snapshots redundant.