A daily job has produced the same dataset for three years and nothing reads it, so what makes retiring it a judgment call?
answer
- green every night, useful to nobody
- a known bill against an unknown risk
- stop producing before deleting anything
- quiet period sized by consumer rhythm
- compute line usually beats storage line
basics
~20 sEach day the job runs charges the pool for compute nobody uses, but stopping it risks a consumer nothing has recorded. The call weighs a known recurring bill against an unknown blast radius, and staging it is what makes it safe.
solid answer
~50 sThe cost side is knowable: one run's metered usage multiplied by runs per year, plus what the accumulated output costs to store. Put that number on the decision — without it, retirement is a tidiness argument and loses to every other priority. The risk side is not knowable from inside the job: establishing who still reads a dataset is a catalog and lineage problem owned elsewhere, and absence of complaints for a week says nothing about a monthly or quarterly consumer. So the judgment is about **reversibility**, not certainty. Stop producing first and leave the existing output in place, which recovers the compute line immediately and is undone by restarting the job; delete only after a quiet period long enough to cover the slowest plausible consumer. The deeper question is why it survived three years: a pool whose bill names no run makes every night of it free at the point of use.
go deeper
Recall that a job producing something nobody reads still costs money every time it runs, and that this never shows up as a failure. Green is not the same as useful.
Explain the two separate lines — compute for each run, and storage of the accumulated output — and why stopping the job recovers the larger one immediately while deleting the data is the part that cannot be undone.
Show a staged retirement: announce, stop producing, hold a quiet period sized by the consumers' cycle, then delete. Bring the arithmetic, and check whether other jobs read this output before touching it.
Argue the policy rather than the instance: owners recorded for every output, labels on every run so the bill reaches someone who can act, and production treated as a commitment with an expiry. Accept that a retirement will occasionally be wrong and make it cheap to reverse.
## Why unread output survives A job producing a dataset nobody reads is not a bug and it never fails. It runs green every night, its throughput is fine, nothing retries, and in a shared pool its money cost lands in a single invoice line that names the pool rather than the run. There is no signal anywhere that says *this was pointless*. Meanwhile the arguments for leaving it alone are cheap and always available: somebody might use it, stopping it is risky, and nobody is complaining. That asymmetry — a known, invisible, recurring cost against an unknown, vivid, one-off risk — is the whole problem, and it is why this is a judgment question rather than a cleanup task. ## Put a number on it first Before any argument about risk, make the cost concrete: 1. **One run's metered usage** — capacity held times duration, or input bytes read, depending on which meter the work sits on. 2. **Runs per period** — nightly for three years is over a thousand runs, and the compute line is almost always the larger half. 3. **Accumulated storage** — what three years of output costs to keep, per month, plus what it costs to keep replicated copies of it if they exist. 4. **The second-order cost** — anything else that reads this output in order to produce more unread output. Retirements cascade, and the cascade is often where the real money is. With a number, the conversation becomes a comparison. Without one, it is an aesthetic preference, and it will lose. ## Two different bills, with different shapes | Line | Recovered by | Typical size | Reversible? | |---|---|---|---| | Compute for each run | stopping the job | usually the larger line | yes — restart the job | | Storage of the output | deleting the data | smaller, but grows forever | no, unless a copy exists | | Reads of this output by other jobs | retiring those too | often the hidden majority | yes, per job | The important property in that table is the last column. Stopping production and deleting data are different decisions with completely different recovery profiles, and treating them as one change is the single most common way this goes wrong. ## Retire in an order that can be undone 1. **Announce and set a date.** The only cheap way to find consumers is to make them identify themselves; establishing who reads a dataset from records rather than from volunteers is a catalog and lineage capability owned outside the job itself. 2. **Stop producing, keep the output.** This recovers the compute line — usually most of the money — immediately, and is undone by restarting the job. 3. **Wait through a quiet period long enough to cover the slowest plausible consumer.** A week proves nothing about a monthly close or a quarterly report. Choose the period from the consumers' rhythms, not from patience. 4. **Then delete**, and only if the output can be rebuilt or is genuinely worthless. If the input it was built from is no longer retained, deleting the output is irreversible in a way stopping the job never was. When the output is cheap to store and expensive to recompute, keep the artefact and stop the job. When it is expensive to store and cheap to recompute, delete it and rebuild on demand. The two lines trade against each other and the answer differs per dataset. ## The organisational question underneath The interesting part of this question is not the retirement; it is the three years. A nightly run that nobody could name, in a pool whose bill nobody could split, was free at the point of use for everyone involved. Three things change that, and all of them are organisational rather than technical: - **Every output has a named owner**, recorded at the point it is created, who is asked periodically whether it is still wanted. - **Every run carries a label** identifying its owner and its pipeline, so the pool's invoice can be split back to the people who can act on it. - **Producing something is a commitment with an expiry**, reviewed on a schedule, rather than a permanent obligation acquired by accident. A lead is also expected to say what the organisation will deliberately *stop* measuring and producing, and to accept the consequence: retire enough datasets and eventually one of them will turn out to have had a consumer. The correct response to that is not to stop retiring things. It is to have made the retirement reversible and the recovery quick, and to have the number that says how much the policy saved against how much that one recovery cost. ## What a weak answer sounds like A weak answer deletes the data first, or treats a silent week as proof of no consumers, or counts only the stored bytes and never the thousand recomputations. The strongest signal in a good answer is that the candidate separates *stop producing* from *delete*, and names the quiet period in terms of the consumers' cycle rather than their own comfort.
- What do you keep after stopping the job, and for how long?The output as it last stood, and the code and inputs needed to rebuild it. Keeping both makes the step reversible by restarting the job, which is what lets you act on weak evidence. Hold them through a quiet period chosen from the consumers' cycle — a monthly close or a quarterly report, not a convenient week — and only then delete.
- How does the decision change when the output is expensive to store but cheap to recompute?Delete and rebuild on demand: the recurring line you are paying is the wrong one. Reverse it when recomputation is the expensive half, or when the input it was derived from is no longer retained, because then deleting the output is irreversible in a way stopping the job never was. The two lines trade per dataset.
- Why did nobody notice for three years?Because nothing surfaced it. The pool's invoice named the pool rather than the run, so no team saw a line for it; the job never failed, so no health signal complained; and the output had no recorded owner who could be asked whether it was still wanted. Each night of it was free at the point of use for everyone involved.
saying these in an interview costs you the question
- Deletes the stored output first and stops the job second
- Treats a week without complaints as proof of no consumers
- Counts only the stored bytes and ignores the daily recompute
- Argues for retirement with no number attached to the decision
- Assumes a shared untagged pool would have surfaced the waste
- Never checks whether other unread outputs are built from this one