Across an estate, how long should a store keep earlier versions of a secret before destroying them?
answer
- history is an undo, not an archive
- the cost is a credential archive
- size the window to detection time
- tighter window for wider reach
- destroy on a schedule, not memory
basics
~20 sLong enough to undo a mistaken write and no longer: days rather than quarters, with destruction on a schedule rather than on memory, and a tighter setting for names whose single value reaches a large part of the estate.
solid answer
~40 sThere is no correct number, only a trade to defend. History is an undo mechanism, and the useful window is set by how long a wrong write takes to be noticed — usually inside a deploy cycle, so days cover it with margin. What history costs is a readable copy of every credential the estate has ever used, sitting behind an ordinary read right on the name. So I would set a short default with destruction scheduled rather than manual, make it tighter still for the widest-reach names, allow a documented longer window only where something genuinely needs it, and treat an exposure as a trigger to destroy that version explicitly instead of waiting it out. Who is permitted to destroy is a separate rights question.
go deeper
Recall that a store can be configured to keep a limited number of earlier values, and that keeping them forever is a choice with a cost.
Explain both sides of the trade: history is what makes a mistaken write cheap to undo, and it is also a readable copy of every credential the name has held.
Set the window against how quickly a wrong write is noticed, retire earlier versions on replacement, and destroy on a schedule rather than when someone remembers.
Own the estate-wide standard: one short default, documented exceptions, tighter settings where a single value reaches the most, and three questions the store itself can answer on demand.
## What history is for, and what it costs Value history exists for one everyday reason: someone writes the wrong value to the right name, and the cheapest fix in existence is to read the previous version and write it back. That is a real operational benefit and it is worth paying something for. What it costs is usually left unstated. Retained history is a readable copy of every credential that name has ever held, reachable by anyone holding the ordinary read right on it — not by an administrator, by a consumer. Multiplied across an estate, unbounded history turns the store into the most complete credential archive the organisation owns, and one that no rotation has ever cleaned out. So the question is not "is history good", it is "how long", and the answer is a number you have to be able to defend. ## The inputs that set the number - **How long a wrong write takes to be noticed.** This is the dominant input, because it is the only thing the window actually buys. If a mistaken value surfaces within a deploy cycle — minutes to hours in most estates — then a window of a few days carries a large margin already, and a quarter carries none of the benefit and all of the cost. - **Whether anything else holds the value authoritatively.** Where the correct value exists elsewhere, history is a convenience rather than the only route back, and the window can be short. - **What the name reaches.** A credential one batch job uses and a credential a whole environment authenticates with are not the same asset. The second deserves the shorter history, even though instinct reaches the other way. - **Why the value was replaced.** A scheduled replacement and a replacement forced by an exposure want opposite treatment of the old version. - **Whether the store gives you a usable retirement operation at all.** Designs differ: some let a read refuse anything below a chosen version, so a whole span retires in one move; others only enable and disable individual versions, which makes retirement per-version work; some names can be configured to keep nothing. A standard that assumes the first and runs on the second silently does nothing. ## A defensible default 1. **Retire on replacement.** After a write, earlier versions should stop being servable straight away, so a plain read is the only read that works. Undo does not need them servable — a recovery is an operation, not a read. 2. **Keep a small number of versions, for a short window.** Size it against detection time with margin, and write the assumption down: "a wrong write is noticed inside a deploy cycle, so this window is days." 3. **Destroy on a schedule, not on memory.** A retention that depends on someone remembering to clean up is not a retention, and it is the most common way an estate ends up with years of history. 4. **Tighter on the widest-reach names.** For the handful of values whose exposure would be an estate-level event, keep minimal or no history and accept that a mistaken write there is a replacement rather than an undo. 5. **Destroy explicitly after an exposure.** If the value was replaced because it got out, the exposed version should be destroyed as part of closing the incident rather than left to age out — and that is only half the job, because the credential also has to stop being accepted where it was used. ## What you standardise and what you leave alone Standardise the default, the scheduled destruction, and the requirement that a longer window be documented with a reason. Leave the exceptions to the teams who own the names, because they are the ones who know what else holds the value. Who may perform a destruction is a separate question about rights and duties, and it is decided in the access rules rather than in the retention number. ## What you must be able to answer afterwards A retention standard is only as good as the questions it lets you answer quickly: - For any name, how many earlier values are currently held, and how many are still servable. - When the oldest of them will be destroyed, and by what mechanism rather than by whom. - Which names are exceptions, and the recorded reason for each. If those three cannot be answered from the store itself, the standard exists on paper only — and the review three months after a replacement will find exactly what it always finds. ## The failure mode to name out loud The common estate-level failure is not a bad number. It is no number: history defaults inherited per name, never revisited, cheap in storage and invisible in every dashboard, until a review asks whether last quarter's password can still be fetched and the honest answer is that every password since the store was installed can.
- Which names deserve the shortest history?The ones whose single value reaches the most — a shared administrative credential, anything a whole environment authenticates with. Their earlier versions are the most valuable material in the store, and the undo they buy is the least valuable, because those values are usually held authoritatively somewhere else anyway.
- What changes when a value was replaced because it leaked?The exposed version should be retired immediately and destroyed rather than aged out, since ending that value was the entire point of the replacement. Destruction is only half of it: the credential also has to stop being accepted at the system that took it, or a copy already out there still works.
- How would you get an existing estate from unbounded history to this standard?Measure first — how many versions each name holds and how many are servable — then retire everything but the current version across the board, which is reversible and low-risk. Destruction comes after, on a schedule, starting with the widest-reach names and with owners given notice that the undo route is closing.
saying these in an interview costs you the question
- Keeps history indefinitely because storage is cheap
- Sets one retention number for every name in the estate
- Assumes the exposed version disappears when a new one is written
- Leaves destruction to whoever remembers, with no schedule
- Treats a long window as free because reads are permissioned