skip to content

A job overwrote every document in a versioned object store with a corrupt render — how do you get the originals back?

level: juniorimportance: must knowfreq 68%

answer

  1. the old bytes did not go anywhere
  2. one key, a stack of objects
  3. the newest version is what a read returns
  4. recovery is promotion, not rewinding
  5. every retained version is still charged

basics

~20 s

Versioning kept each pre-overwrite copy as a previous version under the same key, so recovery is promoting that version back to current, key by key. Nothing is restored from a backup, and every retained version keeps being stored and charged.

solid answer

~40 s

In a versioned object store, writing to an existing key does not replace the bytes — it adds a new version and marks it current, leaving the previous version addressable by its own version identifier. So the corrupt render did not destroy anything: for each affected key, list its versions, pick the newest one written before the job started, and make it current again (usually by copying that version onto the key, which writes a fresh current version). Two caveats decide whether this works. Versioning is not retroactive — only writes made after it was enabled left a previous version. And it is not free: the corrupt versions stay in the history and keep costing storage until something removes them.

code

pseudocode · 12 lines
pseudocode
cutoff = timestamp when the bad job started

for each key under prefix "records/":
    versions = listVersions(key)            // newest first
    candidate = first v in versions where v.writtenAt < cutoff

    if candidate is none:
        report key as unrecoverable           // nothing predates the job
        continue

    copyVersionOntoKey(key, candidate.id)     // writes a NEW current version
    verify(readCurrent(key).checksum == candidate.checksum)

go deeper

for a junior

Recall that with versioning enabled a write adds a version instead of replacing bytes, and that a plain read returns the newest one. Knowing the old copy is still there, addressable by its version identifier, is the whole first-screen answer.

for a middle

Explain the mechanics: what makes a version current, why promotion writes a new version rather than rewinding, and why versioning is not retroactive. Be able to describe the recovery as an enumeration over keys with a cut-off time.

for a senior

Show the operational judgment: stop the writer before promoting, pick the cut-off from evidence, verify with a plain read, and size the promotion job against the outage you are allowed. State plainly that versioning is not a backup and why.

for a principal

Own the trade-off you are setting for every team: how long non-current versions are kept is a bet on how fast a bad write is noticed, priced in storage. Decide it deliberately, and decide where real backups sit outside this store's permission boundary.

## What versioning changes about a write In an object store **without** versioning, a write to a key that already exists is destructive in the plain sense: once the new object is durable, the old bytes are unreachable and the store has no record that they ever existed. The key holds exactly one object, forever, and `PUT` is `REPLACE`. Turn **versioning** on for the store and the same write behaves differently: - Each write to a key creates a **new version** of that key, identified by a version identifier the store assigns. - The newest version becomes the **current** version. A plain read of the key returns it, and a plain listing shows it. - Every earlier version remains stored and remains readable — but only if you ask for it **by its version identifier**, or list the key in a version-aware way. - The key therefore holds a stack of objects rather than one object, and the stack only ever grows on write. That is the whole mechanism, and it is what makes the corrupt re-render survivable. The job did not overwrite the scanned originals; it pushed a corrupt object on top of each of them. | | Store without versioning | Store with versioning | |---|---|---| | Overwrite | previous bytes unreachable | previous bytes kept as a non-current version | | Delete | object removed | a marker hides the key; versions remain | | Unit of recovery | whatever a backup holds | one version of one key | | Storage charged | the current object | every retained version | ## Recovering from the mass overwrite The recovery is mechanical, and it is a job rather than a button: 1. **Stop the writer first.** A re-rendering job still running will simply push a second corrupt version on top of whatever you promote. 2. **Establish the cut-off** — the moment the job's first bad write landed. Every version older than that is a candidate original. 3. **For each affected key, list its versions** and select the newest one written before the cut-off. 4. **Promote it.** The usual move is to copy that version onto its own key, which writes a *new* current version whose content is the old content. The history is not rewound; it is extended. 5. **Verify with a plain read**, not a version-targeted one — a plain read is what your application does, so it is what proves the fix. Note what step 4 does and does not do. Promotion makes the good content current again. It does **not** erase the corrupt version, which stays in the history as one more non-current version. ## What versioning does not give you This is where candidates overclaim, so be precise about the boundaries: - **It is not retroactive.** Enabling versioning today leaves yesterday's overwrite unrecoverable. Nothing was retained before the setting existed. - **It is not a backup.** Every version lives in the same store, inside the same account, under the same permissions. A caller able to purge versions can purge them, and a store-wide misconfiguration reaches the history as easily as the current object. A backup is a copy somewhere else, reachable with different credentials. - **It is not an undo button.** Recovery is per key and per version, so a mass overwrite means a mass promotion job, and that job needs enough throughput to finish inside whatever outage the business is tolerating. - **It does not survive a purge.** A version-targeted delete removes that version's bytes permanently. Versioning protects against overwriting and against an ordinary delete — not against someone deliberately removing versions. ## The cost you just accepted Every retained version is a stored object and is charged like one. After a mass overwrite and a mass promotion, a key that logically holds one document physically holds three: the original, the corrupt render, and the promoted copy. Two consequences follow: - Storage for a versioned store grows with **write volume**, not with the size of the live data set. A service that rewrites the same key hourly accumulates versions relentlessly. - You need a deliberate policy for **non-current versions** — an age-based rule that removes them after some number of days, chosen so it is longer than the time it realistically takes to notice a bad write. Choosing that age is the trade-off: shorter is cheaper, longer is safer, and 'how long until someone notices' is the number that actually matters. The interview-grade summary is that versioning converts a destructive write into an additive one, which turns a data-loss incident into a data-cleanup job — at the price of storing everything you ever wrote until you tell the store to stop.

  • Versioning was switched on the morning after the corrupt job ran. What can you recover?
    Nothing from before it was enabled. Versioning retains previous copies only for writes made while it is on, so the corrupt render is simply the key's only object. Recovery then depends on something outside the store — a backup, or re-deriving the document from the source system that produced it.
  • After promoting the good version, what does the key's version history look like?
    Longer, not shorter. Promotion writes a new current version holding the old content, so the key now carries the original, the corrupt render and the promoted copy. The corrupt version is still stored and still charged until an explicit removal or an age-based rule on non-current versions clears it.
  • Why is there no single store-wide undo for this?
    Because versions are per key and the store has no notion of 'the state at 10:00' across keys — each key's stack is independent. Recovery is therefore an enumeration job over the affected keys. A store that does offer a point-in-time rewind is implementing that job for you, not replacing the mechanism.

saying these in an interview costs you the question

  • Thinks an overwrite replaces the bytes even when versioning is on
  • Assumes versioning was already on before the bad job ran
  • Calls previous versions a backup of the store
  • Expects retained versions to cost nothing
  • Believes enabling versioning now recovers yesterday's overwrite