For a data-owning workload you run yourself, what must a backup cover that redeploying the image never restores?
answer
- code and data recover differently
- the image restores the process only
- a copy of the data, kept elsewhere
- untested backup is only a claim
- measure the restore time and the gap
basics
~20 sEverything written after the image was built. An image restores the process and none of its data, so a data-owning workload needs a separate copy of the data, held outside its failure domain, mapped back to the member that owned it, and proven by an actual timed restore.
solid answer
~50 sCode and data recover by different routes. Redeploying an earlier image gives you the previous process, with every write that happened since still sitting in the backing store - a bad release that corrupted data is not undone by putting the old binary back. So a workload that owns data needs its own recovery artifact: a copy of the data itself, taken by something that understands the data's consistency rules, stored somewhere that does not fail with the original, retained long enough to cover the case where you discover the damage a week later, and recorded with which member it belonged to. The last requirement is the one teams skip: a backup nobody has restored is an untested claim. Rehearse it, time it, and measure how far behind the newest copy is - that gap is the data you lose.
go deeper
Recall that the image contains code, not data, and that a workload's data needs its own backup kept somewhere else.
Explain why a rollback of the image leaves the data where it was, and what a copy of a device taken during writes actually captures.
Show that you run restores, not just backups: the rehearsal, the two numbers it produces, the member mapping, and the retention that outlasts discovery.
Own the durability obligation as a policy - what the organisation promises for data it hosts itself, how often the drill runs, and what that cost implies for buy-or-run.
## The image restores the process, and nothing else The image is the code, its dependencies and its default configuration, fixed at build time. Everything a data-owning workload cares about was written afterwards, by the running process, into storage that outlives the instance. Two consequences follow, and both get missed: - **Rolling the image back does not roll the data back.** The previous build starts against the current data, including whatever the bad build wrote. If the release corrupted rows, truncated a file or wrote a value the old code cannot parse, redeploying makes the situation worse, not better - now you have old code meeting data that moved forward. - **Rebuilding the platform does not rebuild the data.** A cluster can be re-created from its declared specs in minutes. That gets every workload back as an empty shell. The recovery clock for a data-owning workload is set by the data, and nothing about the platform's declarative story shortens it. ## What a real backup of a data-owning workload covers 1. **A copy of the data, taken in a way the data engine can accept.** A copy of a device taken while writes are in flight captures whatever was on the device at that instant - equivalent to the state after a power cut. Some engines recover from exactly that and are fine with it; others need the writer quiesced, or need the copy taken through their own backup mechanism, which flushes and marks a consistent point. Which of the three applies is a property of the engine, and knowing which one is part of owning the workload. 2. **Storage outside the original's failure domain.** A copy that lives on the same device, in the same zone, or under the same account-level blast radius as the original is not a backup; it is a second way to lose the same bytes at the same moment. 3. **Enough retention to outlast discovery.** Corruption is often found long after it happened. A rotation that keeps a day of copies protects you from a device failure and not from a bad write found on Friday. 4. **The mapping from each stored copy to the member that owned it.** Each member holds a different slice. Restoring member 2's data into member 1's store gives a workload that starts, reports healthy and answers from the wrong data - the failure with no alarm attached to it. 5. **A restore procedure someone has actually run.** ## What each artifact actually gives back | Artifact | Restores | Does not restore | |---|---|---| | The image | the process, its dependencies, its defaults | any write made after the build | | The declared workload shape | copy count, identity, storage requests, order | the contents of any store | | A device-level copy of the store | the bytes as of that instant | consistency the engine did not arrange, or the member mapping | | An engine-level backup | a consistent data set the engine will open | writes accepted after it was taken | ## The rehearsal is the only part that proves anything A backup job that reports success proves that bytes were written somewhere. It does not prove they can be read back, that the format is one the current version opens, that the credential for the storage still works, that the copy is complete, or that anyone knows the steps. The rehearsal is what converts the claim into a fact, and it produces two numbers that the job alone never produces: - **How long the restore takes**, end to end, including transferring the data and bringing the workload back up against it. If the data is 400 GB and the path back sustains 100 MB per second, the transfer alone is over an hour before anything starts - the process start is seconds either way, and it is not what you are waiting for. - **How much data you lose**, which is the gap between the newest usable copy and the moment of failure. Nightly copies mean up to a day of accepted writes are gone. If that is unacceptable, the answer is more frequent copies or a continuous stream of changes that lets you recover to a chosen point, not a better nightly job. Run the rehearsal into a scratch environment, against a copy chosen by the process you would use in an incident rather than the one you know works, and have someone who did not build the system follow the written steps. What the rehearsal catches is rarely the data: it is an expired credential, a missing decryption key, a restore path that needs a capacity nobody provisioned, or a step that lives only in one engineer's head. ## The judgment this leads to All of the above is ordinary work, and it is *your* work the moment you decide to run the store yourself. That is the honest cost side of the buy-or-run decision: not the compute, which the platform handles well, but the durability obligation, the retention policy and the drill that has to stay current as the workload changes.
- A copy of the backing device is taken while the workload is accepting writes. What exactly have you captured?The bytes as they stood at that instant, including partially applied writes - the same state the engine would find after a power cut. Some engines replay their own journal and open it cleanly; others need the writer quiesced or the copy taken through their backup mechanism. The copy is useful only if you know which case applies to your engine.
- Why is retaining one day of copies not enough for most data-owning workloads?Because it only covers failures you notice immediately. Corruption from a bad write or a bad release is often found days later, by which time every retained copy already contains it. Retention has to outlast discovery, which means keeping a ladder of older copies even though the recent ones are the ones you usually reach for.
- What does a rehearsal typically catch that the backup job never reports?Almost never the data itself. It catches an expired credential for the storage, a missing decryption key, a restore that needs capacity nobody provisioned, a format the current version will not open, and steps that exist only in one engineer's memory. It also produces the restore duration, which is the number incident planning actually needs.
saying these in an interview costs you the question
- Believes redeploying the previous image rolls the data back
- Calls a job that has never been restored a backup
- Keeps every copy of the data beside the original
- Thinks a device-level copy is automatically consistent for the engine
- Assumes any stored copy can be restored into any member
- Retains one day of copies and calls corruption covered