Reapplying the previous revision put the old ledger writer back in two minutes, but its bad rows are still there - why?
answer
- two different kinds of thing
- the spec describes processes only
- effects already emitted survive
- old code, newer rows
- repair the data separately
basics
~10 sA revision records the desired spec, so reapplying one changes which version runs; it says nothing about rows already committed to a store outside the workload. The spec reverts, the written data does not.
solid answer
~50 sThe revision is a description of what should be running - content reference, copy count, settings, checks. Reapplying it replaces the processes, and that is all it can do, because nothing in the spec describes a ledger, a queue or another service. For an hour the bad version committed rows over the network into a store that keeps its own state, and those rows are effects it already emitted, not state the copies were holding. Replacing the copies discards only what lived inside them - the throwaway writable layer. Everything that left the boundary survives: committed rows, published messages, calls made to downstream services, files written into a volume, and any in-place change the version made to the store's own shape. So the revert ends the bleeding; repairing the data is separate work with its own plan.
go deeper
Remember the asymmetry: going back changes which version runs, and nothing else. Data that was already written to a database or a queue is untouched by it.
Explain why: the revision is a description of what should run, the store holds what already happened, and replacing copies only discards what lived inside them.
Demonstrate the operating instinct - before reverting, ask what the bad version already emitted, and treat the revert as the first of two workstreams whenever the answer names a store or another service.
Own the consequence for design: how much a revert can actually recover is decided long before the incident, by whether a version's effects stay inside its own boundary or land in shared state.
## Two different kinds of thing The surprise here comes from treating one word, "rollback", as covering two categories that a container platform keeps strictly apart. - A **revision** is a stored copy of the declared spec: which image content to run, how many copies, what settings and checks they get. It describes **what should be running**. - A **store** - a ledger, a queue, a search index, a mounted volume - holds **what has already happened**. It is reached over the network or through a mount, it keeps its own state, and no field of the workload spec describes its contents. Reapplying a revision moves the first category and cannot move the second. That is not a gap in the platform; it is the boundary the platform is built on. The control loop's job is to make the running set of copies match a document, and a committed row is not part of any document it holds. ## What the revert covers, and what it does not | | Restored by reapplying the revision | Left exactly as the bad version left it | |---|---|---| | Which image content runs | yes | - | | Copy count, reservations, ceilings | yes | - | | Settings and files delivered to the process | yes | - | | Each copy's throwaway writable layer | discarded with the copy | - | | Rows committed to a ledger or database | - | yes | | Messages published to other services | - | yes | | Calls that already changed a downstream system | - | yes | | Files written into a volume that outlives the copy | - | yes | | An in-place change the version made to the store's shape | - | yes | The fourth row is where a plausible wrong model hides. It is true that a copy's writable layer is thrown away when the copy is replaced, and a candidate who has learned that fact sometimes generalises it into "replacing the copies discards what the bad version wrote". It discards only what the bad version wrote **inside itself**, which for a ledger writer is nothing that matters. Anything it sent across the boundary is gone from its hands and into somebody else's state. ## Which side is old, and why it bites after the revert Once the older version is serving again, it is reading rows the newer version wrote. That is the **old code meeting new data**, the backward-compatibility case, and it is the one a revert creates by construction. Three shapes of trouble follow, in rising order of unpleasantness: 1. **Tolerated.** The older version ignores fields it does not know and keeps working. Nothing to do beyond the data repair. 2. **Rejected.** The older version treats an unknown field or an unexpected value as invalid and errors on those rows. The revert has fixed one failure and produced another, narrower one. 3. **Misread.** The older version parses the row and gets a different meaning from it - the worst case, because it is silent and it writes more bad rows on top. Which of the three you get is a property of the data the bad version wrote, not of the revert, and you find out by looking at the rows rather than at the rollout. ## What the revert is actually for It is a way to stop new bad effects quickly, and it is very good at that: the spec is small, the platform already stores the previous one, and the replacement runs at the speed of the normal rollout. Treat the moment the rollback reports complete as the moment the *rate* of new damage went to zero, not as the end of the incident. Two things are still open at that point, and they belong to different people and different clocks: - **The written record.** Somebody has to decide what the bad rows mean and what to do about them - correct them, compensate them, or leave them and account for them. That is data work, and no rollout action performs it. - **The shape of the store.** If the bad version changed the store's structure in place rather than only adding rows, the older version is now running against a shape it was never written for, and the revert has quietly put mismatched code back into service. ## The habit this teaches Before you press the revert, ask one question: *what did this version already emit?* If the answer is "nothing that left its own boundary", the revert is a complete fix and you can relax. If the answer names a store, a queue or another team's service, the revert is the first of two pieces of work, and the second one starts now rather than after the incident review.
- Other than rows in a database, what else does a revert leave behind?Anything that crossed the boundary: messages published to a queue or topic, calls that already changed another team's system, notifications sent, files written into a volume that outlives the copy, and any in-place change to the store's own shape. The rule of thumb is that a revert reclaims only what lived inside the copies it replaced.
- After the revert the older version rejects rows the newer one wrote. Which compatibility direction is that, and what does it mean for the rollback?Old code meeting new data is the backward-compatibility case, and a revert creates it by construction. It means the rollback is not automatically safe: it exchanges the new version's failure for the old version's inability to read what has been written since. You find out by inspecting the rows, not by watching the rollout finish.
- The rollback finished and the error graph is falling. Is the incident over?Only the rate of new damage is. The written record still contains everything the bad version committed, and the older version is now serving reads over it. Closing the incident means deciding what those rows mean and handling them, which is separate work from the rollout that put the old spec back.
Putting the previous cashier back at the till is fast and it stops new mistakes, but it does not un-ring the sales the last one already put through the register.
saying these in an interview costs you the question
- Says rolling back undoes what the bad version already did
- Thinks committed rows live in the copies' writable layer
- Calls the incident over the moment the revert reports complete
- Assumes the revision history includes the store's state
- Ignores that the older version now reads newer rows