skip to content

Half way through a roll the stored format version was raised; why does reinstalling the previous release now recover nothing?

level: seniorimportance: must knowfreq 55%

answer

  1. name which rollback you mean
  2. new binaries read old data, not the reverse
  3. records written after the raise
  4. the abandon point
  5. after it, restore rather than reinstall

basics

~20 s

Because the previous binaries cannot read records written in the newer layout. Reinstalling them recovers nothing written since the stored format version moved; the cluster has passed its abandon point, and the only way back is a restore or a second cluster, not a reinstall.

solid answer

~50 s

"Rollback" during a roll means two different things. Reinstalling the previous release on a member is cheap while the cluster is still writing the old layout — the binaries are the only thing that changed. Once the **stored format version** moves forward, every record written after that moment is in a layout the previous binaries do not understand, so putting them back leaves a member that cannot read its own data. That step is the **abandon point**: the last moment at which a reinstall is a real answer. The practical discipline is therefore to know, before starting, exactly which step in your roll is the abandon point, to keep the whole roll on the near side of it while you gather evidence, and to recognise that after it the recovery path is a restore or a switch to another cluster — a far more expensive plan that has to exist in advance.

go deeper

for a junior

Recall the asymmetry: a new release is built to read data the old one wrote, but the old release was built before the new layout existed and cannot read it.

for a middle

Explain that records written after the stored format version moves are in a layout the previous binaries do not parse, so a reinstall leaves a member that cannot read its recent data.

for a senior

Locate the abandon point for the change in front of you, keep the roll on its near side while evidence accumulates, and be able to state the far-side recovery plan before crossing.

for a principal

Treat crossing it as a change class of its own — its own approval, its own recovery target, and a standing rule that no cluster crosses without a tested way back existing on paper.

## The word "rollback" is doing three jobs Before anything else, separate them, because an interview answer that uses the bare word is unanswerable: 1. **Reinstalling the previous release** on one or more broker nodes — the subject here. 2. **Backing out a half-finished data move** between members — a membership subject, not a version one. 3. **Reapplying a previously kept deployment revision** on the platform the cluster happens to run on — a container-platform subject, and notably not the same act, because it restores a deployment specification rather than the state on disk. Only the first is what "can we go back?" means during a version change, and it has a hard expiry. ## What the stored format version is A broker node writes records down in some layout: framing, headers, whatever accompanies the payload, and whatever indexing lets a reader find a position in it. The **stored format version** names that layout. A release usually ships able to *read* several past layouts and to *write* one, and which one it writes is a setting held deliberately behind the binaries for exactly the same reason the agreed internal version is. The asymmetry is the whole point: - new binaries reading old data — supported, by design, across a bounded range of past layouts; - old binaries reading new data — **not** supported, because the previous release was built before the layout existed. ## Why a reinstall stops working While the cluster still writes the old layout, a member that goes back to the previous release finds every record on its disk in a layout it has always understood. It rejoins, catches up and serves. Nothing about the upgrade has left a trace it cannot cope with. The moment the stored format version is raised, the members begin appending records in the new layout. Now: - records written **before** the raise are still readable by the previous release; - records written **after** it are not; - so reinstalling recovers nothing that was written since the raise, and the member typically refuses to start, refuses to serve the affected data, or serves an error to readers that reach it. This is why the honest phrasing is not "rollback is impossible" but "a reinstall no longer returns you to a working cluster holding your current data". ## The abandon point The **abandon point** is the last step of the roll at which reinstalling the previous release still leaves the cluster able to read everything on disk. Typically it sits immediately before whichever of these your platform does first: | step | still abandonable? | |---|---| | binaries on some members | yes | | binaries on every member | yes | | agreed internal version raised | in practice, no — old binaries cannot rejoin | | stored format version raised | no — new data is unreadable to the old release | | held-off capability enabled cluster-wide | no — new state presumes it | The operational consequence is that the only question worth asking during a roll is: **can this still be abandoned, and at which wave?** Everything else — which member to take next, how long to wait — is procedure. This is the one thing that changes irreversibly. ## Planning around it 1. **Locate the abandon point before you start.** Read which steps your release ties together: some raise the stored format version automatically with the agreement, some keep them as separate operator actions, and the difference decides how much of the roll is reversible. 2. **Stay on the near side while you gather evidence.** All members new, agreement and format still old, soaking under real traffic, is the most informative reversible state available. 3. **Have the far-side plan written down.** After the abandon point, going back means restoring from a backup, or cutting readers and writers over to another cluster that was never upgraded — including everything a restore does not bring with it. If that plan does not exist, crossing the abandon point is not a decision, it is a hope. 4. **Cross deliberately.** One change, announced, with someone watching, and not bundled with anything else you would later want to blame. ## Where platforms differ The existence of an abandon point is common to the class; its position is not. On platforms that keep records in an on-disk log served directly to readers, the stored format version is explicit and operator-controlled, and it is usually the abandon point. On platforms whose storage is internal and whose records are deleted on acknowledgement, there may be no long-lived record layout to strand — but the same closure arrives through cluster metadata written in a newer form. On a rented cluster the provider owns the format entirely and the tenant has no reinstall to attempt at all; their equivalent of this question is whether the provider offers a point-in-time restore, and to when. The lesson survives all three shapes: some step in the change closes the cheap way back, and you should know which one it is before you take it.

  • Is data written before the stored format version was raised also lost?
    No. Those records are in the layout the previous release has always read, so they remain readable after a reinstall. What is lost is everything appended since the raise. That partial readability is often worse operationally than a clean failure, because a member may start, serve old data and error on recent data.
  • If the second pass is already done, what does going back actually involve?
    Not a reinstall. It means restoring from a backup taken before the abandon point and accepting the loss of everything since, or cutting producers and readers over to a separate cluster still on the previous release. Both are exercises with their own recovery targets, and both have to have been planned before the roll started.
  • Why do releases support reading old layouts but not writing them indefinitely?
    Reading old layouts is a compatibility cost paid once in the reader; keeping the ability to write every past layout would freeze the format permanently and block the improvements the format change existed for. So releases read backwards across a bounded range and write one current layout, which is what makes the upgrade one-way.
  • How should the abandon point change how a roll is scheduled?
    It splits the roll into a reversible phase and a committed one, and they deserve different treatment: the reversible phase can run at normal pace with ordinary review, while crossing the abandon point is scheduled as its own change, with the far-side recovery plan attached and someone watching the signals it could affect.

You can keep the old key while you change the lock's cylinder, but the moment new documents are filed in the new cabinet, taking the old cylinder back does not open it.

saying these in an interview costs you the question

  • Uses the bare word rollback without saying which act is meant
  • Assumes reinstalling the previous release always returns a working cluster
  • Thinks the previous release can read data written in a newer layout
  • Believes crossing the abandon point can be undone by restarting members
  • Plans the restore path only after deciding to go back
  • Assumes the roll can be abandoned at any point right up to the last node